Senior Cloud/Platform Operations Engineer
Bain CapitalAbout the role
Title: Senior Cloud / Platform Operations Engineer
Location: Boston, MA
BAIN CAPITAL OVERVIEW
With approximately $225 billion of assets under management, Bain Capital is one of the world’s leading private investment firms. We create lasting impact for our investors, teams, businesses, and the communities in which we live. Over four decades we have strategically grown our platform to focus on Private Equity, Growth & Venture, Capital Solutions, Credit, and Real Assets. Today, our team includes 1,985+ employees in 24 offices on four continents.
We partner differently to help people and companies embrace possibility and realize potential. Founded as a private partnership in 1984, we have fostered a culture of innovation, entrepreneurialism, and agility, empowering our people to define and own their career trajectories. Today, our partnership approach enables us to pursue strategic growth, build enduring relationships with a robust external network, and collaborate across our integrated platform to connect the deep and diverse expertise that unlocks breakthrough insights.
Our people are the heart of our advantage. Colleagues at all levels have a seat at the table as they tackle business challenges with a principal investor mindset. By asking incisive questions, respectfully challenging one another, and remaining intellectually agile, we work together to achieve exceptional outcomes.
For more information visit: Bain Capital
DESCRIPTION
We are seeking a Senior Cloud / Platform Operations Engineer to play a key role in the day-to-day operations of our AWS and Kubernetes platforms. This is a hands-on technical role focused on ensuring the reliability, security, and operational excellence of our cloud infrastructure.
Unlike traditional platform engineering roles centered on building internal developer platforms or delivering new features, this position is focused on operations, service delivery, and execution. You'll be responsible for maintaining production Kubernetes clusters, enforcing AWS governance, responding to operational issues, and supporting a high volume of internal requests. As well as participating in an On Call off hours rotation.
The ideal candidate thrives in production environments, enjoys solving complex infrastructure problems, and is passionate about building stable, secure, and well-governed cloud platforms.
Responsibilities
Cloud & Platform Operations
- Own the day-to-day operations of AWS and Kubernetes environments, ensuring high availability, reliability, and performance.
- Perform hands-on administration of production Kubernetes clusters and AWS infrastructure.
- Troubleshoot complex infrastructure, networking, and container platform issues.
- Execute platform maintenance activities including upgrades, patching, scaling, and lifecycle management.
- Continuously improve operational processes, automation, and platform stability.
Kubernetes Operations
- Operate and maintain production Kubernetes clusters.
- Perform cluster upgrades, patching, and version management.
- Troubleshoot Kubernetes control plane, worker nodes, networking, storage, ingress, and workload issues.
- Optimize cluster performance, resource utilization, and resilience.
- Support containerized application deployments and resolve runtime issues.
- Implement Kubernetes operational best practices around security, reliability, and scalability.
AWS Operations & Governance
- Support AWS account administration, IAM, networking, and infrastructure management.
- Implement and maintain AWS governance guardrails, security controls, and access policies.
- Ensure compliance with organizational standards for cloud infrastructure.
- Assist with cloud cost optimization and resource management.
- Partner with security teams to remediate infrastructure risks and vulnerabilities.
Incident Response & Operational Support
- Participate in production incident response and on-call rotations.
- Troubleshoot and resolve complex production issues across AWS and Kubernetes environments.
- Perform root cause analysis and implement corrective actions.
- Support a ticket-driven operational model with a strong focus on responsiveness and customer service.
- Document operational procedures and contribute to knowledge sharing across the team.
Monitoring & Reliability
- Maintain monitoring, logging, and alerting for cloud infrastructure and Kubernetes platforms.
- Proactively identify reliability risks before they impact
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s