About the Role
We are seeking a hands‑on, operationally focused Cloud / Platform Operations Manager to lead the day‑to‑day management of our AWS and Kubernetes environments. This leader will bring deep technical expertise in operating, troubleshooting, and optimizing production Kubernetes platforms while ensuring the reliability, security, and governance of our cloud infrastructure.
Responsibilities
- Operational Leadership & Service Delivery
- Own day‑to‑day operations of AWS and Kubernetes environments, ensuring reliability, availability, and performance.
- Provide technical leadership and hands‑on support for complex Kubernetes and cloud infrastructure issues.
- Lead a ticket‑driven support model, prioritizing and ensuring timely fulfillment of internal user requests.
- Establish and enforce SLAs/SLOs for platform support and operational responsiveness.
- Drive a culture of execution, accountability, and customer service within the team.
- AWS Governance, Guardrails & Policy Enforcement
- Define, implement, and enforce AWS guardrails, IAM policies, and governance frameworks.
- Ensure proper account structure, access controls, cost controls, and compliance standards.
- Partner with security teams to enforce best practices and reduce risk across the platform.
- Continuously audit and improve cloud policy adherence.
- Kubernetes Operations & Stability
- Own the operational health and lifecycle management of production Kubernetes clusters.
- Maintain and troubleshoot Kubernetes control planes, worker nodes, networking, storage, ingress, and containerized workloads.
- Perform cluster administration, upgrades, patching, capacity planning, and performance tuning.
- Ensure robust monitoring, logging, alerting, and incident response.
- Drive operational best practices for cluster reliability, security, and resiliency.
- Partner with application teams to troubleshoot deployment, scaling, networking, and workload issues.
- Incident & Escalation Management
- Participate in a rotating after‑hours on‑call schedule to provide operational support for critical production incidents.
- Act as the primary escalation point for AWS and Kubernetes operational issues.
- Lead technical troubleshooting during production incidents and guide the team through resolution.
- Handle and resolve escalations effectively before they reach senior leadership.
- Lead incident response, root‑cause analysis, and post‑incident improvements.
- Build strong communication channels with stakeholders during incidents.
- Team Leadership & Development
- Lead, coach, and develop a team of cloud/platform engineers with an operational mindset.
- Mentor engineers on Kubernetes operations, AWS best practices, and operational excellence.
- Set clear expectations around ownership, responsiveness, and execution.
- Manage performance, provide feedback, and address gaps in delivery or behavior.
- Foster a culture of accountability, urgency, and continuous improvement.
- Cross‑Functional Support
- Partner with engineering, product, and business teams to support platform needs and unblock users.
- Balance operational workload with longer‑term improvements without compromising service delivery.
- Act as a bridge between users and platform capabilities, ensuring needs are met efficiently.
Required Qualifications
- 8+ years of experience in cloud infrastructure or platform operations, with at least 2 years in a leadership role.
- Strong hands‑on experience administering and supporting production Kubernetes environments, including cluster operations, troubleshooting, upgrades, networking, storage, ingress, and workload management.
- Proven ability to independently diagnose and resolve complex Kubernetes platform issues in production.
- Strong hands‑on experience with AWS, including IAM, networking, account structure, and governance controls.
- Proven experience managing AWS guardrails, policies, and access controls at scale.
- Experience with Kubernetes observability and operational tooling (e.g., Prometheus, Grafana, CloudWatch, Fluent Bit).
- Deep operational experience running and supporting Kubernetes environments in production.
- Experience leading ticket‑based support or operations teams with high request volume.
- Strong incident management and escalation handling experience.
- Demonstrated ability to lead teams focused on execution, reliability, and service delivery while remaining technically engaged.
- Excellent communication and stakeholder management skills.
Compensation and Benefits
Expected annual base salary: $180k–$210k, subject to role, level, experience, and location. The position may also be eligible for a discretionary annual bonus based on individual impact, team, and firm performance. Bain Capital offers a competitive benefits package designed to support employees’ health, financial security, family needs, and overall well‑being.
Equal Employment Opportunity
Bain Capital is an equal‑opportunity employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, or veteran status. For job applicants in the United States, Bain Capital participates in the E‑Verify program and will use E‑Verify to confirm work authorization.