Manager - Production Operations & Site Reliability Engineering

Alcon MX

Lake Forest (CA)

On-site

USD 140,000 - 182,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
PTO
Retirement plan

Job summary

Alcon is seeking a Principal Engineer/Manager to lead Production Operations and Site Reliability Engineering for its Digital Health Cloud platform. You will own reliability, observability, and deployment governance across SaaMD and customer-facing apps, guiding incident response, capacity planning, and disaster recovery.

The role requires deep cloud expertise (AWS, Kubernetes, Istio) and strong collaboration with Security, Architecture, and Product teams to deliver resilient, compliant

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 5+ years of experience in Production Operations or Site Reliability Engineering.
  • Experience with regulated healthcare or SaMD platforms is a plus.
  • Strong knowledge of cloud architectures, security, and high availability.

Responsibilities

  • Lead production operations for cloud-based healthcare platforms to ensure high availability.
  • Drive reliability improvements, incident management, and governance across environments.
  • Govern release readiness, validation, staging, and production deployment processes.
  • Champion SRE practices, automation, and observability to reduce toil and MTTR.
  • Collaborate with Engineering, Security, and Operations to scale, secure, and optimize platforms.

Skills

Production Ops
SRE
Cloud Ops

Education

Bachelor's Degree in CS/Engineering

Tools

AWS
Kubernetes
Istio
Datadog
CI/CD
Terraform

Job description

Manager - Production Operations & Site Reliability EngineeringSkip to main content#Manager - Production Operations & Site Reliability Engineering page is loaded## Manager - Production Operations & Site Reliability EngineeringApplyremote type: Not applicablelocations: Lake Forest, Californiatime type: Full timeposted on: Posted Todaytime left to apply: End Date: August 31, 2026 (17 days left to apply)job requisition id: R-2026-48834**Manager – Production Operations & Site Reliability Engineering**At Alcon, we are driven by the meaningful work we do to help people see brilliantly. We innovate boldly, champion progress, and act with speed as the global leader in eye care. Here, you’ll be recognized for your commitment and contributions and see your career like never before. Together, we go above and beyond to make an impact in the lives of our patients and customers.We foster an inclusive culture and are looking for diverse, talented people to join Alcon. As a **Principal Engineer** you will provide technical leadership for the reliability, availability, security, and continuous improvement of Alcon’s Digital Health Cloud platform supporting Software as a Medical Device (SaMD) and customer-facing digital health applications.This role serves as the senior technical authority for production operations and Site Reliability Engineering (SRE), driving platform reliability, observability, automation, incident management, release governance, and cloud optimization. Partnering across Product Engineering, Architecture, Security, Infrastructure, and Operations, the Principal Engineer enables resilient, compliant, and scalable healthcare platforms with predictable, high-quality software delivery. In this role, a typical day will include:**Production Operations & Reliability*** Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability, reliability, and performance.* Establish and continuously improve operational standards, SLIs/SLOs, readiness reviews, and service excellence practices.* Drive platform resilience, capacity planning, disaster recovery, and governance through Alcon’s Steady State Operations Framework (SSOF).**Release & Deployment Governance*** Lead production readiness reviews, ensuring applications meet operational, security, monitoring, compliance, and supportability requirements.* Govern the Road to Production process across Validation, Staging, and Production environments, validating deployment readiness, infrastructure qualification, rollback plans, and acceptance criteria.* Partner with Engineering and DevOps teams to improve release quality, deployment reliability, and change success through standardized governance and automation.**Site Reliability Engineering (SRE)*** Champion SRE best practices, leveraging automation and self-healing capabilities to reduce operational toil.* Improve service reliability and customer experience by reducing MTTD and MTTR and increasing deployment success rates.* Lead incident investigations, root cause analyses, and long-term corrective actions for critical production events.**Cloud Platform Operations*** Provide technical leadership across AWS platforms including EKS, EC2, RDS, S3, ElastiCache, AWS MQ, Route53, and Kubernetes/Istio.* Optimize cloud infrastructure for scalability, resilience, security, performance, and cost efficiency.**Observability & Automation*** Define enterprise observability strategies using Datadog, CloudWatch, distributed tracing, synthetic monitoring, centralized logging, and executive dashboards.* Lead automation initiatives across deployments, monitoring, health validation, incident response, and operational workflows.**Security & Compliance*** Ensure compliance with HIPAA, GDPR, FDA, and enterprise cybersecurity standards.* Partner with Security teams to strengthen cloud security architecture, identity management, network segmentation, and operational controls.**Technical Leadership*** Serve as the senior escalation point for major incidents, production events, and Hypercare operations.* Mentor engineering teams on operational excellence, production engineering, and SRE best practices.* Influence platform and architectural decisions that enhance operability, maintainability, resilience, and long-term service reliability.* Collaborate across Engineering, Architecture, Infrastructure, Security, and Global Operations to advance platform stability and operational maturity.**WHAT YOU’LL BRING TO ALCON:**Bachelor’s Degree or Equivalent years of directly related experience (or high school +13 yrs; Assoc.+9 yrs; M.S.+2 yrs; PhD+0 yrs)The ability to fluently read, write, understand and communicate in English5 Years of Relevant Experience**PREFERRED QUALIFICATIONS:*** Bachelor’s Degree in degree in Computer Science, Engineering, or related field (Master’s preferred)* Experience in Production Operations, Site Reliability Engineering, Platform Engineering, or Cloud Operations* Experience supporting regulated healthcare or medical device platforms* AWS Solutions Architect or DevOps Professional certification* Kubernetes (CKA/CKS) certification* Experience with healthcare interoperability standards (HL7, FHIR, DICOM)* Experience implementing Site Reliability Engineering practices within large-scale cloud environments* AWS (EKS, EC2, RDS, S3, Route53, IAM, Load Balancers, CloudWatch)* Kubernetes, Istio, Docker* Datadog, APM, logging, distributed tracing, synthetic monitoring* CI/CD, Infrastructure as Code, automation* HIPAA, GDPR, FDA compliance* Strong understanding of cloud networking, security, and high-availability architectures**HOW YOU CAN THRIVE AT ALCON:*** Benefit from working in a highly collaborative environment.* Join Alcon’s mission to provide top-tier, innovative products to enhance sight, enhance lives, and grow your career.* Alcon provides robust benefits package including health, life, retirement, PTO, and much more!**Alcon Careers**See your impact at alcon.com/careers**ATTENTION: Current Alcon Employee/Contingent Worker**If you are currently an active employee/contingent worker at Alcon, please click the appropriate link below to apply on the Internal Career site.Find Jobs for EmployeesFind Jobs for Contingent Worker**Compensation and Benefits**Alcon’s Total Rewards programs are designed to align incentives with business objectives, support our values, and deliver long‐term value. Our compensation approach includes a combination of fixed and variable pay, with short‐term and long‐term incentive opportunities for eligible roles. Our benefits offerings are designed to support associates and their families across key life events, including programs that promote health and well‐being, provide financial security, and support retirement planning.*The salary range posted represents the anticipated hiring range for this role. Actual compensation may vary based on factors such as experience, skills, location, and internal equity, and may fall outside the posted range.***Pay Range**140,250.00 - 181,500.00**Pay Frequency**AnnualAlcon is an Equal Opportunity Employer and participates in E-Verify. Alcon takes pride in maintaining a diverse environment and our policies are not to discriminate in recruitment, hiring, training, promotion or other employment practices for reasons of race, color, religion, gender, national origin, age, sexual orientation, gender identity, marital or veteran status, disability, or any other legally protected status. Alcon is also committed to working with and providing reasonable accommodation to individuals with disabilities. If, because of a medical condition or disability, you need a reasonable accommodation for any part of the application process, or in order to perform the essential functions of a position, please send an email to ****alcon.recruitment@alcon.com**** and let us know the nature of your request and your contact information.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager - Production Operations & Site Reliability Engineering
Manager - Production Operations & Site Reliability Engineering

Alcon • Lake Forest (CA)

On-site
USD 140,000 - 182,000
Health insurance
Retirement plan
Paid time off
Manager - Production Operations & Site Reliability Engineering
Manager - Production Operations & Site Reliability Engineering

Alcon MX • Lake Forest (IL)

On-site
USD 140,000 - 182,000
Health benefits
Retirement plan
Paid time off (PTO)
Principal Software Engineer (Medical Device Software/C#, .NET, Windows)
Principal Software Engineer (Medical Device Software/C#, .NET, Windows)

Alcon MX • Lake Forest (CA)

On-site
USD 128,000 - 167,000
Job Profile Name Sr. Manager, Cloud & Middleware Architecture (M)
Job Profile Name Sr. Manager, Cloud & Middleware Architecture (M)

Alcon MX • Lake Forest (CA)

On-site
USD 168,000 - 218,000
Principal, Software Engineering - Digital Health
Principal, Software Engineering - Digital Health

Alcon • Fort Worth (TX)

On-site
USD 140,000 - 190,000
Relocation assistance
Sponsorship available
Software Delivery Lead
Software Delivery Lead

Alcon MX • Lake Forest (CA)

On-site
USD 128,000 - 167,000
Senior Regulatory Affairs Associate
Senior Regulatory Affairs Associate

Alcon MX • Lake Forest (CA)

Hybrid
USD 103,000 - 133,000
Sr. Principal Manufacturing Data Architect
Sr. Principal Manufacturing Data Architect

Alcon MX • Fort Worth (TX)

On-site
USD 100,000 - 150,000
Robust benefits package
Flexible time off
Principal Engineer - Software Life Cycle Management
Principal Engineer - Software Life Cycle Management

Alcon • Lake Forest (CA)

On-site
USD 114,000 - 172,000
Robust benefits package
Collaborative work environment
Opportunities for professional development
Technical Program Manager - Digital Twin of Eye - R&D
Technical Program Manager - Digital Twin of Eye - R&D

Alcon MX • Lake Forest (CA)

On-site
USD 130,000 - 170,000
Health insurance
Retirement plan
Flexible time off