Director, Site Reliability Operations

Omnicell

Austin (TX)

Hybrid

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Omnicell is seeking a Director of Site Reliability Operations to build and lead a modern, AI-enabled global SRO. You will own the production cloud strategy, operate 24×7 production services, and coordinate major incident management with executive stakeholders.

You will champion automation, observability, and standardized processes across engineering, security, product, and support teams, driving reliability and improved customer experience at scale.

Qualifications

  • Bachelor’s degree or higher in a technical field; 12+ years in cloud/production ops or related disciplines.
  • 7+ years leading 24×7 global production operations teams.
  • Experience with 24×7 incident management, MIM, and operational governance.

Responsibilities

  • Build and mentor a globally distributed Site Reliability Operations organization.
  • Lead 24×7 production operations, event management, and service restoration.
  • Drive operational governance, incident communications, and executive reporting.
  • Lead automation, AI-assisted operations, and runbooks to reduce toil.

Skills

Leadership
Executive communication
Crisis leadership
Automation
Observability
Incident management
Cross-functional collaboration
Strategic thinking

Education

Bachelor’s degree in Computer Science / Information Technology / Engineering
Master’s degree preferred

Job description

  • Noise Reduction
Operational Governance & Service Management

Provide governance across enterprise operational processes.

  • Enterprise Architecture
Director, Site Reliability Operations (SRO)
Department:

Global Cloud Operations

Reports To:

Vice President, Global Cloud Operations

Location:

Remote (U.S.) / Hybrid (Preferred)

Travel:

Up to 20% (Domestic and International)

Why Join Omnicell?

At Omnicell, our mission is to transform the pharmacy and nursing care experience through intelligent automation, cloud-native software, and operational excellence. As our healthcare platforms continue to scale globally, ensuring reliable, secure, and uninterrupted service delivery is essential to our customers and the patients they serve.

Site Reliability Operations (SRO) is the operational heartbeat of Omnicell’s cloud platform. This organization is responsible for operating production services, maintaining situational awareness across the enterprise, coordinating incident response, restoring service, and continuously improving operational excellence.

As Director of Site Reliability Operations, you will build and lead a modern, AI-enabled operations organization that moves beyond the traditional Network Operations Center (NOC). Your team will leverage automation, observability, operational intelligence, and disciplined operational processes to proactively manage production environments and ensure exceptional service reliability for customers around the world.

This is an opportunity to redefine cloud operations by building a world-class Site Reliability Operations organization that partners closely with Site Reliability Engineering, Cloud Platform Engineering, Cloud Security, Product Engineering, and Technical Support.

About This Opportunity

This is not a traditional NOC leadership role.

Reporting directly to the Vice President of Global Cloud Operations, the Director of Site Reliability Operations will establish Omnicell’s global operational strategy for cloud services, leading the teams responsible for 24×7 production operations, event management, service restoration, major incident management, operational governance, and operational readiness.

Working closely with the Director of Site Reliability Engineering, this organization will execute the operational strategy while SRE continuously engineers improvements to reliability, automation, and resilience. Together, these organizations create a modern cloud operating model where reliability is both engineered and operationally sustained.

The successful candidate will build a proactive, automation-driven operations organization that reduces operational risk, accelerates service restoration, and continuously improves the customer experience.

Purpose

The Director of Site Reliability Operations is responsible for leading Omnicell’s global Site Reliability Operations organization.

This role owns the operational execution, governance, and continuous improvement of production cloud services, ensuring that customer-facing platforms remain available, resilient, and operationally efficient.

The organization serves as the central coordination point for operational awareness, incident response, service restoration, operational communications, and production readiness while driving automation that reduces manual effort and improves operational maturity.

Success in this role is measured through operational excellence, service availability, incident response effectiveness, customer experience, and continuous operational improvement.

Primary Impact

As Director of Site Reliability Operations, you will define how Omnicell operates its production cloud environment.

Your leadership will establish a proactive operational culture that leverages automation, observability, AI-assisted operations, and disciplined operational processes to detect issues early, minimize customer impact, accelerate service restoration, and continuously improve operational performance.

Working alongside Site Reliability Engineering, Cloud Platform Engineering and Cloud Security, your organization will ensure that production systems remain stable, resilient, and ready to support Omnicell’s continued growth.

What You’ll Do
Build and Lead a World-Class Site Reliability Operations Organization

Build and scale Omnicell’s global Site Reliability Operations organization.

You Will:
  • Recruit, mentor, and develop Operations Managers and Site Reliability Operations Engineers.
  • Build a high-performing global operations organization.
  • Establish operational standards, career development frameworks, and leadership expectations.
  • Foster a culture of ownership, accountability, continuous improvement, and operational excellence.
  • Develop a follow-the-sun operational model supporting global cloud services.
24×7 Production Operations

Lead Omnicell’s enterprise production operations organization responsible for continuous monitoring and operational management of cloud services.

Responsibilities Include:
  • Production Operations
  • Operational Monitoring
  • Service Health Management
  • Event Management
  • Alert Triage
  • Operational Escalation
  • Shift OperationsOperational Command Center
  • Customer Impact Assessment
  • Operational Coordination

Ensure operational teams maintain continuous awareness of platform health while proactively identifying and mitigating production risks.

Major Incident Management

Own Omnicell’s enterprise Major Incident Management (MIM) program.

Responsibilities Include:
  • Executive Incident Command
  • Major Incident Coordination
  • Cross-Functional War Rooms
  • Service Restoration
  • Stakeholder Communications
  • Executive Communications
  • Customer Communications (in partnership with Support)
  • Incident Documentation
  • Incident Timeline Management
  • Escalation Governance

Lead high-severity incidents with urgency, structure, and transparency while minimizing business and customer impact.

Event & Operational Management

Establish enterprise operational governance for production events.

Own Processes Supporting:
  • Event Correlation
  • Alert Quality
  • Noise Reduction
  • Event Prioritization
  • Operational Dashboards
  • Monitoring Effectiveness
  • Operational KPIs
  • Operational Health Reviews

Partner with Site Reliability Engineering to continuously improve monitoring quality and operational intelligence.

Operational Readiness

Develop operational readiness standards for new services entering production.

Responsibilities Include:
  • Operational Readiness Reviews
  • Runbook Validation
  • Playbook Development
  • Operational Acceptance
  • Monitoring Validation
  • Escalation Readiness
  • Support Readiness
  • Disaster Recovery Readiness
  • Service Transition

Partner with Product Engineering and Site Reliability Engineering to ensure services are operationally prepared before production deployment.

Operational Excellence & Continuous Improvement

Drive continuous improvement across production operations.

Lead Initiatives Focused On:
  • Reducing MTTD (Mean Time to Detect)
  • Reducing MTTR (Mean Time to Restore)
  • Improving Incident Quality
  • Eliminating Recurring Operational Issues
  • Standardizing Operational Processes
  • Increasing Automation Adoption
  • Improving Service Availability
  • Enhancing Customer Experience

Establish a culture that uses operational metrics and post-incident learning to drive measurable improvements.

Operational Automation & AIOps

Champion automation and AI-assisted operations across the production environment.

Lead Initiatives Supporting:
  • Automated Incident Response
  • Self-Healing Workflows
  • Event Correlation
  • Intelligent Alerting
  • Automated Runbooks
  • Operational ChatOps
  • AI-Assisted Root Cause Analysis
  • Predictive Operations
  • Operational Knowledge Management

Partner with Site Reliability Engineering and Cloud Platform Engineering to engineer automation that reduces operational toil and improves operational consistency.

Operational Governance & Service Management

Provide governance across enterprise operational processes.

Own:
  • Operational Policies
  • Incident Governance
  • Problem Management
  • Change Coordination
  • Service Health Reporting
  • Operational Risk Reviews
  • SLA Compliance
  • Operational Metrics
  • Executive Operational Reporting

Ensure consistent operational practices across all production environments.

Managed Service Provider (MSP) Governance

Lead operational governance for Omnicell’s strategic managed service providers.

Responsibilities Include:
  • Operational Performance Management
  • SLA Governance
  • Vendor Escalation Management
  • Operational Reviews
  • Performance Scorecards
  • Continuous Improvement Programs
  • Contractual Operational Alignment

Ensure third-party operational partners consistently meet Omnicell’s operational standards and customer expectations.

Cross-Functional Partnership

Develop Trusted Partnerships Across:

  • Site Reliability Engineering
  • Cloud Platform Engineering
  • Cloud Security
  • Product Engineering
  • Technical Support
  • Enterprise Architecture
  • Customer Success
  • Engineering Leadership

Serve as the operational bridge between engineering organizations and customer-facing support teams.

Organizational Leadership

Provide Strategic Leadership By:

  • Building a globally distributed Site Reliability Operations organization.
  • Developing future operational leaders.
  • Driving employee engagement and professional growth.
  • Creating a culture centered on operational discipline and customer focus.
  • Encouraging innovation and operational excellence.
Executive Leadership

Partner With Executive Leadership To:

  • Present operational health and service performance.
  • Recommend investments that improve operational maturity.
  • Communicate production risks and mitigation strategies.
  • Lead executive incident communications during major outages.
  • Support enterprise cloud transformation initiatives.
  • Influence enterprise operational strategy.
What Success Looks Like
First 90 Days
  • Assess current production operations and operational maturity.
  • Build relationships with Engineering, Product, Security, and Support leaders.
  • Evaluate monitoring, incident management, and operational processes.
  • Identify opportunities to improve operational effectiveness.
  • Develop a multi-year Site Reliability Operations strategy.
First Six Months
  • Standardize enterprise incident management processes.
  • Improve operational dashboards and executive reporting.
  • Implement operational governance standards.
  • Launch automation initiatives to reduce manual operational effort.
  • Recruit key operational leadership positions.
  • Establish operational KPIs and scorecards.
First Year
  • Build a high-performing global Site Reliability Operations organization.
  • Improve MTTD, MTTR, and overall service availability.
  • Reduce alert fatigue through improved monitoring quality.
  • Expand operational automation and AI-assisted operations.
  • Establish a mature, scalable operational model supporting Omnicell’s global cloud platform.
  • Become a trusted operational partner across Engineering, Product, Support, and Executive Leadership.
Who You Are

You are a servant leader who thrives in dynamic production environments and believes operational excellence is built through discipline, teamwork, automation, and continuous learning.

You excel at coordinating complex operational events, building high-performing operations teams, and transforming reactive operational processes into proactive, data-driven capabilities.

You understand that exceptional cloud operations require strong partnerships with engineering organizations, customer-facing teams, and executive leadership. Most importantly, you are passionate about creating resilient operational systems that enable engineers to innovate while ensuring customers experience reliable, always-available services.

Minimum Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (Master’s degree preferred).
  • 12+ years of experience in cloud operations, production operations, NOC, SRE, IT Operations, or related disciplines.
  • 7+ years of progressive leadership experience managing managers and operational teams.
  • Experience leading 24×7 global production operations organizations.
  • Strong understanding of cloud-native architectures, observability platforms, incident management, and operational governance.
  • Experience driving operational transformation through automation and process improvement.
  • Exceptional executive communication and crisis leadership skills.
Preferred Qualifications
  • Experience building modern Site Reliability Operations or cloud operations organizations.
  • Experience with AIOps, observability platforms, and operational analytics.
  • Expertise with ITIL, Incident Management, Problem Management, and Change Management frameworks.
  • Experience supporting regulated healthcare, SaaS, or enterprise cloud platforms.
  • Demonstrated success leading globally distributed operations organizations.
  • Experience managing strategic managed service providers.
Leadership Expectations
As Director Of Site Reliability Operations, You Will:
  • Establish the operational vision for Omnicell’s production cloud environment.
  • Build and mentor a world-class Site Reliability Operations organization.
  • Champion operational excellence through disciplined execution, automation, and continuous improvement.
  • Lead with urgency, transparency, and accountability during production events.
  • Foster trusted partnerships across Engineering, Security, Product, and Support.
  • Serve as Omnicell’s executive authority on production operations, incident management, and operational governance.
How You’ll Elevate Omnicell

Your leadership will transform production operations into a strategic capability that enables Omnicell to deliver reliable, resilient, and secure cloud services at scale. By building a modern Site Reliability Operations organization grounded in automation, operational intelligence, and continuous improvement, you will ensure every customer benefits from exceptional service availability and operational excellence while providing engineering teams with the confidence to innovate rapidly and safely.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Site Reliability Operations
Director, Site Reliability Operations

Omnicell, Inc. • United States

Hybrid
USD 180,000 - 280,000
Director, Site Reliability Operations
Director, Site Reliability Operations

RXinsider LTD. • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 240,000
Director, Strategy & Operations
Director, Strategy & Operations

RXinsider LTD. • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 240,000
Director, Strategy & Chief of Staff
Director, Strategy & Chief of Staff

Omnicell, Inc. • United States

Hybrid
USD 180,000 - 240,000
Engineer III, Site Reliability
Engineer III, Site Reliability

Omnicell • Cranberry Township

Hybrid
USD 120,000 - 180,000
Engineer III, Site Reliability
Engineer III, Site Reliability

Omnicell • Austin (TX)

Hybrid
USD 130,000 - 185,000
Remote or hybrid work
Up to 10% travel
Director of Global Site Reliability & AI-Driven Ops
Director of Global Site Reliability & AI-Driven Ops

Omnicell, Inc. • United States

Hybrid
USD 180,000 - 280,000
Director, Global Site Reliability Operations (Remote)
Director, Global Site Reliability Operations (Remote)

Omnicell • Austin (TX)

Hybrid
USD 180,000 - 260,000
Director, Global Site Reliability & AI-Driven Ops
Director, Global Site Reliability & AI-Driven Ops

RXinsider LTD. • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 240,000
Technical Director, Global Cloud Operations
Technical Director, Global Cloud Operations

Omnicell • Austin (TX)

On-site
USD 173,000 - 240,000
Travel up to 20%
Remote-friendly work arrangement