Engineer III, Site Reliability

RXinsider LTD.

Cranberry Township (Butler County)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Omnicell is seeking a Site Reliability Engineer to own reliability and operational health of cloud services enabling critical pharmacy automation. You will define SLIs/SLOs, build dashboards, and drive automation to reduce toil.

You will participate in on‑call rotations and lead blameless post‑incident reviews while partnering with MSPs for escalation paths. The role emphasizes CI/CD across tools like GitHub Actions and Octopus Deploy, IaC with Terraform, and a strong focus on observability and

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related technical field.
  • 5+ years of experience in software or platform engineering, including 3+ years in an SRE/DevOps or reliability-focused role.
  • Strong hands‑on experience with cloud platforms (AWS, Azure, or GCP).
  • Proficiency in Python or another OO language for automation.
  • Production experience with Kubernetes, Docker, and Helm.
  • Experience Infrastructure as Code using Terraform or similar frameworks.
  • Working knowledge of observability tools across metrics, logs, and tracing.
  • Real‑world incident response experience including on‑call participation.
  • Solid Linux system administration skills.
  • Collaborative, coachable mindset with senior mentorship.

Responsibilities

  • Own reliability outcomes for assigned services with instrumentation, alerts, dashboards and runbooks.
  • Define and implement SLOs/SLIs; surface reliability in reviews.
  • Identify toil and design automation to reduce manual work.
  • Drive improvements in observability, automation, and resilience.
  • Participate in on‑call rotation and lead post‑incident reviews.
  • Coordinate with MSPs to ensure clean escalation paths into SRE ownership.
  • Design and operate CI/CD pipelines for cloud-native deliveries.
  • Automate infrastructure via code (Terraform).
  • Contribute to the evolution of observability platforms and diagnostics.
  • Support architecture reviews with a reliability lens.

Skills

Cloud platforms (AWS/Azure/GCP)
Python
Observability
Linux
On-call incident experience
CI/CD tooling
GitOps (ArgoCD/Flux)
CI/CD pipelines

Education

Bachelor’s degree in Computer Science/Engineering or related field

Tools

Kubernetes
Docker
Helm
Terraform

Job description

Site Reliability Engineer (SRE)

Location: Hybrid (U.S.)

Department: Global Cloud Operations

Reports to: VP, Global Cloud Operations

Why Join Omnicell?

At Omnicell, we’re transforming how medications and supplies are managed across the entire healthcare continuum—helping clinicians deliver safer, more efficient patient care every day. As we evolve our portfolio from on-premise systems to a cloud-native SaaS platform, reliability has become a core product feature, not a back-office function. This is a rare opportunity to join a newly formed Site Reliability Engineering practice as its second hire. You’ll work side‑by‑side with a Senior SRE in a true player‑coach model, gaining hands‑on mentorship while helping define how reliability, observability, and incident response are built at Omnicell. The scope is broad, the impact is real, and the growth path is intentional.

What You’ll Do
Primary Impact:

As a Site Reliability Engineer, you will own the reliability, scalability, and operational health of a defined set of cloud services that support mission‑critical pharmacy automation systems used by healthcare providers worldwide.

Service Reliability & Automation
  • Own reliability outcomes for assigned services, ensuring strong instrumentation, actionable alerts, meaningful dashboards, and up‑to‑date runbooks.
  • Define and implement SLIs and SLOs in partnership with product and engineering teams, and surface reliability performance in regular Cloud Operations reviews.
  • Identify operational toil and design automation to eliminate repetitive manual work.
  • Drive continuous improvement initiatives that increase observability, automation coverage, and system resilience.
Incident Response & Operational Excellence
  • Participate in the SRE on‑call rotation, progressing from secondary to primary ownership as readiness increases.
  • Command Sev-2 and Sev-3 incidents independently over time, with pairing and coaching from a Senior SRE; act as technical lead during Sev-1 incidents.
  • Lead blameless post‑incident reviews and own follow‑up actions through completion.
  • Partner closely with managed services providers (IBM, HCL) to ensure clean escalation paths from L1/L2 monitoring into SRE ownership.
Platform, CI/CD & Observability
  • Design, build, and operate CI/CD pipelines supporting cloud-native application delivery using tools such as GitHub Actions, CodeFresh, TeamCity, and Octopus Deploy.
  • Automate infrastructure and platform services using Infrastructure as Code (Terraform preferred).
  • Contribute to the evolution of Omnicell’s observability platform, including intelligent alerting, ML-based anomaly detection, and automated diagnostics.
  • Participate in architecture and launch readiness reviews, bringing a reliability lens to system design.
  • Help establish reference implementations and "golden paths" that enable product teams to launch services with reliability built in from day one.
Who You Are
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of experience in software or platform engineering, including 3+ years in an SRE, DevOps, or reliability-focused role.
  • Strong hands‑on experience with at least one major public cloud platform (AWS, Azure, or GCP).
  • Proficiency in Python or another object‑oriented programming language for automation and tooling.
  • Production experience with Kubernetes, Docker, and Helm.
  • Experience implementing Infrastructure as Code using Terraform or similar frameworks.
  • Working knowledge of modern observability tools across metrics, logs, and tracing.
  • Real‑world incident response experience, including on‑call participation and post‑incident write‑ups.
  • Solid Linux system administration skills.
  • Collaborative, coachable mindset with a desire to grow under senior mentorship.
Preferred Qualifications
  • Experience working in regulated environments such as healthcare, financial services, or government (HIPAA, SOC 2, or similar).
  • Familiarity with managed service provider models for L1/L2 operations.
  • Exposure to AIOps, ML-based anomaly detection, or LLM-assisted incident triage.
  • Understanding of GitOps principles and tools such as ArgoCD or Flux.
  • Experience operating secure, compliant Kubernetes platforms.
  • Familiarity with chaos engineering, messaging systems (Kafka, RabbitMQ), or stateful services in Kubernetes.
How You’ll Elevate at Omnicell
About
  • Collaborate: Partner closely with product engineering, security, and operations teams to build shared ownership of reliability.
  • Inspire: Influence reliability best practices across teams by modeling calm, structured incident leadership.
  • Develop: Continuously build your technical depth while learning directly from a senior SRE mentor.
  • Execute: Take ownership of services, incidents, and follow‑through—turning lessons learned into measurable improvements.
  • Impact: Help shape foundational SRE practices and introduce modern reliability and AIOps capabilities that scale with the business.
Growth & Career Path

This role is intentionally designed as a growth role. With strong performance and increasing ownership, the natural progression is into a Senior Site Reliability Engineer position as the practice scales. Omnicell also supports lateral growth into platform engineering, security engineering, or product engineering for SREs who discover adjacent passions.

Work Conditions
  • Remote or hybrid work environment supported.
  • Up to 10% travel as needed.
  • Participation in an SRE on‑call rotation is required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Director, Site Reliability Operations
Director, Site Reliability Operations

RXinsider LTD. • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer - Cloud Platform & Growth
Senior Site Reliability Engineer - Cloud Platform & Growth

RXinsider LTD. • Cranberry Township

Hybrid
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice By Numbers, Inc. • Bellevue (WA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice By Numbers, Inc. • Redmond (WA)

On-site
USD 130,000 - 160,000
High impact work on healthcare infrastructure
Strong engineering culture focused on automation
Small team with high ownership and autonomy