Senior Site Reliability Engineer

Abbott

Sunnyvale (CA)

On-site

USD 90,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Abbott is hiring a Senior Site Reliability Engineer to join our DevOps team on-site in Sunnyvale, CA or Sylmar, CA. You will ensure the reliability, scalability, and security of Merlin.net, a remote patient data monitoring platform, by building resilient systems with SLIs/SLOs and robust incident response.

You will collaborate with software, security, and compliance teams to automate delivery, improve observability, and guide capacity planning in a fast-paced environment with patient health data

Qualifications

  • Bachelor's in Computer Science, Software Engineering, Systems Engineering, or related field.
  • Minimum 7 years of site reliability or related experience.

Responsibilities

  • Design, implement, and maintain highly available, fault‑tolerant systems.
  • Identify and eliminate performance bottlenecks for low latency and high throughput.
  • Define and monitor SLIs/SLOs and error budgets.
  • Develop monitoring, logging, tracing, and alerting frameworks.
  • Automate provisioning, deployment, and recovery tasks.
  • Scale services and infrastructure for cloud deployments (Azure).
  • Collaborate with security, compliance, and QA teams to ensure data integrity and confidentiality.
  • Create runbooks and post‑mortem processes to prevent recurrence.

Skills

SRE principles
Incident management
Distributed systems
Cross-functional collaboration
Troubleshooting & debugging
Automation scripting (Python/Go/Bash)
Performance optimization

Education

Bachelor's in CS/Engineering or related field

Tools

Azure
AKS
Azure Monitor
Azure DevOps
Kubernetes
Docker
Prometheus
Grafana
Datadog
ELK/EFK stack

Job description

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life‑changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.

About The Role

This Senior Site Reliability Engineer position works on‑site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.

We are seeking a highly skilled and mission‑driven Senior Site Reliability Engineer (SRE) to join our DevOps team. In this critical role, you will be responsible for ensuring the reliability, scalability, performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists, and care teams automatically collect and review data from patients with implanted cardiac devices.

This isn’t just about keeping servers up; it’s about building and maintaining the resilient backbone for systems where failure is not an option, and where our success directly impacts patient care around the world. You will embed within our DevOps team, acting as a bridge between development and operations.

What You’ll Do

Design, implement, and maintain highly available, fault‑tolerant, and resilient systems that meet demanding uptime and safety requirements. Identify and eliminate performance bottlenecks in software and infrastructure, ensuring low‑latency, high‑throughput, and real‑time responsiveness for customer‑facing services. Define, monitor, and uphold Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Develop and implement comprehensive monitoring, logging, tracing, and alerting solutions to provide deep insights into system health and behavior at scale. Automate away manual operational tasks, from provisioning and deployment to testing and recovery. Develop and implement strategies for scaling our services and infrastructure to meet evolving business demands, including distributed systems and cloud deployments in Azure. Work closely with software engineering, security, quality, and compliance teams to integrate SRE best practices into our operational processes and infrastructure, ensuring the integrity, availability, and confidentiality of our systems that carry sensitive patient health data. Create clear, concise, and comprehensive documentation, runbooks, and playbooks for operational procedures. Lead blameless post‑mortem processes following incidents and drive systematic follow‑through on action items. Work with a multi‑disciplinary team on challenging problems in a fast‑paced environment, contributing across architecture reviews, incident response, capacity planning, and reliability roadmap planning.

Required Qualifications
  • Bachelor's in Computer Science, Software Engineering, Systems Engineering, or a related technical discipline; equivalent professional experience will be considered.
  • Minimum 7 years of experience working in site reliability, software engineering, systems engineering or a related technical discipline.
  • Excellent communication skills with the demonstrated ability to work effectively in cross‑functional teams, translate technical complexity for non‑technical stakeholders, and collaborate with development, quality, security, marketing, and regulatory teams.
  • Strong analytical, problem‑solving, and debugging skills with a methodical and structured approach to diagnosing complex, distributed system issues under pressure — including production incidents with patient safety implications.
  • Proficiency in at least one systems or automation programming language (e.g., Python, Go, Bash, PowerShell) for building tooling, automation, and operational systems.
  • Demonstrated expertise with Microsoft Azure — including Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, and related managed services.
  • Container orchestration expertise — hands‑on production experience with Kubernetes and Docker at scale, including deployment strategies, resource management, and cluster operations.
  • Observability platform experience with tools such as Prometheus, Grafana, the ELK/EFK stack, Datadog, Azure Monitor, or similar enterprise‑grade monitoring and tracing platforms.
  • Experience designing and operating CI/CD pipelines for continuous delivery of software in production environments, including safe deployment strategies such as blue/green, canary, and feature‑flag‑gated rollouts.
  • Deep understanding of distributed systems — including load balancing, service meshes, microservices, message queues, and fault‑tolerant design patterns.
  • Solid Linux & networking fundamentals — DNS, TCP/IP, HTTP/S, TLS, load balancing, and networking in cloud environments.
  • Incident management experience — including on‑call rotation participation, structured incident response, root cause analysis (RCA), and systematic prevention of recurrence.
Preferred Qualifications
  • Experience working in a regulated healthcare, medical device, or life sciences environment, with familiarity in compliance frameworks such as HIPAA.
  • Relevant professional certifications such as: Microsoft Certified Azure DevOps Engineer Expert, Azure Solutions Architect Expert, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Security Specialist (CKS).
  • Background in cost optimization for cloud‑native architectures, including FinOps practices for Azure environments.

Experience contributing to or driving disaster recovery (DR) design, business continuity planning, and tabletop exercises.

The base pay for this position is $90,000.00 – $180,000.00. In specific locations, the pay range may vary from the range posted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer — Healthcare Platform
Senior Site Reliability Engineer — Healthcare Platform

Abbott • Sunnyvale (CA)

On-site
USD 90,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Govcio LLC • United States

Hybrid
USD 210,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice by Numbers • Bellevue (WA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer - 7 Month Contract
Site Reliability Engineer - 7 Month Contract

Orion Health group • Fort Worth (TX), Town of Texas (WI)

On-site
USD 110,000 - 170,000
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5