Site Reliability Engineer II

Abbott Laboratories

Los Angeles (CA)

On-site

USD 82,000 - 141,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Abbott Laboratories in the Los Angeles area is seeking a Site Reliability Engineer II to join the DevOps team in the Cardiac Rhythm Management Division. This on-site role focuses on reliability, scalability, and operational excellence for Merlin.net, our remote patient data monitoring platform.

Design and maintain highly available systems, implement robust monitoring and automation, and collaborate with software, security, and compliance teams.

Qualifications

  • Bachelor's degree or equivalent in a technical field.

Responsibilities

  • Design, implement, and maintain highly available, fault-tolerant systems with uptime and safety requirements.

Skills

Excellent communication skills
Analytical problem-solving
Automation programming
Azure expertise
Linux fundamentals
Incident management
CI/CD knowledge
Observability tooling

Education

Associate's Degree
Bachelor's degree in Computer Science/Engineering or related field

Tools

AKS
Azure Monitor
Azure DevOps
Prometheus
Grafana
Datadog
ELK/EFK stack
Kubernetes
Docker

Job description

Position Title: Site Reliability Engineer II

Team:CRM DevOpsEmployment Type:Full-Time

About the Role

This Site Reliability Engineer II position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.

We are seeking a highly skilled and mission-drivenSite Reliability Engineer (SRE)to join our DevOps team. In this critical role, you will be responsible for ensuring the reliability, scalability, performance, and operational excellence ofMerlin.net— a remote monitoring platform designed to help doctors, cardiologists, and care teams automatically collect and review data from patients with implanted cardiac devices.

This isn't just about keeping servers up; it's about building and maintaining the resilient backbone for systems where failure is not an option, and where our success directly impacts patient care around the world. You will embed within our DevOps team, acting as a bridge between development and operations.

What You'll Do
  • Design, implement, and maintain highly available, fault-tolerant, and resilient systems that meet demanding uptime and safety requirements.
  • Identify and eliminate performance bottlenecks in software and infrastructure, ensuring low-latency, high-throughput, and real-time responsiveness for customer-facing services.
  • Define, monitor, and uphold Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Develop and implement comprehensive monitoring, logging, tracing, and alerting solutions to provide deep insights into system health and behavior at scale.
  • Automate away manual operational tasks, from provisioning and deployment to testing and recovery.
  • Develop and implement strategies for scaling our services and infrastructure to meet evolving business demands, including distributed systems and cloud deployments in Azure.
  • Work closely with software engineering, security, quality, and compliance teams to integrate SRE best practices into our operational processes and infrastructure, ensuring the integrity, availability, and confidentiality of our systems that carry sensitive patient health data.
  • Create clear, concise, and comprehensive documentation, runbooks, and playbooks for operational procedures. Lead blameless postmortem processes following incidents and drive systematic follow-through on action items.
  • Work with a multi-disciplinary team on challenging problems in a fast-paced environment, contributing across architecture reviews, incident response, capacity planning, and reliability roadmap planning.
Required Qualifications
  • Associates Degree
  • Minimum of one (1) year of full-time related work experience
  • Excellent communication skillswith the demonstrated ability to work effectively in cross-functional teams, translate technical complexity for non-technical stakeholders, and collaborate with development, quality, security, marketing, and regulatory teams.
  • Strong analytical, problem-solving, and debugging skillswith a methodical and structured approach to diagnosing complex, distributed system issues under pressure — including production incidents with patient safety implications.
  • Proficiency in at least one systems or automation programming language(e.g., Python, Go, Bash, PowerShell) for building tooling, automation, and operational systems.
  • Demonstrated expertise with Microsoft Azure— including Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, and related managed services.
  • Solid Linux & networking fundamentals— DNS, TCP/IP, HTTP/S, TLS, load balancing, and networking in cloud environments.
  • Incident management experience— including on-call rotation participation, structured incident response, root cause analysis (RCA), and systematic prevention of recurrence.
Preferred Qualifications
  • Bachelor's in Computer Science, Software Engineering, Systems Engineering, or a related technical discipline; equivalent professional experience will be considered.
  • Minimum two (2) years of full-time related work experience
  • Experience working in aregulated healthcare, medical device, or life sciences environment, with familiarity in compliance frameworks such as HIPAA.
  • Relevant professional certificationssuch as: Microsoft Certified Azure DevOps Engineer Expert, Azure Solutions Architect Expert, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Security Specialist (CKS).
  • Background incost optimization for cloud-native architectures, including FinOps practices for Azure environments.
  • Deep understanding of distributed systems— including load balancing, service meshes, microservices, message queues, and fault-tolerant design patterns.
  • Container orchestration expertise— hands-on production experience with Kubernetes and Docker at scale, including deployment strategies, resource management, and cluster operations.
  • Experience designing and operating CI/CD pipelinesfor continuous delivery of software in production environments, including safe deployment strategies such as blue/green, canary, and feature flag-gated rollouts.
  • Observability platform experiencewith tools such as Prometheus, Grafana, the ELK/EFK stack, Datadog, Azure Monitor, or similar enterprise-grade monitoring and tracing platforms.
  • Experience contributing to, or drivingdisaster recovery (DR)design, business continuity planning, and tabletop exercises.

The base pay for this position is $81,500.00 – $141,300.00. In specific locations, the pay range may vary from the range posted.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Talentify • Los Angeles (CA)

On-site
USD 82,000 - 141,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Talentify • Los Angeles (CA)

On-site
USD 90,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Fabric Labs, Inc. • New York (NY), Northern (KY)

On-site
USD 135,000 - 160,000
Medical, dental, vision
Unlimited PTO
401(k) plan
+2
Site Reliability Engineer
Site Reliability Engineer

Cosm Inc. • El Segundo (CA), Northern (KY)

On-site
USD 110,000 - 145,000
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

LexisNexis Risk Solutions • San Jose (CA), Northern (KY)

Hybrid
USD 105,000 - 175,000
401(k) with match
Wellbeing programs
Life Insurance
+1
Site Reliability Engineer II
Site Reliability Engineer II

Thornton Tomasetti • Los Angeles (CA)

On-site
USD 82,000 - 141,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Senior System Reliability Engineer
Senior System Reliability Engineer

On-Demand Group • Eagan (MN)

On-site
USD 229,233,000 - 286,541,000