Senior DevOps Engineer

Qualified Health PBC

Palo Alto (CA)

Hybrid

USD 170,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical/Dental/Vision insurance
Hybrid work options
Flexible working hours

Job summary

Qualified Health PBC is seeking a Senior DevOps Engineer / SRE to ensure reliability, performance, and operational excellence of production environments powering AI healthcare solutions. You will partner with engineering to ship features safely, own observability, incident response, and drive reliability as we scale.

You will design zero-trust networks, build CI/CD pipelines, and automate with Python and Terraform, while collaborating with security and compliance to meet HIPAA/HITRUST controls

Qualifications

  • 6+ years in SRE/DevOps or Infra Eng with production workloads.
  • Terraform mastery including modules, state, multi-env setups.
  • Production Kubernetes experience with troubleshooting and networking.
  • GCP and Azure services experience.
  • Zero trust security concepts and secrets management.
  • Temporal or similar workflow orchestration in production.
  • Python for automation and tooling.
  • Observability stack design: metrics, logs, tracing, alerts.
  • Incident response leadership and postmortems.
  • Strong communication and release documentation skills.

Responsibilities

  • Ensure services are production-ready before release, review patterns and rollback plans.
  • Design and maintain observability with metrics, logging, tracing, dashboards.
  • Define alerting, SLIs/SLOs, and on-call rotations.
  • Lead incident response, perform root cause analysis and hotfix coordination.
  • Maintain release docs, runbooks, postmortems, and operational playbooks.
  • Provide day-to-day support, unblock deployments and improve developer experience.
  • Design and manage zero-trust networks across multi-cloud environments.
  • Build and improve CI/CD pipelines for safer, faster deployments.
  • Automate operations with Python and Terraform.
  • Manage production Kubernetes workloads and Temporal workflows.

Skills

SRE/DevOps experience
Python scripting
Incident response
Cloud networking
Cross-team collaboration

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

Terraform & Terragrunt
Kubernetes (GKE)
Temporal workflow orchestration
Python
Go / Shell scripting
GitOps (ArgoCD/Rancher Fleet/Flux)

Job description

Transform healthcare with us.

At Qualified Health, we’re redefining what’s possible with Generative AI in healthcare. Our infrastructure provides the guardrails for safe AI governance, healthcare-specific agent creation, and real-time algorithm monitoring—working alongside leading health systems to drive real change.

This is more than just a job. It’s an opportunity to build the future of AI in healthcare, solve complex challenges, and make a lasting impact on patient care. If you’re ambitious, innovative, and ready to move fast, we’d love to have you on board.

Join us in shaping the future of healthcare.
Job Summary

We're looking for a Senior DevOps Engineer / Site Reliability Engineer to ensure the reliability, performance, and operational excellence of our production environments powering AI solutions for major health systems. You'll partner closely with engineering teams to make services production-ready, own observability and incident response, and drive the practices that keep our platform stable as we scale. As a key member of our infrastructure team, you'll be the connective tissue between development and production, ensuring new features ship safely while maintaining the reliability standards required for healthcare workloads.

Key Responsibilities
  • Partner with engineering teams to ensure services are production-ready before release, including reviewing deployment patterns, failure modes, resource requirements, and rollback strategies

  • Design and maintain observability infrastructure including metrics, logging, distributed tracing, and dashboards across multi-cloud environments

  • Define and manage alerting policies, SLIs/SLOs, and on-call rotations to ensure timely response to production issues

  • Lead and support incident response for production issues, drive root cause analysis, and coordinate hotfix deployments when needed

  • Author and maintain release documentation, runbooks, incident postmortems, and operational playbooks

  • Provide day-to-day operational support to engineering teams, unblocking deployments, debugging production issues, and improving developer experience around shipping to production

  • Design and maintain zero trust network architectures, ensuring secure connectivity across multi-cloud environments and tenant boundaries

  • Build and improve CI/CD pipelines and release processes to make production deployments safer, faster, and more predictable

  • Develop automation in Python and Terraform to reduce toil and codify operational best practices

  • Manage Kubernetes-based workloads in production, including troubleshooting cluster issues, optimizing resource utilization, and maintaining workload reliability

  • Operate Temporal workflows in production, including monitoring, scaling, and troubleshooting long-running workflow executions

  • Collaborate with security and compliance teams to maintain HIPAA and HITRUST controls across production environments

Required Qualifications
  • 6+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, with at least 3 years directly managing production workloads

  • Strong proficiency with Terraform including module development, state management, and multi-environment architectures

  • Deep experience operating production Kubernetes environments, including troubleshooting, networking, workload management, and cluster operations

  • Hands-on experience with both Google Cloud Platform and Microsoft Azure services

  • Strong networking and security knowledge, including zero trust architectures, network segmentation, private connectivity, identity-based access controls, and secrets management

  • Production experience with Temporal or comparable workflow orchestration systems

  • Strong proficiency in Python for automation, tooling, and operational scripting

  • Demonstrated experience designing and operating observability stacks including metrics, logging, tracing, and alerting

  • Experience leading incident response, including on-call rotation management, runbook development, and postmortem processes

  • Track record of partnering with engineering teams to improve production readiness and release practices

  • Excellent written communication skills for authoring runbooks, postmortems, and release documentation

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience

Desirable Skills
  • Experience in healthcare industry with understanding of HIPAA compliance requirements

  • Familiarity with HITRUST or similar compliance frameworks

  • Experience operating LLM-based systems, agentic workflows, or RAG pipelines in production

  • Experience with GitOps workflows (Rancher Fleet, ArgoCD, or Flux)

  • Experience building and operating multi-tenant SaaS infrastructure

  • Familiarity with chaos engineering and reliability testing practices

  • Prior experience as a founding or early SRE/Platform hire at a startup

Technical Environment

Our infrastructure is built on modern cloud technologies including:

  • Google Cloud Platform (primary) and Microsoft Azure

  • Google Kubernetes Engine (GKE)

  • Terraform and Terragrunt

  • Temporal for workflow orchestration

  • Python, Go, Shell scripting

  • GitOps-based deployment workflows

  • Modern monitoring and observability tools

Why Join Qualified Health?

This is an opportunity to join a fast-growing company and a world-class team, that is poised to change the healthcare industry. We are a passionate, mission-driven team that is building a category-defining product. We are backed by premier investors and are looking for founding team members who are excited to do the best work of their careers.

Our employees are integral to achieving our goals so we are proud to offer competitive salaries with equity packages, robust medical/dental/vision insurance, flexible working hours, hybrid work options and an inclusive environment that fosters creativity and innovation.

Our Commitment to Diversity

Qualified Health is an equal opportunity employer. We believe that a diverse and inclusive workplace is essential to our success, and we are committed to building a team that reflects the world we live in. We encourage applications from all qualified individuals, regardless of race, color, religion, gender, sexual orientation, gender identity or expression, age, national origin, marital status, disability, or veteran status.

Pay & Benefits: The pay range for this role is between $170,000 and $220,000, and will depend on your skills, qualifications, experience, and location. This role is also eligible for equity and benefits.

Join our mission to revolutionize healthcare with AI. To apply, please send your resume through the application below.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Senior DevOps Engineer
Senior DevOps Engineer

Qualified Health • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Dental insurance
+3
Senior Platform Engineer
Senior Platform Engineer

Qualified Health • Palo Alto (CA)

On-site
USD 160,000 - 240,000
Equity
Medical/dental/vision insurance
Flexible working hours
+1
Technical Lead Platform Engineer
Technical Lead Platform Engineer

Qualified Health • Palo Alto (CA)

Hybrid
USD 180,000 - 300,000
Health insurance
Dental insurance
Vision insurance
+2
Healthcare AI Solutions Engineer
Healthcare AI Solutions Engineer

Qualified Health • Northern (KY)

Hybrid
USD 140,000 - 200,000
Equity
Hybrid work options
Medical/dental/vision insurance
Forward Deployed Engineer
Forward Deployed Engineer

Qualified Health • United States

Hybrid
USD 140,000 - 200,000
Competitive salaries
Equity packages
Robust medical/dental/vision insurance
+2
Senior Backend Engineer
Senior Backend Engineer

Qualified Health PBC • Palo Alto (CA)

Hybrid
USD 160,000 - 240,000
Competitive salaries
Equity packages
Robust medical/dental/vision insurance
+2
Staff / Principal Forward-Deployed Architect, Data Modernization
Staff / Principal Forward-Deployed Architect, Data Modernization

Qualified Health • United States

Hybrid
USD 190,000 - 250,000
Equity packages
Robust medical/dental/vision insurance
Flexible working hours
+1
Senior AI PM for Data and Governance
Senior AI PM for Data and Governance

Qualified Health PBC • Palo Alto (CA)

Hybrid
USD 170,000 - 200,000
Senior AI PM for Data and Governance
Senior AI PM for Data and Governance

Qualified Health • United States

On-site
USD 170,000 - 200,000
Equity
Medical/Dental/Vision insurance
Flexible working hours