Staff Site Reliability Engineer

Hippocratic AI

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hippocratic AI is seeking a Senior Site Reliability Engineer to own the GPU management and scheduling platform that runs a fleet of ~30 models on heterogeneous hardware.

You will design metrics, admission control, and autoscaling, build infrastructure automation with Terraform and CI/CD, and operate secure production systems on AWS, GCP, or Azure.

Join a team building healthcare AI at scale, mentoring engineers and collaborating with researchers to ensure reliability, performance, and safety.

Qualifications

  • 10+ years of experience across site reliability, DevOps, and software engineering.
  • Computer Science degree from a top program.
  • Strong software engineering fundamentals with Python and/or Go for orchestration and scheduling.
  • Experience designing systems using operational metrics for autoscaling and admission control.
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD).
  • Hands-on production experience with AWS, GCP, or Azure.
  • Strong knowledge of Docker and Kubernetes.
  • Experience with monitoring/logging stacks (ELK, Grafana, Datadog).
  • Secrets management and security tooling (Vault, KMS, Key Vault).
  • Excellent problem-solving and communication skills.

Responsibilities

  • Design and build GPU management and scheduling platform across ~30 models on heterogeneous hardware.
  • Build metrics pipeline for GPU load and utilization and translate signals into decisions.
  • Implement admission control to protect capacity and manage inference requests.
  • Develop autoscaling for model replicas in real time.
  • Develop cloud orchestration in Python and Go to manage the model fleet.
  • Operate scalable, secure production systems on AWS, GCP, or Azure.
  • Create infrastructure automation and deployment pipelines (Terraform, CI/CD).
  • Establish monitoring, logging, and alerting to ensure reliability.
  • Enforce security/compliance for healthcare AI.
  • Collaborate with researchers to diagnose complex issues; mentor teammates.

Skills

SRE / DevOps
Python / Go
Cloud platforms AWS/GCP/AZ
Docker / Kubernetes
CI/CD tooling Terraform/GitLab
Monitoring & Logging
Security tooling Vault/KMS
Communication
Problem-solving
Autonomous collaboration

Education

Computer Science Degree from a top CS program

Tools

Terraform
GitLab CI/CD
Docker
Kubernetes
ELK
Grafana
Datadog
HashiCorp Vault
AWS KMS
Azure Key Vault

Job description

About The Role

We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models. We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts. This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

What You'll Do
  • Design and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware
  • Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions
  • Implement admission control to protect capacity — deciding when to accept, queue, or shed inference requests so we operate within fleet limits
  • Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization
  • Develop cloud orchestration systems and operators in Python and Go to manage the model fleet
  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure
  • Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software
  • Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant
  • Develop and enforce security and compliance policies appropriate to a healthcare AI platform
  • Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues
  • Mentor engineers and raise the technical bar across the team
What You Bring
Must-Have
  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering
  • Computer Science Degree Required from a top CS program.
  • Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools
  • Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
  • Hands‑on production experience with at least one major cloud platform (AWS, GCP, or Azure)
  • Strong knowledge of containerization and orchestration (Docker, Kubernetes)
  • Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)
  • Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)
  • Excellent problem‑solving skills and the ability to work both independently and collaboratively
  • Strong communication and interpersonal skills
Nice-to-Have
  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
  • Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)
  • Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)
  • Experience implementing HIPAA and SOC 2 compliance
  • Experience operating in an HPC environment
  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field

Join our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.

Why Join Hippocratic AI

We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact people@hippocraticai.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

AI Chopping Block • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer, Research
Senior Software Engineer, Research

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 230,000
Forward Deployed Engineer (Mid/Senior)
Forward Deployed Engineer (Mid/Senior)

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 260,000
Senior Product Data Analyst
Senior Product Data Analyst

Precision Labs • Menlo Park (CA)

Hybrid
USD 140,000 - 190,000
Senior Product Data Analyst
Senior Product Data Analyst

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 120,000 - 170,000
Lead Product Data Analyst
Lead Product Data Analyst

Hippocratic AI Inc. • Menlo Park (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Product Data Analyst
Senior Product Data Analyst

Hippocratic AI • Menlo Park (CA)

On-site
USD 140,000 - 190,000
Associate Chief Medical Officer (East Coast)
Associate Chief Medical Officer (East Coast)

Hippocratic AI • United States

On-site
USD 210,000 - 360,000
Director of Quality
Director of Quality

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 210,000 - 320,000