Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hippocratic AI is seeking a Senior Site Reliability Engineer to own the GPU management and scheduling platform that runs a fleet of ~30 models on heterogeneous hardware.

You will design metrics, admission control, and autoscaling, build infrastructure automation with Terraform and CI/CD, and operate secure production systems on AWS, GCP, or Azure.

Join a team building healthcare AI at scale, mentoring engineers and collaborating with researchers to ensure reliability, performance, and safety.

Qualifications

  • 10+ years of experience across site reliability, DevOps, and software engineering.
  • Computer Science degree from a top program.
  • Strong software engineering fundamentals with Python and/or Go for orchestration and scheduling.
  • Experience designing systems using operational metrics for autoscaling and admission control.
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD).
  • Hands-on production experience with AWS, GCP, or Azure.
  • Strong knowledge of Docker and Kubernetes.
  • Experience with monitoring/logging stacks (ELK, Grafana, Datadog).
  • Secrets management and security tooling (Vault, KMS, Key Vault).
  • Excellent problem-solving and communication skills.

Responsibilities

  • Design and build GPU management and scheduling platform across ~30 models on heterogeneous hardware.
  • Build metrics pipeline for GPU load and utilization and translate signals into decisions.
  • Implement admission control to protect capacity and manage inference requests.
  • Develop autoscaling for model replicas in real time.
  • Develop cloud orchestration in Python and Go to manage the model fleet.
  • Operate scalable, secure production systems on AWS, GCP, or Azure.
  • Create infrastructure automation and deployment pipelines (Terraform, CI/CD).
  • Establish monitoring, logging, and alerting to ensure reliability.
  • Enforce security/compliance for healthcare AI.
  • Collaborate with researchers to diagnose complex issues; mentor teammates.

Skills

SRE / DevOps
Python / Go
Cloud platforms AWS/GCP/AZ
Docker / Kubernetes
CI/CD tooling Terraform/GitLab
Monitoring & Logging
Security tooling Vault/KMS
Communication
Problem-solving
Autonomous collaboration

Education

Computer Science Degree from a top CS program

Tools

Terraform
GitLab CI/CD
Docker
Kubernetes
ELK
Grafana
Datadog
HashiCorp Vault
AWS KMS
Azure Key Vault

Job description

Hippocratic AI is seeking a Senior Site Reliability Engineer to own the GPU management and scheduling platform that runs a fleet of ~30 models on heterogeneous hardware.

You will design metrics, admission control, and autoscaling, build infrastructure automation with Terraform and CI/CD, and operate secure production systems on AWS, GCP, or Azure.

Join a team building healthcare AI at scale, mentoring engineers and collaborating with researchers to ensure reliability, performance, and safety.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE — GPU Fleet Orchestration & Autoscaling
Senior SRE — GPU Fleet Orchestration & Autoscaling

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior GPU SRE: GPU Scheduling & Autoscaling Platform
Senior GPU SRE: GPU Scheduling & Autoscaling Platform

AI Chopping Block • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Remote GPU Infra SRE for Large-Scale AI Training
Remote GPU Infra SRE for Large-Scale AI Training

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

AI Chopping Block • Menlo Park (CA)

On-site
USD 180,000 - 240,000
SRE - Compute & Hyperscale GPU Fleet Reliability
SRE - Compute & Hyperscale GPU Fleet Reliability

Fluidstack • New York (NY)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Equity participation
Retirement plan
+1
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays