Senior SRE — GPU Fleet Orchestration & Autoscaling

Hippocratic-Ai

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hippocratic AI in Menlo Park, CA, seeks a Senior Site Reliability Engineer who writes production software and runs the infrastructure it lives on. You will own a GPU management and scheduling platform for ~30 models on heterogeneous hardware, shaping admission control, autoscaling, and reliability.

You'll mentor engineers, design scalable systems on AWS/GCP/Azure, implement Terraform pipelines and CI/CD, and ensure HIPAA-ready security and observability.

Qualifications

  • 10+ years in SRE/DevOps and software engineering.
  • CS degree from a top program is required.
  • Strong Python/Go orchestration and scheduling experience.
  • Experience building systems from operational metrics to control loops.
  • Deep infra automation and CI/CD experience with Terraform or similar.
  • Hands-on production cloud experience (AWS/GCP/Azure).

Responsibilities

  • Design and build GPU management and scheduling platform for ~30 models.
  • Build metrics pipeline and decision logic from GPU signals.
  • Implement admission control to protect capacity and autoscaling.
  • Develop cloud orchestration in Python/Go and automate deployments.
  • Architect scalable, fault-tolerant production systems on cloud.
  • Establish monitoring, logging, and alerting for reliability.
  • Ensure HIPAA/compliance readiness and security best practices.
  • Mentor engineers and raise technical bar across the team.

Skills

10+ years experience
Strong CS fundamentals
Scheduling systems in Python/Go
Autoscaling & load shedding
CI/CD & infra automation
Cloud platforms (AWS/GCP/Azure)
Containerization & orchestration
Monitoring & logging stacks
Security tooling / secrets management
Problem solving & communication

Education

Computer Science Degree

Tools

Terraform
GitLab CI/CD
Docker
Kubernetes
ELK
Grafana
Datadog
HashiCorp Vault
AWS
GCP
Azure
KMS

Job description

Hippocratic AI in Menlo Park, CA, seeks a Senior Site Reliability Engineer who writes production software and runs the infrastructure it lives on. You will own a GPU management and scheduling platform for ~30 models on heterogeneous hardware, shaping admission control, autoscaling, and reliability.

You'll mentor engineers, design scalable systems on AWS/GCP/Azure, implement Terraform pipelines and CI/CD, and ensure HIPAA-ready security and observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: GPU Fleet Orchestration & Auto-Scaling
Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior GPU SRE: GPU Scheduling & Autoscaling Platform
Senior GPU SRE: GPU Scheduling & Autoscaling Platform

AI Chopping Block • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
SRE - Compute & Hyperscale GPU Fleet Reliability
SRE - Compute & Hyperscale GPU Fleet Reliability

Fluidstack • New York (NY)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Equity participation
Retirement plan
+1
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000