Senior GPU SRE: GPU Scheduling & Autoscaling Platform

AI Chopping Block

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hippocratic AI is seeking a Senior Site Reliability Engineer to own a complex GPU management platform that schedules inference calls across a fleet of ~30 models on heterogeneous hardware. You will build the metrics, admission control, and autoscaling systems, while ensuring secure, scalable, and compliant production environments on cloud providers.

A decade in SRE/DevOps with strong software skills is required.

Qualifications

  • 10+ years of professional experience across SRE/DevOps and software engineering.
  • CS degree from a top program.
  • Strong software engineering fundamentals in Python and/or Go.
  • Experience designing decision systems from operational metrics.
  • Deep experience with infrastructure automation and CI/CD tooling.
  • Production cloud experience (AWS, GCP, or Azure).
  • Hands-on knowledge of containers and orchestration (Docker, Kubernetes).
  • Experience with monitoring/logging stacks (ELK, Grafana, Datadog).
  • Secrets management and security tooling (HashiCorp Vault, KMS).
  • Excellent problem solving and collaboration.

Responsibilities

  • Design and build GPU management and scheduling platform for ~30 models on heterogeneous hardware.
  • Build metrics pipeline and logic to interpret signals for decisions.
  • Implement admission control to protect capacity and queue or shed requests.
  • Create autoscaling for model replicas in real time.
  • Develop cloud orchestration in Python and Go to manage the fleet.
  • Architect scalable, fault-tolerant production systems on cloud platforms.
  • Design infrastructure automation and deployment pipelines (Terraform, CI/CD).
  • Stand up monitoring, logging, and alerting for reliability.
  • Develop security/compliance policies for healthcare AI platform.
  • Collaborate with engineers and researchers to diagnose issues and mentor teammates.

Education

Computer Science degree
Bachelor's or Master's in CS

Tools

Terraform
GitLab CI/CD
Docker
Kubernetes
ELK
Grafana
Datadog
HashiCorp Vault
AWS KMS
Azure Key Vault

Job description

Hippocratic AI is seeking a Senior Site Reliability Engineer to own a complex GPU management platform that schedules inference calls across a fleet of ~30 models on heterogeneous hardware. You will build the metrics, admission control, and autoscaling systems, while ensuring secure, scalable, and compliant production environments on cloud providers.

A decade in SRE/DevOps with strong software skills is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: GPU Fleet Orchestration & Auto-Scaling
Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — GPU Fleet Orchestration & Autoscaling
Senior SRE — GPU Fleet Orchestration & Autoscaling

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE: Global GPU Inference Infra & Reliability
Senior SRE: Global GPU Inference Infra & Reliability

Kindredventures • San Mateo (CA)

On-site
USD 150,000 - 240,000
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Remote GPU Infra SRE for Large-Scale AI Training
Remote GPU Infra SRE for Large-Scale AI Training

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Senior Infrastructure Engineer, AI GPU Cloud Orchestration
Senior Infrastructure Engineer, AI GPU Cloud Orchestration

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000