Senior SRE: Global GPU Inference Infra & Reliability

Kindredventures

San Mateo (CA)

On-site

USD 150,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

You’ll own production reliability end-to-end, from capacity planning to incident response, with a focus on scalable, observable, and secure infrastructure as inference workloads grow.

Qualifications

  • Experience building and operating production infrastructure or distributed systems with real ownership of reliability.
  • Strong Linux fundamentals and practical knowledge of networking, storage, and containers.
  • Hands-on experience running Kubernetes in production.
  • The ability to write maintainable software and automation to solve infrastructure problems.
  • A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries.
  • Good judgment about when to move quickly, when to simplify, and where reliability matters most.
  • The initiative to take a problem from investigation through implementation and work closely with teammates along the way.

Responsibilities

  • Scale a global GPU fleet by building and improving the Kubernetes infrastructure behind provisioning, networking, storage, and service deployment across providers and regions.
  • Design better isolation, failover, and recovery so hardware and infrastructure failures have less impact on customers.
  • Automate capacity expansion, deployments, and maintenance, eliminating manual work and making changes safer.
  • Develop observability and diagnostics that reveal bottlenecks, surface failures, and help engineers act quickly.
  • Respond to incidents, diagnose root causes, and translate learnings into stronger systems.
  • Collaborate across the stack to improve performance, utilization, security, and reliability as inference demand grows.

Skills

Production infra ownership
Linux fundamentals
Networking & storage
Kubernetes in prod
Automation & scripting
System observability

Tools

Kubernetes
CI/CD tools
Infrastructure as Code
Monitoring tools

Job description

You’ll own production reliability end-to-end, from capacity planning to incident response, with a focus on scalable, observable, and secure infrastructure as inference workloads grow.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior GPU SRE: GPU Scheduling & Autoscaling Platform
Senior GPU SRE: GPU Scheduling & Autoscaling Platform

AI Chopping Block • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Remote GPU Infra SRE for Large-Scale AI Training
Remote GPU Infra SRE for Large-Scale AI Training

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE: GPU Fleet Orchestration & Auto-Scaling
Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior SRE — GPU Fleet Orchestration & Autoscaling
Senior SRE — GPU Fleet Orchestration & Autoscaling

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 180,000 - 240,000
GPU Infrastructure Sourcing & Procurement Lead
GPU Infrastructure Sourcing & Procurement Lead

Kindredventures • San Mateo (CA)

On-site
USD 180,000 - 260,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000