Head of Inference Infrastructure

General Compute

San Francisco (CA)

On-site

USD 190,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

General Compute is building an infrastructure-first inference cloud in San Francisco. You will own the control plane, gateway, and observability, shaping how we route requests across ASICs and GPUs while managing capacity and reliability.

The role grows to manage partnerships with ASIC vendors and expand across multiple hardware platforms. The ideal candidate has 7+ years in infra/SRE, hands-on production Kubernetes, and a proven track record with tail-latency optimization.

Qualifications

  • 7+ years in infrastructure, SRE, or platform engineering for inference/ML/HPC.
  • Hands-on production Kubernetes at scale with debugging skills.
  • Strong tail-latency instincts focusing on p99 and utilization.
  • Comfortable owning vendor relationships and production issues.
  • Proven observability practices that catch problems, not just dashboards.
  • Experience on-call through real incidents and learning outcomes.
  • Excited to be the first infra hire at a growing startup.

Responsibilities

  • Own the inference control plane, including possible replacement of components.
  • Own the gateway and load balancer for model placement and routing.
  • Own end-to-end observability with per-request tracing and SLO dashboards.
  • Run capacity planning across a distributed traffic mix of open-weight models.
  • Own operational aspects of ASIC partnerships and partner engineering touchpoints.
  • Bring up the pre-fill side of the disaggregated architecture on new hardware.
  • Build on-call and incident response practices and grow the team.

Skills

SRE leadership
Tail latency optimization
Incident response
Vendor management
Observability design
On-call experience
Early-stage infra

Tools

Kubernetes
OpenTelemetry
Load balancing
RoCE/InfiniBand

Job description

General Compute is building an infrastructure-first inference cloud in San Francisco. You will own the control plane, gateway, and observability, shaping how we route requests across ASICs and GPUs while managing capacity and reliability.

The role grows to manage partnerships with ASIC vendors and expand across multiple hardware platforms. The ideal candidate has 7+ years in infra/SRE, hands-on production Kubernetes, and a proven track record with tail-latency optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of Infrastructure
Head of Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Founding Platform Engineer
Founding Platform Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding AI Inference Architect
Founding AI Inference Architect

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Inference Engineer
Inference Engineer

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000