Head of Infrastructure

General Compute

San Francisco (CA)

On-site

USD 190,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

General Compute is building an infrastructure-first inference cloud in San Francisco. You will own the control plane, gateway, and observability, shaping how we route requests across ASICs and GPUs while managing capacity and reliability.

The role grows to manage partnerships with ASIC vendors and expand across multiple hardware platforms. The ideal candidate has 7+ years in infra/SRE, hands-on production Kubernetes, and a proven track record with tail-latency optimization.

Qualifications

  • 7+ years in infrastructure, SRE, or platform engineering for inference/ML/HPC.
  • Hands-on production Kubernetes at scale with debugging skills.
  • Strong tail-latency instincts focusing on p99 and utilization.
  • Comfortable owning vendor relationships and production issues.
  • Proven observability practices that catch problems, not just dashboards.
  • Experience on-call through real incidents and learning outcomes.
  • Excited to be the first infra hire at a growing startup.

Responsibilities

  • Own the inference control plane, including possible replacement of components.
  • Own the gateway and load balancer for model placement and routing.
  • Own end-to-end observability with per-request tracing and SLO dashboards.
  • Run capacity planning across a distributed traffic mix of open-weight models.
  • Own operational aspects of ASIC partnerships and partner engineering touchpoints.
  • Bring up the pre-fill side of the disaggregated architecture on new hardware.
  • Build on-call and incident response practices and grow the team.

Skills

SRE leadership
Tail latency optimization
Incident response
Vendor management
Observability design
On-call experience
Early-stage infra

Tools

Kubernetes
OpenTelemetry
Load balancing
RoCE/InfiniBand

Job description

About Us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the role

You'll own the infrastructure layer of our inference cloud end-to-end. Today that means the control plane, the gateway in front of our ASIC fleet, and the observability stack that tells us where every millisecond goes. Over the next 6-8 months, it will grow into a heterogeneous fleet: ASICs for decode, GPUs for pre-fill, and the physical-layer ownership that comes with it.

The first six months are hands-on: k8s manifests, dashboards, oncall, and a direct line to our ASIC partner's engineering team when production behaves strangely. The team grows under you from there.

What you'll do:
  • Own the inference control plane. At the moment, it's built on configuration provided by our ASIC partner; you'll be the person who understands it deeply enough to modify, extend, and eventually replace pieces of it.

  • Own the gateway and load balancer that fronts the fleet. Model placement, request routing, and tail-latency engineering live here, driven by live utilization and per-model SLOs.

  • Own observability end-to-end. Per-request tracing from OpenRouter ingress through to the accelerator, with p50/p95/p99 dashboards, SLOs, and alerting that wakes the right person.

  • Run capacity planning against a real, distributed traffic mix across the open-weight models we serve.

  • Own the operational side of the ASIC partnership. Most weird production issues route through their engineering team until we build that expertise in-house, and you'll be our technical face in those conversations.

  • Bring up the pre-fill side of our disaggregated architecture on a second hardware platform as it comes online. Different vendor, different fabric, different kernels.

  • Build the on-call and incident response practice from zero. Hire and grow the team underneath you.

What we need from you:
  • 7+ years in infrastructure, SRE, or platform engineering, with at least some of it at a serious inference, ML, or HPC shop.

  • Hands-on with Kubernetes at production scale — not just deploying, but debugging the weird stuff.

  • Strong instincts for tail latency. You think about p99 and utilization as the same problem, not different ones.

  • Comfortable owning a vendor relationship where the vendor's bugs are now your production issues.

  • Track record of building observability practices that actually catch problems, not just generate dashboards.

  • Have been on-call through real incidents and can talk about what you learned.

  • Want to be the first infra hire at something early, not the tenth at something big.

NIce-to-Haves:
  • Experience operating non-NVIDIA accelerators in production — TPUs, ASICs, or alternative GPU vendors.

  • Background with model-serving stacks (vLLM, TGI, TensorRT-LLM, SGLang).

  • Network fabric experience at data-center scale (RoCE, InfiniBand).

  • Have hired and managed an infra team before.

  • Comfort at the hardware boundary — firmware, drivers, thermals — for when the roadmap takes us there.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Infrastructure
Head of Infrastructure

General Compute Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Model Bring-up Engineer / ML Compiler Engineer
Model Bring-up Engineer / ML Compiler Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 240,000
Infrastructure Engineer, Perimeter Compute
Infrastructure Engineer, Perimeter Compute

Montauk Capital • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation
Equity options
Studio support from Montauk Capital
Infrastructure Engineer
Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 390,000
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Inference Engineer
Inference Engineer

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000