Head of Inference

Blackhornvc

San Francisco (CA)

On-site

USD 220,000 - 420,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Canyon Code is looking for a Head of Inference to own the end-to-end inference stack, from procuring bare-metal GPU capacity to serving production endpoints. You will work directly with the CEO and Chief Architect, hire and mentor the team, and shape the hardware mix, providers, and serving stack.

This is a hands-on leadership role in a fast-moving, early-stage environment, requiring deep expertise in multi-model inference and Kubernetes on bare metal.

Qualifications

  • End-to-end experience procuring bare-metal GPU capacity and bringing open-weight models to production endpoints.
  • Hands-on expertise with multi-model inference systems and cost-per-token optimization.
  • Deep systems/ops skills: Kubernetes, containers, networking, storage, autoscaling, observability on bare metal.

Responsibilities

  • Own the entire inference path from capacity procurement to serving endpoints.Build and mentor the inference team as the motion scales.Collaborate with the CEO and Chief Architect to select providers, hardware mix, and serving stack.

Job description

Why This Role, Why Now

Running enterprise AI workflows is one of the four pillars of Canyon Code, alongside Design, Deploy, and Optimize. Run is where we take open-weight models, put them on capacity we have procured, and serve the endpoints our customers point their workflows at.

We are ready for someone to own that pillar full time. This is a founding-team role, and the person in it will own the inferencing motion at Canyon Code and build the team underneath them.

About the Role

We are looking for a Head of Inference to own the whole path, working directly with Canyon Code's CEO and Chief Architect: procuring bare-metal GPU capacity from neoclouds and colocation facilities, bringing open-weight models up on that capacity, and serving the endpoints customers actually run against. This is hands‑on. You will be in the system, not only directing it, and you will hire the team under you as the motion grows.

You're a fit if you…
  • You have done this end to end, more than once, within the last couple of years: procured bare-metal GPU capacity from a neocloud or colocation facility, brought open-weight models up on it, and served production endpoints off it.

  • You understand this concretely rather than in the abstract. You know the nuances, where it goes wrong, and how long each step actually takes.

  • You have procured bare-metal GPU capacity yourself: evaluating providers, negotiating terms, taking delivery, and getting it production-ready.

  • Deep hands‑on experience with vLLM, SGLang, TensorRT-LLM, or NVIDIA Dynamo.

  • You understand KV-cache management, continuous batching, paged attention, speculative decoding, and quantization, and you know which of them actually pay off for a given workload.

  • You have owned a cost-per-token number and moved it.

  • Strong systems and operations skills: Kubernetes, containers, networking, storage, autoscaling, and observability. On bare metal you own more of the stack than you would on a hyperscaler, and you are comfortable with that.

  • Comfortable across heterogeneous accelerators: NVIDIA plus at least one of AMD, Groq, Cerebras, Intel, or TPU.

  • You can hire, mentor, and build a team under you.

  • Self‑driven; thrives in ill‑defined, ambiguous, early‑stage work.

  • You have done this inside an inference provider such as Baseten, Fireworks, Lightning, Together, or Modal. Plus

  • Existing relationships with neoclouds or colocation providers. Plus

  • You have worked on multi‑agentic workloads, where a single request fans out across many model calls. Plus

  • Published or open-source work in inference optimization. Plus

Why Join

Own a pillar

Run is one of the four things Canyon Code is built on. This is not a supporting function; it is the pillar that has to work for the rest of the vision to hold.

Founding team

You own the inferencing motion and build the team under you. The architecture, the providers, the hardware mix, and the hiring are yours.

A proven advantage to compound

CanyonOS is already 3x better at harness scaling and 3x better at LLM serving. You start from a measured edge rather than a hypothesis.

The buy-versus-build calls are yours

Capacity, providers, and serving stack decisions sit with you, with the money in the bank to act on them.

Work alongside the founding team

Daily collaboration with a CEO and Chief Architect who have deep experience in ML systems research and AI infrastructure platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Business Development
Head of Business Development

Blackhornvc • San Francisco (CA)

On-site
USD 180,000 - 270,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Head of Research
Head of Research

Blackhornvc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of Inference, Perimeter Compute
Head of Inference, Perimeter Compute

Montauk Capital • New York (NY)

On-site
USD 150,000 - 200,000
Competitive compensation + equity
Studio support from Montauk Capital’s network
Head of AI Inference & Systems Platform
Head of AI Inference & Systems Platform

Blackhornvc • San Francisco (CA)

On-site
USD 220,000 - 420,000
Founding Platform Engineer
Founding Platform Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Head of Infrastructure
Head of Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Founding Engineer - ML Performance
Founding Engineer - ML Performance

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000