Head of AI Inference & Systems Platform

Blackhornvc

San Francisco (CA)

On-site

USD 220,000 - 420,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Canyon Code is looking for a Head of Inference to own the end-to-end inference stack, from procuring bare-metal GPU capacity to serving production endpoints. You will work directly with the CEO and Chief Architect, hire and mentor the team, and shape the hardware mix, providers, and serving stack.

This is a hands-on leadership role in a fast-moving, early-stage environment, requiring deep expertise in multi-model inference and Kubernetes on bare metal.

Qualifications

  • End-to-end experience procuring bare-metal GPU capacity and bringing open-weight models to production endpoints.
  • Hands-on expertise with multi-model inference systems and cost-per-token optimization.
  • Deep systems/ops skills: Kubernetes, containers, networking, storage, autoscaling, observability on bare metal.

Responsibilities

  • Own the entire inference path from capacity procurement to serving endpoints.Build and mentor the inference team as the motion scales.Collaborate with the CEO and Chief Architect to select providers, hardware mix, and serving stack.

Job description

Canyon Code is looking for a Head of Inference to own the end-to-end inference stack, from procuring bare-metal GPU capacity to serving production endpoints. You will work directly with the CEO and Chief Architect, hire and mentor the team, and shape the hardware mix, providers, and serving stack.

This is a hands-on leadership role in a fast-moving, early-stage environment, requiring deep expertise in multi-model inference and Kubernetes on bare metal.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Inference
Head of Inference

Blackhornvc • San Francisco (CA)

On-site
USD 220,000 - 420,000
Head of Inference Infrastructure
Head of Inference Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Chief Architect, Cluster-Scale AI Inference
Chief Architect, Cluster-Scale AI Inference

The Consensus • San Jose (CA)

On-site
USD 260,000 - 520,000
Medical, dental, and vision packages
Housing subsidy near Santana Row
Relocation support to San Jose
+3
Founding AI Inference Architect
Founding AI Inference Architect

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Principal Engineer, Inference Cloud
Principal Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture
Director, AI Inference & Model Scaling
Director, AI Inference & Model Scaling

Cerebras • Sunnyvale (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000