VP of Engineering

Hyperbolic

San Francisco (CA)

On-site

USD 200,000 - 300,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hyperbolic is in search of a Vice President of Infrastructure to lead the development of foundational infrastructure for our AI cloud platform in San Francisco. This hands-on executive role requires overseeing infrastructure strategy while actively participating in technical design and engineering initiatives.

The ideal candidate will have extensive experience in building cloud platforms, especially GPU-native infrastructure, and will be involved in critical architectural decisions and team leadership to drive innovation and performance.

Qualifications

  • 12+ years in infrastructure systems development and operations.
  • Experience in building or operating cloud platforms at scale.
  • Expertise in Kubernetes and GPU infrastructure.

Responsibilities

  • Build and scale infrastructure for AI cloud platform.
  • Engage in architecture and design directly.
  • Define architecture for GPU orchestration and distributed systems.

Skills

Infrastructure scaling
Kubernetes
Cloud platform experience
AI/ML compute platforms
Reliability engineering
Linux and networking

Job description

Requirements
  • 12+ years building and operating large-scale infrastructure systems
  • Experience leading infrastructure organizations while remaining hands‑on technically
  • Previous experience building or operating a cloud platform at scale
  • Experience building GPU infrastructure or AI/ML compute platforms
  • Proven track record scaling infrastructure in high‑growth startup environments
  • Expert-level Kubernetes knowledge
  • Experience designing and operating multi‑region cloud infrastructure
  • Strong understanding of Linux, networking, distributed systems, and storage architecture
  • Experience with Infrastructure‑as‑Code and automation frameworks
  • Deep expertise in observability, monitoring, and reliability engineering
  • Experience building highly available production systems
  • (Desirable) Experience with GPU scheduling, Slurm, Kubernetes GPU operators, Ray, or distributed training systems
  • (Desirable) Experience managing thousands of GPUs in production environments
  • (Desirable) Background supporting AI training and inference platforms
What the job involves
  • We are seeking a highly technical Vice President of Infrastructure to build and scale the foundational infrastructure powering our AI cloud platform
  • This is a hands‑on executive leadership role
  • While you will own infrastructure strategy, organizational growth, and executive‑level decision making, we expect you to remain deeply engaged in architecture, design, and engineering execution
  • You should expect to spend approximately 30-40% of your time directly contributing to technical design, architecture reviews, debugging critical production issues, and partnering with engineers on implementation
  • The ideal candidate has previously built and scaled cloud platforms, preferably GPU‑native cloud infrastructure supporting AI training and inference workloads
  • You have experience operating at the intersection of executive leadership and hands‑on engineering and are excited to help build both the technology and the team
  • Lead the design and evolution of our AI cloud platform
  • Define the architecture for GPU orchestration, compute scheduling, networking, storage, and distributed systems
  • Make critical decisions regarding cloud infrastructure, bare‑metal deployments, and platform scalability
  • Personally participate in architecture reviews and key technical initiatives
  • Build and scale large GPU clusters supporting customer workloads
  • Design systems for GPU provisioning, scheduling, utilization optimization, and capacity management
  • Drive platform reliability and performance for AI training and inference workloads
  • Partner closely with engineering teams on infrastructure requirements for next generation AI systems
  • Remain deeply involved in engineering decisions and technical direction
  • Contribute directly to infrastructure design and implementation efforts
  • Review architecture proposals, system designs, and major infrastructure changes
  • Act as the technical escalation point for complex infrastructure challenges
  • Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence
  • Build SRE and Platform Engineering functions from the ground up
  • Define reliability standards including SLOs, SLIs, incident response processes, and capacity planning
  • Drive automation across infrastructure operations
  • Recruit and develop world‑class Infrastructure, Platform, and SRE teams
  • Build a high‑performance engineering culture focused on ownership and execution
  • Partner with executive leadership on company strategy and infrastructure investments
  • Manage infrastructure budgets, vendor relationships, and capacity planning
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
VP AI - Core Infrastructure
VP AI - Core Infrastructure

kadence • New York (NY)

Hybrid
USD 250,000 - 350,000
Data Center Business Technical SME
Data Center Business Technical SME

Talentstra • Town of Texas (WI)

On-site
USD 200,000 - 280,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
VP of AI Cloud Infrastructure & GPU Compute
VP of AI Cloud Infrastructure & GPU Compute

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000