Head of GPU Cloud

Blue Signal Search

United States

On-site

USD 200,000 - 300,000

Full time

6 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Blue Signal Search is seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming GPU capacity into high-performance cloud services for AI workloads. This role shapes platform architecture, production inference, developer experiences, and engineering strategy at a pivotal growth stage.

You will own Kubernetes-based infrastructure, model serving, and optimization across vLLM, TensorRT LLM, and related stacks, balancing throughput, latency, and cost across

Qualifications

  • Extensive leadership of engineering teams delivering mission-critical cloud infrastructure.
  • Proven track record optimizing large-scale AI/model serving workloads.
  • Deep experience with GPU accelerators, memory management, and high-throughput scheduling.

Responsibilities

  • Lead the engineering org behind a nationwide GPU cloud platform.
  • Design orchestration policies matching workloads to accelerators by memory, performance, and cost.
  • Own Kubernetes-based infrastructure strategy including multi-tenant concerns and reliability.
  • Guide product and platform decisions affecting inference services and developer UX.

Skills

GPU cloud leadership
Cloud platforms
Distributed infrastructure
Kubernetes expertise
Model serving optimization
High-performance computing
AI infrastructure
Resource allocation
Memory-aware scheduling
Platform reliability

Tools

Kubernetes
TensorRT LLM
vLLM
TGI
Infra tooling

Job description

Our client is building the software foundation for a next generation accelerated computing platform designed to make large-scale AI infrastructure easier to consume, operate, and scale. They are seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming significant GPU capacity into high-performance cloud services for AI workloads. This is an opportunity to shape platform architecture, production inference, developer experiences, and engineering strategy at a stage where technical decisions will directly influence customer adoption, infrastructure economics, and long-term growth.

This Role Offers

  • Opportunity to define the architecture and operating model behind large scale AI inference services.
  • Direct influence over GPU utilization, platform economics, customer experience, and technical strategy.
  • Close collaboration with leaders across infrastructure, networking, product, operations, and commercial functions.
  • A highly technical environment where software engineering intersects with accelerated computing, distributed systems, and AI infrastructure.

Focus

  • Build and scale the engineering organization behind a nationwide GPU cloud platform, with ownership across inference services, orchestration, APIs, platform reliability, and technical execution.
  • Set the architecture for moving AI workloads efficiently from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory aware scheduling.
  • Lead production model serving and optimization across technologies such as vLLM, TensorRT LLM, and TGI, improving throughput, latency, availability, accelerator utilization, and cost.
  • Own Kubernetes based GPU infrastructure strategy across workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience.
  • Develop platform capabilities for hosted models, private inference environments, customized deployments, usage measurement, customer controls, and developer facing services.
  • Design orchestration policies that match workloads to accelerators based on memory requirements, performance objectives, capacity, cluster topology, and infrastructure economics.
  • Partner with networking, systems, data center, product, and commercial leaders to align software decisions with accelerator architecture, fabric performance, storage, and customer requirements.
  • Establish engineering standards, service objectives, capacity planning, incident readiness, and team accountability while recruiting and developing senior technical talent.

Skill Set

  • 12 or more years of progressive software engineering experience, including substantial leadership responsibility across cloud platforms, distributed infrastructure, HPC, or similarly complex production systems.
  • 5 or more years leading engineering teams responsible for business critical infrastructure, platform services, or other mission critical technology products.
  • Demonstrated production experience operating large scale model inference using vLLM, TensorRT LLM, TGI, or equivalent serving stacks.
  • Strong expertise in model serving optimization, including dynamic batching, decoding acceleration, reduced precision execution, compilation, memory reuse, and request scheduling.
  • Advanced knowledge of Kubernetes and containerized infrastructure, including scheduling, elasticity, telemetry, deployment practices, and production reliability.
  • Practical accelerator infrastructure knowledge covering GPU memory behavior, high bandwidth networking, InfiniBand, RoCE, storage performance, and cluster topology.
  • Experience architecting highly available distributed platforms with automated resource allocation, programmatic interfaces, multi customer support, and detailed consumption measurement.
  • Strong technical and leadership judgment with the ability to balance performance, reliability, security, customer experience, and infrastructure economics across multidisciplinary teams.

Additional Experience That Stands Out

  • Leadership experience in GPU cloud, AI infrastructure, hosted model platforms, or accelerated computing environments.
  • Experience delivering elastic inference services, dedicated AI capacity, model customization workflows, or managed AI products.
  • Familiarity with open model ecosystems and the operational differences among model families, serving configurations, and hardware profiles.
  • Track record improving accelerator utilization, workload density, capacity forecasting, and compute economics across multiple GPU generations.
  • Experience scaling engineering organizations in fast moving environments where software requirements and infrastructure capacity evolve together.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VP of Engineering
VP of Engineering

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000