HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 235,000 - 315,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Founding engineer equity
Full benefits package

Job summary

Hamilton Barnes Associates Limited is building a founding engineering team for a VC-backed GPU cloud platform in San Francisco. This hands-on role owns core infrastructure spanning Kubernetes, Slurm, GPU orchestration, and high-speed storage, collaborating with founders to shape architecture and scale production AI systems.

You will influence the platform's design, reliability, and performance while helping define engineering culture and hiring as an early team member.

Qualifications

  • 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production.
  • Proven experience operating large-scale Kubernetes and Slurm clusters.
  • Experience building and managing GPU orchestration layers on top of core schedulers.
  • Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking.
  • Deep knowledge of storage architectures for large-scale AI infrastructure.
  • Comfort operating with founding-level ownership across the full infrastructure stack.
  • Based in or willing to relocate to San Francisco.

Responsibilities

  • Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
  • Build and manage GPU orchestration layers on top of core scheduling infrastructure
  • Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
  • Design and manage high-bandwidth networking supporting distributed training and inference at scale
  • Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
  • Build telemetry, observability, and automated remediation across the GPU fleet
  • Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
  • Own operational health, reliability, and performance of the platform end to end
  • Work directly with founders on architecture, roadmap, and technical strategy
  • Help define engineering culture, standards, and hiring as one of the first technical team members

Skills

Kubernetes
Slurm
GPU orchestration
Distributed storage
NVMe storage
High-bandwidth networking
Telemetry/observability
Founding-level ownership
Willing to relocate to SF

Tools

Kubernetes
Slurm
NVMe storage
Distributed storage systems
Telemetry tooling
Networking knowledge

Job description

Ready to take the next step in your career?

Join a VC-backed GPU cloud company at the founding stage, building the next generation of managed GPU compute and inference infrastructure. Backed by leading venture investors, the organization delivers fully managed AI infrastructure and offers the opportunity to help shape the technical foundation of a rapidly growing GPU cloud platform.

This company is seeking a Founding Engineer to take ownership of core GPU cloud infrastructure across Kubernetes, Slurm, GPU orchestration, distributed storage, high-bandwidth networking, and inference platforms. This hands-on role offers founding-level ownership, significant influence over platform architecture, and the opportunity to work closely with the founders to build and scale production AI infrastructure from the ground up.

Responsibilities:
  • Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
  • Build and manage GPU orchestration layers on top of core scheduling infrastructure
  • Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
  • Design and manage high-bandwidth networking supporting distributed training and inference at scale
  • Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
  • Build telemetry, observability, and automated remediation across the GPU fleet
  • Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
  • Own operational health, reliability, and performance of the platform end to end
  • Work directly with founders on architecture, roadmap, and technical strategy
  • Help define engineering culture, standards, and hiring as one of the first technical team members
Skills/Must Have:
  • 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production
  • Proven experience operating large-scale Kubernetes and Slurm clusters
  • Experience building and managing GPU orchestration layers on top of core schedulers
  • Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking
  • Deep knowledge of storage architectures for large-scale AI infrastructure
  • Comfort operating with founding-level ownership across the full infrastructure stack
  • Based in or willing to relocate to San Francisco
Benefits:
  • Founding engineer equity
  • Full benefits package
Salary:
  • $275,000 Base
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Cash bonus
Founding engineer equity
Benefits
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Solution Architect - AI Infrastructure
Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco