Infrastructure Product Engineer - AI Infrastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 300,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

A stealth-mode startup is seeking a Senior Infrastructure Product Engineer to architect large-scale GPU platforms for training and inference. The candidate will build automation frameworks and identify systemic bottlenecks. The ideal applicant should have over 7 years of experience in relevant engineering roles, with expertise in Kubernetes and High Performance Computing. This is a unique opportunity to influence how AI infrastructure is built. The position offers a salary range of $300,000 to $350,000 per year with equity benefits.

Qualifications

  • 7+ years of experience in Infrastructure Engineering, Platform Engineering, SRE, or Systems Architecture roles.
  • Proven experience designing and operating large-scale GPU or HPC platforms.
  • Deep hands-on expertise with Kubernetes and Slurm, including scheduler behaviour and workload optimisation.

Responsibilities

  • Architect and evolve large-scale GPU platforms to support training, inference, and emerging AI workloads.
  • Build scalable automation frameworks for provisioning, scheduling, and lifecycle management.
  • Identify systemic bottlenecks across compute, network, and storage layers.

Skills

Infrastructure Engineering
Platform Engineering
SRE
Systems Architecture
Kubernetes
Linux systems
Python
Go
Bash
Observability platforms

Job description

Join a stealth-mode startup building a next-generation AI and cloud platform powered by thousands of H100s, H200s, and B200s, designed for rapid experimentation, full-scale model training, and production inference. As a Senior Infrastructure Product Engineer, you’ll sit at the intersection of platform architecture, product thinking, and large-scale systems engineering, shaping how AI infrastructure is exposed, consumed, and scaled.

This role goes beyond keeping systems running. You’ll architect the underlying primitives that power new infrastructure products, defining how compute, networking, scheduling, and observability come together as a coherent platform. You’ll work closely with product, ML, and hardware teams to turn raw GPU capacity into reliable, developer-friendly capabilities.

If you want to architect infrastructure as a product, define the building blocks behind frontier AI platforms, and influence how thousands of GPUs are consumed at scale, this is a rare chance to do it from first principles.

Get in touch and apply today!

Responsibilities:
  • Architect and evolve large-scale GPU platforms (H100/H200/B200) to support training, inference, and emerging AI workloads.
  • Design infrastructure abstractions and platform primitives that enable new AI and cloud products.
  • Build scalable automation frameworks for provisioning, scheduling, and lifecycle management across Slurm, Kubernetes, and bare-metal environments.
  • Partner with product and ML teams to translate user requirements into infrastructure architecture and platform capabilities.
  • Define reliability, scalability, and performance standards as architectural constraints rather than reactive fixes.
  • Develop observability and capacity models that inform platform design, roadmap decisions, and customer-facing SLAs.
  • Identify systemic bottlenecks across compute, network, and storage layers and drive architectural improvements.
Skills/Must have:
  • 7+ years of experience in Infrastructure Engineering, Platform Engineering, SRE, or Systems Architecture roles.
  • Proven experience designing and operating large-scale GPU or HPC platforms.
  • Deep hands-on expertise with Kubernetes and Slurm, including scheduler behaviour and workload optimisation.
  • Strong Linux systems and networking fundamentals in high-performance environments.
  • Proficiency in Python, Go, or Bash for building platform tooling and automation.
  • Experience treating infrastructure as a product, with a focus on usability, interfaces, and scalability.
  • Familiarity with observability platforms (Prometheus, Grafana, Loki) and performance analysis at scale.
Benefits:
  • Equity
Salary:
  • $300,000 to $350,000 gross per year
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Product Manager (AI Infrastructure) - Hosting
Product Manager (AI Infrastructure) - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock Options
Company Bonus
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Head of Product - AI Infrastructure
Head of Product - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 360,000 - 440,000
Equity
Bonus
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Solution Architect - AI Infrastructure
Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 283,500 - 346,500
Equity (RSUs)
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000