Staff AI Infrastructure Engineer — Orchestration & Inference

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 213,000 - 288,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Early-stage equity
Direct access to leadership

Job summary

Hamilton Barnes Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration, scheduling, and agentic operations across bare metal, Slurm, Kubernetes, and InfiniBand in a live production environment.

You’ll contribute to inference platforms, model serving, and high-utilization microservices, while building IaC tooling to automate deployment across thousands of servers.

Qualifications

  • Strong systems software background at the intersection of distributed systems and infrastructure automation.
  • Experience building orchestration, scheduling, or cluster management software at scale (Kubernetes, Slurm, or equivalent).
  • Solid understanding of GPU infrastructure (InfiniBand, NVLink, NCCL) and related software considerations.
  • Production engineering mindset with shipped software on real infrastructure.
  • Proficiency in Go, Python, or Rust; Linux systems programming.

Responsibilities

  • Build and evolve the core AI infrastructure software stack, including orchestration, scheduling, cluster management, and the agentic operations layer for placement, healing, and recovery.
  • Contribute to the inference platform, including token gateway, microVM pooling, and model serving infra for high utilization and low latency.
  • Work across bare metal, Slurm, Kubernetes, and InfiniBand environments; write software that abstracts complexity for customers while remaining debuggable by engineers.
  • Develop IaC tooling to automate deployment at scale and track changes across 10,000+ servers.
  • Integrate with NVIDIA inference microservices, custom model pipelines, and customer workloads for simple consumption of GPU infra.
  • Contribute to the SRE tooling layer: monitoring, optimization, GPU utilization balancing, and automated remediation.

Skills

Distributed systems
Orchestration software
Kubernetes
Slurm
GPU infrastructure
Go
Python
Rust
Linux systems programming
InfiniBand
NCCL
NVLink
Go/Python/Rust

Job description

Hamilton Barnes Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration, scheduling, and agentic operations across bare metal, Slurm, Kubernetes, and InfiniBand in a live production environment.

You’ll contribute to inference platforms, model serving, and high-utilization microservices, while building IaC tooling to automate deployment across thousands of servers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Network Architect for Large-Scale GPU Clusters
Senior AI Network Architect for Large-Scale GPU Clusters

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior AI Storage Engineer — High-Performance GPU Data Fabric
Senior AI Storage Engineer — High-Performance GPU Data Fabric

Hamilton Barnes Associates Limited • United States

On-site
USD 170,000 - 230,000
Stock options
Bonus 10%
Staff Engineer, AI Cloud Orchestration
Staff Engineer, AI Cloud Orchestration

Lambda • United States

Remote
USD 180,000 - 240,000
Senior AI Infra Engineer - Hyperscale GPU Platform (Remote)
Senior AI Infra Engineer - Hyperscale GPU Platform (Remote)

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Remote AI Platform Engineer — Kubernetes & GPU, Equity
Remote AI Platform Engineer — Kubernetes & GPU, Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 250,000 - 300,000
Meaningful equity
Fully remote across North America
Full insurance coverage for you and你的依
Senior AI Systems Infra Engineer — Remote & Equity
Senior AI Systems Infra Engineer — Remote & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 250,000 - 300,000
Meaningful equity
Fully remote across North America
Full insurance coverage for you and你的依
+1
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Staff AI Research Infrastructure Engineer
Staff AI Research Infrastructure Engineer

Doist • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
AI Inference Orchestration - Distributed Systems Engineer
AI Inference Orchestration - Distributed Systems Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000