AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 225,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options

Job summary

Hamilton Barnes Associates Limited is seeking a Senior ML Infrastructure Engineer to help build and scale Kubernetes-based machine learning platforms. This role focuses on workload orchestration, GPU scheduling, and ensuring system reliability, working with highly technical teams in the AI space.

The ideal candidate will have a strong ML engineering background, as well as hands-on experience with both training and inference infrastructure. The position offers a competitive salary and stock options.

Qualifications

  • Strong ML engineering background with hands‑on experience supporting both training and inference infrastructure.
  • Experience with infrastructure engineering, platform engineering, or software engineering environments.
  • Deep experience with Kubernetes, including operators and GPU scheduling.

Responsibilities

  • Build and scale internal ML infrastructure platforms focused on AI training and inference workloads.
  • Develop systems for workload orchestration and job scheduling across Kubernetes environments.
  • Collaborate with infrastructure teams to maximise GPU utilisation and operational reliability.

Skills

ML engineering background
Experience with Kubernetes
Strong programming skills in Python
Linux operating skills
CI/CD pipeline experience
Distributed systems design

Job description

Ready to take the next step in your career?

Join a rapidly growing AI cloud infrastructure provider building high-performance compute platforms for large-scale AI training and inference workloads. With expanding GPU infrastructure across Europe and the United States, the organisation enables AI teams to access scalable compute environments without traditional infrastructure limitations.

As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale Kubernetes-based machine learning platforms supporting large-scale training and inference systems. The role focuses on workload orchestration, GPU scheduling, inference optimisation, and distributed systems reliability, working alongside highly technical teams at the intersection of machine learning, cloud infrastructure, and high-performance computing.

If you would like to learn more about this opportunity, feel free to reach out and apply today!

Responsibilities
  • Build and scale internal ML infrastructure platforms focused on AI training and inference workloads
  • Develop systems for workload orchestration, job scheduling, and reliable execution across Kubernetes environments
  • Improve and maintain inference infrastructure, including model packaging, deployment, and serving optimisation
  • Collaborate with infrastructure and platform teams to maximise GPU utilisation, hardware performance, and operational reliability
  • Design scalable systems and reusable platform capabilities that improve developer experience and operational efficiency
  • Support CI/CD, GitOps, and infrastructure automation workflows across ML platform environments
  • Troubleshoot GPU performance, distributed systems behaviour, networking, and storage bottlenecks
  • Contribute to platform architecture discussions and long-term infrastructure scalability initiatives
Skills/Must Have
  • Strong ML engineering background with hands‑on experience supporting both training and inference infrastructure
  • Experience with infrastructure engineering, platform engineering, or software engineering environments
  • Strong programming skills in Python (Go experience is a plus)
  • Deep experience with Kubernetes, including operators, CRDs, workload orchestration, and GPU scheduling
  • Comfortable operating in Linux environments and debugging GPU‑related issues, including CUDA, drivers, networking, and filesystems
  • Strong systems thinking and ability to design scalable, reliable, distributed infrastructure
  • Experience with CI/CD pipelines, GitOps workflows, and infrastructure automation
Desirable Skills
  • Familiarity with orchestration and scheduling platforms such as Kueue, Flyte, Ray, or Slurm
  • Experience with PyTorch or JAX environments
  • Hands‑on experience deploying inference workloads using vLLM, SGLang, TensorRT‑LLM, or Triton
  • Knowledge of GPU networking and performance optimisation, including InfiniBand, NVLink, and NCCL
  • Experience working within HPC or large-scale distributed systems environments
Benefits
  • Stock options
Salary
  • $250,000 base salary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solution Architect - AI Infrastructure
Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 283,500 - 346,500
Equity (RSUs)
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Inference Platform Engineer - Data Center
Senior Inference Platform Engineer - Data Center

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
Senior Software Engineer (Cloud Platform) - AI Infrastructure
Senior Software Engineer (Cloud Platform) - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 190,000 - 225,000
Flexible PTO
Full Healthcare for Primary
401k
AI Infrastructure / ML Infrastructure Engineer
AI Infrastructure / ML Infrastructure Engineer

DeWinter Group • Campbell (CA)

On-site
Product Manager (AI Infrastructure) - Hosting
Product Manager (AI Infrastructure) - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock Options
Company Bonus
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000