AI Infra/HPC Engineer

Blue Signal Search

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
Equity participation
Comprehensive benefits
Ownership of critical infra

Job summary

Blue Signal is seeking an infrastructure engineer to help design, deploy, and operate large-scale GPU cloud infrastructure for AI training and inference. You will own production reliability and work with a founding team to shape architecture, tooling, and operational standards from day one.

You will collaborate across engineering and customer teams to deliver scalable, high-performance infrastructure, with equity participation and a path to technical leadership in a fast-growing company.

Qualifications

  • 3–6 years building and operating large Kubernetes/Slurm clusters.
  • Experience designing or managing GPU orchestration platforms, including custom scheduling or topology-aware placement.
  • Deep understanding of distributed object storage, NVMe storage clusters, and high bandwidth networking for AI infra.
  • Experience operating production GPU infrastructure at scale.
  • Strong systems engineering background supporting distributed AI workloads.
  • Comfortable owning production reliability, troubleshooting, and infra operations.
  • Strong Linux systems administration and infrastructure automation experience.
  • Excellent troubleshooting, communication, and collaboration skills.

Responsibilities

  • Design, deploy, and operate large-scale Kubernetes and/or Slurm environments for production AI workloads.
  • Build and enhance GPU orchestration capabilities, including scheduling optimization and topology-aware placement.
  • Architect and manage distributed storage platforms using object storage, NVMe clusters, and high performance networking.
  • Develop reliable infrastructure supporting large-scale distributed training and inference environments.
  • Partner with customers and internal teams to design infrastructure solutions meeting AI workload requirements.
  • Own production reliability, incident response, performance optimization, and observability.

Skills

Kubernetes clusters
Slurm clusters
GPU orchestration
Distributed storage
NVMe storage
High bandwidth networking
Linux administration
Automation / IaC
Troubleshooting
Collaboration

Job description

Are you passionate about building the infrastructure powering the next generation of artificial intelligence? Our client is an emerging, venture backed technology company developing advanced GPU cloud infrastructure designed for large scale AI training and inference workloads. As one of the organization's early infrastructure engineers, you will play a foundational role in architecting and operating the systems that enable cutting edge AI applications. This is a unique opportunity to help shape a rapidly growing platform, influence core technical decisions, and work alongside an experienced founding team tackling some of the most complex infrastructure challenges in AI.

This Role Offers

  • Competitive base salary, annual bonus, comprehensive benefits, and meaningful equity participation.
  • Opportunity to join an early stage, well-funded company at a pivotal stage of growth.
  • Direct ownership of mission critical AI infrastructure supporting enterprise customers.
  • Highly collaborative engineering culture with significant technical autonomy.
  • Opportunity to influence architecture, tooling, and operational standards from the ground up.
  • Exposure to some of the industry's most advanced GPU computing technologies.

Focus

  • Design, deploy, and operate large scale Kubernetes and or Slurm environments supporting production AI workloads.
  • Build and enhance GPU orchestration capabilities, including scheduling optimization, topology aware placement, and GPU lifecycle automation.
  • Architect and manage distributed storage platforms utilizing object storage, NVMe clusters, and high performance networking.
  • Develop reliable infrastructure supporting large scale distributed training and inference environments.
  • Partner directly with customers and internal engineering teams to design infrastructure solutions that meet demanding AI workload requirements.
  • Own production reliability, incident response, performance optimization, and infrastructure observability.
  • Create operational documentation, automation, and best practices that improve platform scalability and maintainability.
  • Continuously evaluate emerging infrastructure technologies that improve efficiency, performance, and operational resilience.

Required Qualifications

  • 3 to 6 years of experience building and operating large scale Kubernetes and or Slurm clusters.
  • Experience designing, implementing, or managing GPU orchestration platforms, including custom scheduling, topology aware placement, or GPU lifecycle automation.
  • Deep understanding of distributed object storage, NVMe storage clusters, storage architectures, and high bandwidth networking for AI infrastructure.
  • Experience operating production GPU infrastructure at scale.
  • Strong systems engineering background supporting distributed AI workloads.
  • Comfortable owning production reliability, troubleshooting, and infrastructure operations.
  • Strong Linux systems administration and infrastructure automation experience.
  • Excellent troubleshooting, communication, and collaboration skills.

Preferred Qualifications

  • Experience supporting large scale AI inference platforms.
  • Familiarity with GPU performance optimization and workload scheduling.
  • Experience with infrastructure as code and production automation.
  • Exposure to power aware scheduling, GPU power management, or energy optimized infrastructure is a plus.

About Blue Signal:

Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Client Partner Executive – AI Infrastructure
Client Partner Executive – AI Infrastructure

Blue Signal Search • San Diego (CA)

On-site
USD 150,000 - 200,000
Competitive base salary
Equity participation
Full medical, dental, and vision coverage
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Cash bonus
Founding engineer equity
Benefits
Infrastructure Engineer - AI, Kubernetes & Edge Systems
Infrastructure Engineer - AI, Kubernetes & Edge Systems

UMATR • Austin (TX)

On-site
USD 120,000 - 190,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus
HPC Solution Architect - AI Infrastructure
HPC Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Healthcare + Dental + Vision
401(k)
+1