AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Annual bonus
Equity participation
Comprehensive benefits
Ownership of critical infra

Job summary

Blue Signal is seeking an infrastructure engineer to help design, deploy, and operate large-scale GPU cloud infrastructure for AI training and inference. You will own production reliability and work with a founding team to shape architecture, tooling, and operational standards from day one.

You will collaborate across engineering and customer teams to deliver scalable, high-performance infrastructure, with equity participation and a path to technical leadership in a fast-growing company.

Qualifications

  • 3–6 years building and operating large Kubernetes/Slurm clusters.
  • Experience designing or managing GPU orchestration platforms, including custom scheduling or topology-aware placement.
  • Deep understanding of distributed object storage, NVMe storage clusters, and high bandwidth networking for AI infra.
  • Experience operating production GPU infrastructure at scale.
  • Strong systems engineering background supporting distributed AI workloads.
  • Comfortable owning production reliability, troubleshooting, and infra operations.
  • Strong Linux systems administration and infrastructure automation experience.
  • Excellent troubleshooting, communication, and collaboration skills.

Responsibilities

  • Design, deploy, and operate large-scale Kubernetes and/or Slurm environments for production AI workloads.
  • Build and enhance GPU orchestration capabilities, including scheduling optimization and topology-aware placement.
  • Architect and manage distributed storage platforms using object storage, NVMe clusters, and high performance networking.
  • Develop reliable infrastructure supporting large-scale distributed training and inference environments.
  • Partner with customers and internal teams to design infrastructure solutions meeting AI workload requirements.
  • Own production reliability, incident response, performance optimization, and observability.

Skills

Kubernetes clusters
Slurm clusters
GPU orchestration
Distributed storage
NVMe storage
High bandwidth networking
Linux administration
Automation / IaC
Troubleshooting
Collaboration

Job description

Blue Signal is seeking an infrastructure engineer to help design, deploy, and operate large-scale GPU cloud infrastructure for AI training and inference. You will own production reliability and work with a founding team to shape architecture, tooling, and operational standards from day one.

You will collaborate across engineering and customer teams to deliver scalable, high-performance infrastructure, with equity participation and a path to technical leadership in a fast-growing company.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
GPU Cloud Platform Lead for AI Inference
GPU Cloud Platform Lead for AI Inference

Blue Signal LLC. • United States

On-site
USD 250,000 - 350,000
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader

Vcluster • Northern (KY)

Hybrid
USD 140,000 - 165,000
Competitive Salary
Equity participation
Health, dental, vision, life Insurance
+2
AI Kernel & GPU HPC Cluster Engineer
AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior MLOps Engineer: GPU AI Infra & Production
Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Senior AI Infra Engineer - GPU & Kubernetes
Senior AI Infra Engineer - GPU & Kubernetes

SB Telecom America Corp. • Sunnyvale (CA)

On-site
USD 150,000 - 250,000
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000