AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
Equity participation
Comprehensive benefits
Ownership of critical infra

Job summary

Blue Signal is seeking an infrastructure engineer to help design, deploy, and operate large-scale GPU cloud infrastructure for AI training and inference. You will own production reliability and work with a founding team to shape architecture, tooling, and operational standards from day one.

You will collaborate across engineering and customer teams to deliver scalable, high-performance infrastructure, with equity participation and a path to technical leadership in a fast-growing company.

Qualifications

  • 3–6 years building and operating large Kubernetes/Slurm clusters.
  • Experience designing or managing GPU orchestration platforms, including custom scheduling or topology-aware placement.
  • Deep understanding of distributed object storage, NVMe storage clusters, and high bandwidth networking for AI infra.
  • Experience operating production GPU infrastructure at scale.
  • Strong systems engineering background supporting distributed AI workloads.
  • Comfortable owning production reliability, troubleshooting, and infra operations.
  • Strong Linux systems administration and infrastructure automation experience.
  • Excellent troubleshooting, communication, and collaboration skills.

Responsibilities

  • Design, deploy, and operate large-scale Kubernetes and/or Slurm environments for production AI workloads.
  • Build and enhance GPU orchestration capabilities, including scheduling optimization and topology-aware placement.
  • Architect and manage distributed storage platforms using object storage, NVMe clusters, and high performance networking.
  • Develop reliable infrastructure supporting large-scale distributed training and inference environments.
  • Partner with customers and internal teams to design infrastructure solutions meeting AI workload requirements.
  • Own production reliability, incident response, performance optimization, and observability.

Skills

Kubernetes clusters
Slurm clusters
GPU orchestration
Distributed storage
NVMe storage
High bandwidth networking
Linux administration
Automation / IaC
Troubleshooting
Collaboration

Job description

Blue Signal is seeking an infrastructure engineer to help design, deploy, and operate large-scale GPU cloud infrastructure for AI training and inference. You will own production reliability and work with a founding team to shape architecture, tooling, and operational standards from day one.

You will collaborate across engineering and customer teams to deliver scalable, high-performance infrastructure, with equity participation and a path to technical leadership in a fast-growing company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior MLOps Engineer: GPU AI Infra & Production
Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
AI Kernel & GPU HPC Cluster Engineer
AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
AI Infra Solutions Architect — GPU & Cloud Deployments
AI Infra Solutions Architect — GPU & Cloud Deployments

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000