ML Infrastructure Engineer - Scale GPU ML Platforms

SpaceXAI

United States

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical, vision, dental
401(k) retirement plan
Disability insurance

Job summary

SpaceXAI is hiring an ML Infrastructure Engineer to design and scale a high-performance ML platform powering recommendations. You will own GPU compute infrastructure, training frameworks, and experimentation tools to accelerate ML research and production throughput.

You will collaborate with ML teams to productionize models, ensure reliability, and optimize the stack from data ingestion to inference. The role includes mentoring junior engineers and shaping the future of our infrastructure.

Qualifications

  • Bachelors/Masters/PhD in CS or quantitative field or equivalent experience (BSc/BEng not required).
  • 2+ years in high-traffic or large-scale production environments.
  • 2+ years with ML platforms, training infra, or modeling collaboration.
  • Strong Python, plus C or Rust is a plus.

Responsibilities

  • Design, build, and scale GPU compute infrastructure and training frameworks.
  • Develop data pipelines for large-scale data, training, and inference.
  • Collaborate with ML teams to productionize models across the stack.
  • Ensure scalability, reliability, and efficiency of ML systems.
  • Work across the full stack to solve problems independently.
  • Mentor junior engineers and contribute to team growth.

Skills

Python programming
Distributed systems
ML platforms
GPU infrastructure
Strong communication

Education

Bachelor/Master/PhD in CS or related field

Tools

CUDA toolkits
NVIDIA drivers
Slurm
Puppet/Ansible

Job description

SpaceXAI is hiring an ML Infrastructure Engineer to design and scale a high-performance ML platform powering recommendations. You will own GPU compute infrastructure, training frameworks, and experimentation tools to accelerate ML research and production throughput.

You will collaborate with ML teams to productionize models, ensure reliability, and optimize the stack from data ingestion to inference. The role includes mentoring junior engineers and shaping the future of our infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer — Scale GPU ML Platform & Equity
ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
GPU-Powered ML Platform Engineer
GPU-Powered ML Platform Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+5
ML Infrastructure Engineer - Scalable GPU Platform + Equity
ML Infrastructure Engineer - Scalable GPU Platform + Equity

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and/or
Dental coverage
+4
ML Platform Engineer: Build Scalable AI Infrastructure
ML Platform Engineer: Build Scalable AI Infrastructure

Neura Market • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Staff ML Engineer – Large-Scale AI & Data Systems
Staff ML Engineer – Large-Scale AI & Data Systems

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
401(k) retirement plan
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and/or
Dental coverage
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

SpaceXAI • United States

On-site
USD 180,000 - 440,000
Equity
Medical, vision, dental
401(k) retirement plan
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1