ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical coverage
Vision coverage
Dental coverage
401(k)
Short/Long disability insurance
Life insurance
Discounts

Job summary

SpaceXAI seeks an ML Infrastructure Engineer to design and scale a high-performance ML platform powering recommendations. You will build GPU compute infra, data pipelines, and tooling to accelerate experimentation and productionization of models.

You will collaborate with ML teams to integrate models across the stack, ensure reliability, and mentor junior engineers while tackling complex systemic challenges.

Qualifications

  • Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
  • 2+ years of industry experience with large-scale production environments, distributed systems, GPU infra, and/or deep learning applications
  • 2+ years experience with ML platforms, training infrastructure, or collaboration with modeling engineers and data scientists
  • Strong proficiency with Python and experience with C++ or Rust

Responsibilities

  • Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
  • Developing data pipelines and integrating large-scale data, training, and inference systems
  • Collaborating with ML teams to productionize models and ensure seamless integration across the stack
  • Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
  • Working across the full stack to solve complex problems independently
  • Mentoring junior engineers and contributing to the growth of the team

Skills

Python
C++/Rust
Distributed systems
GPU infrastructure
ML platforms
Linux
Slurm

Education

Bachelor/Master/PhD in CS/ML

Tools

CUDA toolkits
NVIDIA drivers
Slurm
Puppet/Ansible

Job description

SpaceXAI seeks an ML Infrastructure Engineer to design and scale a high-performance ML platform powering recommendations. You will build GPU compute infra, data pipelines, and tooling to accelerate experimentation and productionization of models.

You will collaborate with ML teams to integrate models across the stack, ensure reliability, and mentor junior engineers while tackling complex systemic challenges.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer - Scale GPU ML Platforms
ML Infrastructure Engineer - Scale GPU ML Platforms

SpaceXAI • United States

On-site
USD 180,000 - 440,000
Equity
Medical, vision, dental
401(k) retirement plan
+1
ML Infrastructure Engineer - Scalable GPU Platform + Equity
ML Infrastructure Engineer - Scalable GPU Platform + Equity

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and/or
Dental coverage
+4
GPU-Powered ML Platform Engineer
GPU-Powered ML Platform Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and/or
Dental coverage
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

SpaceXAI • United States

On-site
USD 180,000 - 440,000
Equity
Medical, vision, dental
401(k) retirement plan
+1
AI Infrastructure Engineer — Scale ML Serving & Equity
AI Infrastructure Engineer — Scale ML Serving & Equity

Scale • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 225,000
Health, dental & vision coverage
Equity-based compensation
Retirement benefits
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+5
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
ML Platform Engineer: Scale AI Infra, Deploy & Optimize
ML Platform Engineer: Scale AI Infra, Deploy & Optimize

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000