ML Infra Engineer: Scale & Optimize Large-Scale Training

Goliath Partners Inc.

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

27 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Goliath Partners Inc. in San Francisco is hiring an ML Infrastructure Engineer to build the infrastructure for training, experimentation, and deployment of large-scale models.

You will own challenging ML systems work at the intersection of machine learning, distributed systems, and infrastructure, collaborating with researchers and engineers to push production-ready capabilities. Expect to optimize GPU utilization, design data pipelines, develop tooling, and improve experiment management as you

Qualifications

  • Strong Python and software engineering fundamentals.
  • Experience with PyTorch and modern ML training stacks.
  • Strong understanding of distributed computing and GPU-based workloads.
  • Experience building reliable systems for ML research or production.
  • Comfortable in a fast-moving environment with technical ownership.

Responsibilities

  • Build and scale infrastructure for training large machine learning models.
  • Develop distributed training systems and improve training efficiency, reliability, and throughput.
  • Build data pipelines and infrastructure supporting large-scale ML workloads.
  • Improve GPU utilization, compute orchestration, checkpointing, and experiment management.
  • Develop tooling that enables researchers and ML engineers to iterate faster.
  • Diagnose performance bottlenecks across training, data, and compute systems.
  • Help take ML systems from experimentation through production and real-world deployment.

Skills

Strong Python
Distributed computing understanding
System design fundamentals
Experience with ML pipelines
Ownership in fast-moving environ

Tools

PyTorch
CUDA
Kubernetes

Job description

Goliath Partners Inc. in San Francisco is hiring an ML Infrastructure Engineer to build the infrastructure for training, experimentation, and deployment of large-scale models.

You will own challenging ML systems work at the intersection of machine learning, distributed systems, and infrastructure, collaborating with researchers and engineers to push production-ready capabilities. Expect to optimize GPU utilization, design data pipelines, develop tooling, and improve experiment management as you

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Ops Architect: Scale ML Infra & CI/CD (Hybrid SF)
ML Ops Architect: Scale ML Infra & CI/CD (Hybrid SF)

Rise Technical Recruitment Limited • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
Equity
Healthcare
401(k)
+1
ML Infrastructure Engineer: Scalable Training and Deployment
ML Infrastructure Engineer: Scalable Training and Deployment

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Infra Engineer: Scale GPU Training & Data Pipelines
ML Infra Engineer: Scale GPU Training & Data Pipelines

Humble Robotics • United States

On-site
USD 150,000 - 230,000
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infra Architect: Build Scalable Data & Training Pipelines
ML Infra Architect: Build Scalable Data & Training Pipelines

Humble Robotics • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Goliath Partners Inc. • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

On-site
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2