Machine Learning Engineer

Goliath Partners Inc.

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

35 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Goliath Partners Inc. in San Francisco is hiring an ML Infrastructure Engineer to build the infrastructure for training, experimentation, and deployment of large-scale models.

You will own challenging ML systems work at the intersection of machine learning, distributed systems, and infrastructure, collaborating with researchers and engineers to push production-ready capabilities. Expect to optimize GPU utilization, design data pipelines, develop tooling, and improve experiment management as you

Qualifications

  • Strong Python and software engineering fundamentals.
  • Experience with PyTorch and modern ML training stacks.
  • Strong understanding of distributed computing and GPU-based workloads.
  • Experience building reliable systems for ML research or production.
  • Comfortable in a fast-moving environment with technical ownership.

Responsibilities

  • Build and scale infrastructure for training large machine learning models.
  • Develop distributed training systems and improve training efficiency, reliability, and throughput.
  • Build data pipelines and infrastructure supporting large-scale ML workloads.
  • Improve GPU utilization, compute orchestration, checkpointing, and experiment management.
  • Develop tooling that enables researchers and ML engineers to iterate faster.
  • Diagnose performance bottlenecks across training, data, and compute systems.
  • Help take ML systems from experimentation through production and real-world deployment.

Skills

Strong Python
Distributed computing understanding
System design fundamentals
Experience with ML pipelines
Ownership in fast-moving environ

Tools

PyTorch
CUDA
Kubernetes

Job description

I’m working with a rapidly growing, well-funded AI company building some of the most technically ambitious real-world AI systems today.

They’re looking for an ML Infrastructure Engineer to join a highly technical team responsible for building the infrastructure that powers large-scale model training, experimentation, and deployment.

This is a hands-on engineering role for someone who enjoys solving difficult systems problems at the intersection of machine learning, distributed systems, and infrastructure.

What you’ll work on:
  • Build and scale infrastructure for training large machine learning models
  • Develop distributed training systems and improve training efficiency, reliability, and throughput
  • Build data pipelines and infrastructure supporting large-scale ML workloads
  • Improve GPU utilization, compute orchestration, checkpointing, and experiment management
  • Develop tooling that enables researchers and ML engineers to iterate faster
  • Diagnose performance bottlenecks across training, data, and compute systems
  • Help take ML systems from experimentation through production and real-world deployment
What they’re looking for:
  • Strong Python and software engineering fundamentals
  • Experience working with PyTorch and modern ML training stacks
  • Strong understanding of distributed computing and GPU-based workloads
  • Experience building reliable systems that support ML research or production ML
  • Comfortable working in a fast-moving environment with significant technical ownership
Especially interesting backgrounds include:
  • Distributed training and large-scale model training
  • GPU infrastructure and compute orchestration
  • Large-scale data infrastructure and pipelines
  • Performance optimization for ML workloads
  • Infrastructure supporting robotics, embodied AI, computer vision, or other real-world ML systems
  • Experience taking ML systems beyond research and into production or physical-world environments

This is a great opportunity for an engineer who wants to work on hard ML systems problems at scale while being much closer to the models and real-world applications than you would be on a traditional infrastructure team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman Ltd • New York (NY)

On-site
USD 120,000 - 170,000
Health insurance
Competitive equity
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

On-site
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

Mach9 Robotics Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Strativ Group • Palo Alto (CA)

On-site
USD 450,000 - 550,000
Founding equity
Direct exposure to founders
Competitive compensation
Machine Learning Engineer
Machine Learning Engineer

Evlo AI • Miami (FL)

On-site
USD 110,000 - 170,000
None
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
Machine Learning Engineer
Machine Learning Engineer

Prodigy Resources • Denver (CO)

On-site
USD 160,000 - 210,000