ML Infra Engineer: Scale GPU Training & Inference

Reducto

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Free lunch
Reimbursed transportation
Generous health insurance
Health and wellness budget
Parental leave

Job summary

Reducto, a fast-growing AI company in San Francisco, is hiring a Machine Learning Infra Engineer. This role involves building and maintaining the training and inference frameworks necessary for optimal performance. Ideal candidates should possess strong Python skills, have a background in systems engineering, and experience with Kubernetes. The position offers numerous benefits, including unlimited PTO, free lunches, and health insurance coverage, all focused on fostering a productive and supportive work environment.

Qualifications

  • Strong Python skills and a background in systems engineering.
  • Experience with Kubernetes and distributed training frameworks.
  • Ability to work in fast-changing, high-growth environments.

Responsibilities

  • Build and maintain our training and inference stack.
  • Develop benchmarks for training and inference to identify bottlenecks.
  • Design systems for scaling model training across multi-node, multi-GPU environments.

Skills

Python skills
Systems engineering background
Kubernetes
Distributed training frameworks
Problem-solving

Job description

Reducto, a fast-growing AI company in San Francisco, is hiring a Machine Learning Infra Engineer. This role involves building and maintaining the training and inference frameworks necessary for optimal performance. Ideal candidates should possess strong Python skills, have a background in systems engineering, and experience with Kubernetes. The position offers numerous benefits, including unlimited PTO, free lunches, and health insurance coverage, all focused on fostering a productive and supportive work environment.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Goliath Partners Inc. • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
ML Infrastructure Engineer — On-Site SF, GPU Pipelines
ML Infrastructure Engineer — On-Site SF, GPU Pipelines

Objective Partners • San Francisco (CA)

On-site
USD 180,000 - 250,000
Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
+1
ML Infrastructure Engineer: GPU Training & Serving
ML Infrastructure Engineer: GPU Training & Serving

Character.AI • San Francisco (CA)

On-site
USD 130,000 - 207,000
Health insurance
Flexible working hours
Professional development opportunities
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8