ML Infrastructure Engineer

Objective Partners

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
Commuter reimbursement

Job summary

Objective Partners in San Francisco is seeking a Machine Learning Infrastructure Engineer to architect and scale compute, training, and inference pipelines powering next-gen AI models.

You will own distributed systems across large GPU clusters, collaborate with core platform teams, and drive low-latency serving with high reliability in a fast-paced startup environment.

The role is full-time and on-site, with comprehensive health benefits, flexible PTO, catered lunches, and commuter reimbursement.

Qualifications

  • 3+ years of production experience in software development, backend systems, or ML platforms.
  • Advanced proficiency in Python and strong background in low-level systems design.
  • Hands-on experience with container orchestration (Kubernetes) and distributed compute frameworks.
  • Ability to own projects from planning to deployment.
  • Comfort in a fast-paced, high-rigor environment with rapid iteration.

Responsibilities

  • Architect and sustain high-throughput pipelines for rapid experimentation and ultra-low latency serving.
  • Establish benchmarking suites to identify bottlenecks and optimize resources.
  • Evaluate model optimization developments and integrate them into production.
  • Build multi-node environments that support parallelized training.
  • Maximize hardware efficiency across large GPU clusters.
  • Develop tooling and telemetry to streamline research-to-production transitions.

Skills

Python
Distributed systems
Backend development
First principles thinking
Fast-paced environments

Tools

Kubernetes
Docker
GPU clusters
Telemetry tooling

Job description

Job Title: Machine Learning Infrastructure Engineer

Location: San Francisco, CA Metro Area (100% On-Site)

About the Opportunity

An ultra-high-growth artificial intelligence platform company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks powering next-generation model performance. Backed by top-tier venture capital firms and serving elite technology enterprises, this team is scaling rapidly to solve complex unstructured data challenges. In this role, you will hold direct ownership over scaling distributed systems across massive hardware clusters while collaborating directly with core platform teams.

Responsibilities
  • Architect and sustain high-throughput execution and training pipelines optimized for rapid experimentation, flexibility, and ultra-low latency serving.

  • Establish comprehensive benchmarking suites across system stacks to identify performance bottlenecks and optimize resource efficiency.

  • Evaluate cutting-edge developments in model optimization and integrate state-of-the-art research into production systems.

  • Engineer resilient, highly observable multi-node environments that support seamless parallelized training tasks.

  • Maximize hardware efficiency and operational reliability while running large-scale workloads across extensive GPU clusters.

  • Build developer abstractions, internal tooling, and telemetry frameworks that streamline the transition from research prototype to production.

Requirements (Must-Have)
  • 3+ years of production experience in software development, backend systems, or machine learning platforms (exceptional early-career candidates with extraordinary technical foundations will be evaluated).

  • Advanced proficiency in Python alongside a solid background in low-level systems architectural design.

  • Hands-on experience working with container orchestration (e.g., Kubernetes) and modern distributed compute frameworks.

  • Demonstrated ability to build from first principles and take full ownership of projects from strategic planning to deployment.

  • Comfortable operating in a fast-paced, high-rigor setting with continuous iteration cycles.

Preferred Qualifications (Nice-to-Have)
  • Prior experience working within an early-stage or rapidly scaling tech startup.

  • Meaningful open-source contributions to established distributed training or inference engines.

  • Practical experience managing multi-node inference setups across hundreds or thousands of GPUs.

  • Strong passion for aligning deep technical excellence with direct business results.

Compensation & Benefits
  • Work Arrangement: Full-time, 5 days per week on-site in San Francisco, CA.

  • Health & Wellness: Fully covered medical, dental, and vision coverage, plus a $150 monthly stipend for wellness and fitness expenses.

  • Time Off & Flexibility: Flexible PTO policy and adaptable parental leave programs.

  • Perks: Daily catered lunches in the office and full commuter reimbursement.

  • Equal Opportunity Employer: All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status.

#LI-ML1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infrastructure Engineer — On-Site SF, GPU Pipelines
ML Infrastructure Engineer — On-Site SF, GPU Pipelines

Objective Partners • San Francisco (CA)

On-site
USD 180,000 - 250,000
Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
+1
Machine Learning Engineer
Machine Learning Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 130,000 - 180,000
Competitive salary guidance
Performance bonus up to 20%
Equity options
+4
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Segment (Twilio) • San Francisco (CA)

On-site
USD 175,000 - 220,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision
401(k) program
Generous PTO
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Software Engineer AI/ML Systems - USA Onsite (Santa Clara, CA)
Software Engineer AI/ML Systems - USA Onsite (Santa Clara, CA)

Dover • Santa Clara (CA), Northern (KY)

Hybrid
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Stock Purchase Program (ESPP)
+3
Senior Machine Learning Engineer, AI Infra
Senior Machine Learning Engineer, AI Infra

United States Digital Space LLC • Bellevue (CA)

On-site
USD 209,000 - 245,000
Health insurance
Equity ownership
401k matching
+2