Backend Software Engineer (ML Infra)

Rockstar

San Francisco (CA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A dynamic digital product studio is seeking a Backend Software Engineer (ML Infrastructure) to design and build core systems for training and deploying ML models. This early-career role involves collaborating with ML engineers and focuses on distributed training pipelines and cloud-native infrastructure. The ideal candidate has backend engineering experience, strong foundations in distributed systems, and is comfortable working in Python or Go. The position offers an exciting opportunity to work on real ML infrastructure in a fast-paced environment in San Francisco.

Qualifications

  • 1–3 years of backend engineering experience in production systems.
  • Strong fundamentals in distributed systems, networking, and backend architecture.
  • Experience building scalable systems under load.
  • Comfortable working in Python and/or Go.
  • Excited to work on-site in San Francisco.

Responsibilities

  • Design and implement backend systems for large-scale ML workloads.
  • Build efficient, fault-tolerant, and observable training pipelines.
  • Develop tools for ML engineers to train and deploy models.
  • Optimize systems for performance and cost efficiency.
  • Implement monitoring and observability for production services.

Skills

Distributed systems
Backend architecture
Python
Go
Kubernetes
Docker

Tools

Ray
vLLM
SGLang

Job description

Rockstar is recruiting for a mobile-first digital product studio that turns ideas into extraordinary experiences. They are a team of dynamic and savvy professionals who know how to create killer digital products. Our lean structure and remote team mean we can move fast while still delivering top-notch technology and design.

Our client is building the AI backbone for the next generation of intelligent products. They help fast-growing AI startups design, fine-tune, evaluate, deploy, and maintain specialized models across text, vision, and embeddings.

Think of them as “AWS for AI models”—not data or raw compute, but a full-stack backend for fine-tuning, reinforcement learning, inference, and long-term model maintenance.

Their customers are Series A–C AI companies building enterprise-grade products. Their promise is simple: they make your AI system better.

They are hiring a Backend Software Engineer (ML Infrastructure) to help design, build, and scale the core systems that power large-scale model training and deployment.

The candidate will work on distributed training pipelines, cloud-native infrastructure, and internal developer platforms that support fine-tuning, reinforcement learning, and inference at scale. This role sits at the intersection of backend engineering and ML systems—the candidate will collaborate closely with ML engineers while owning production-grade infrastructure.

This is an ideal role for an early-career engineer who wants to work on real distributed systems, GPU workloads, and modern ML infrastructure—not dashboards or CRUD apps.

What You’ll Do
Build & Scale Core Infrastructure
  • Design and implement backend systems that support large-scale ML workloads, including fine-tuning and reinforcement learning.
  • Build distributed training and inference pipelines that are efficient, fault-tolerant, and observable.
  • Develop internal developer tools and platforms that make it easier for ML engineers to train, evaluate, and deploy models.
Cloud & Systems Engineering
  • Work on cloud-native systems using containers and orchestration (e.g., Kubernetes).
  • Optimize systems for performance, reliability, and cost efficiency, especially for GPU-heavy workloads.
  • Implement monitoring, logging, and observability for long-running training jobs and production services.
Collaborate with ML Engineers
  • Partner closely with ML engineers to support evolving model architectures, training workflows, and evaluation needs.
  • Translate ML requirements into scalable backend and infrastructure solutions.
Who You Are
Required
  • 1–3 years of backend engineering experience, ideally working on production systems.
  • Strong fundamentals in distributed systems, networking, and backend architecture.
  • Experience building systems that scale under real load.
  • Comfortable working in Python and/or Go (or similar backend languages).
  • Excited to work on-site in San Francisco with a fast-moving early-stage team.
Strongly Preferred
  • Experience with or exposure to ML infrastructure or ML platforms.
  • Familiarity with GPU workloads, training pipelines, or inference systems.
  • Experience with containerization and orchestration (Docker, Kubernetes).
  • Contributions to or deep familiarity with ML infrastructure libraries such as:
  • Ray
  • vLLM
  • SGLang
  • or similar distributed ML systems
Bonus
  • Computer science background from a top-tier program or equivalent demonstrated excellence.
  • Open-source contributions, research projects, or side projects in systems or ML infrastructure.
  • A track record of high ownership and technical curiosity.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Member of Technical Staff
Member of Technical Staff

kadence • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

Hybrid
USD 233,000 - 275,000
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)

Match Group • West Hollywood (CA)

On-site
USD 190,000 - 246,000
Backend Engineer
Backend Engineer

Space Executive • Berkeley (CA)

Remote
USD 125,000 - 225,000
Medical, dental, vision
401(k)
Unlimited PTO
+2
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Senior ML Engineer
Senior ML Engineer

Next Ventures • New York (NY)

On-site
USD 130,000 - 160,000
Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

Rebar • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive medical, dental, and vision coverage
Free lunches and dinners
Senior Software Engineer - Infrastructure
Senior Software Engineer - Infrastructure

InCommon • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Echo • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including stock options
Comprehensive benefits package
401(k) program with matching contributions