Staff ML Systems Engineer, Distributed Systems

Medium

Irvine (CA)

On-site

USD 170,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits
Equity participation
Opportunity for advancements in AI and robotics

Job summary

Medium is looking for a Senior / Staff ML Systems Engineer in Irvine, California, to design and build distributed infrastructures that power large-scale machine learning workflows. The ideal candidate has over 5 years of experience in distributed systems and strong Python proficiency.

You will work closely with ML and data engineering teams to create efficient production-grade workflows and establish best practices for machine learning infrastructure. The role offers a competitive salary range of $170,000 - $200,000 annually, along with comprehensive benefits.

Qualifications

  • 5+ years of experience in distributed systems or large-scale data processing.
  • Strong Python skills with performance optimization experience.
  • Expertise in system design for scalability and reliability.

Responsibilities

  • Design scalable distributed machine learning pipelines.
  • Architect distributed execution systems for workload scheduling.
  • Develop abstractions and frameworks to simplify pipeline development.

Skills

Distributed systems
Python programming
Data pipeline design
Machine learning platforms

Tools

Ray
Spark
Kubernetes
PyTorch

Job description

FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk‑aware, reliable, field‑ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data‑driven approaches or pure transformer‑only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field.

We are seeking a Senior / Staff ML Systems Engineer to architect and build the distributed infrastructure that powers large‑scale machine learning workflows across the organization.

This role sits at the intersection of machine learning, distributed systems, and platform engineering. You will be responsible for designing scalable systems that support data processing, model training, evaluation, and post‑processing pipelines while enabling ML teams to efficiently develop, operate, and scale production‑grade workflows.

You will play a critical role in defining the architectural patterns, tooling, and infrastructure that underpin our machine learning platform.

What You’ll Get To Do
  • Design and build scalable distributed machine learning pipelines across data processing, model training, evaluation, and post‑processing workflows.
  • Architect distributed execution systems, including parallelization strategies, workload scheduling, resource allocation, and fault tolerance mechanisms.
  • Develop reusable abstractions, frameworks, and libraries that simplify distributed pipeline development.
  • Optimize performance across distributed CPU and GPU environments, improving throughput, utilization, and reliability.
  • Design systems that effectively manage data partitioning, memory utilization, serialization overhead, and compute efficiency.
  • Partner closely with ML engineers, data engineers, and infrastructure teams to productionize research workflows and enable large‑scale model development.
  • Establish best practices and engineering standards for distributed machine learning infrastructure.
  • Evaluate and guide decisions around distributed computing frameworks, infrastructure technologies, and system design trade‑offs.
  • Improve observability, debugging, monitoring, and operational tooling for distributed systems at scale.
What You Have
  • 5+ years of experience building distributed systems, backend infrastructure, machine learning platforms, or large‑scale data processing systems.
  • Strong Python programming skills, including experience with concurrency, performance optimization, and systems development.
  • Experience with distributed computing frameworks such as Ray, Spark, Dask, Flink, or similar technologies.
  • Experience designing and scaling data pipelines or machine learning workflows.
  • Strong system design skills with demonstrated expertise in scalability, reliability, and performance optimization.
  • Experience diagnosing and resolving bottlenecks in distributed environments.
  • Ability to work cross‑functionally and drive technical decisions across multiple teams.
The Extras That Set You Apart
  • Experience building infrastructure for machine learning training and inference systems.
  • Familiarity with modern ML frameworks such as PyTorch or TensorFlow.
  • Experience with multi‑node or multi‑GPU training architectures, including DDP, FSDP, DeepSpeed, or similar technologies.
  • Experience operating Kubernetes‑based infrastructure and large‑scale cloud systems.
  • Deep understanding of distributed systems concepts including data locality, serialization costs, scheduling, and resource management.
  • Experience with distributed debugging, observability, and workflow orchestration platforms.
  • Proven ability to establish technical direction and influence architecture across organizations.

$170,000 - $200,000 a year

Our salary range is highly competitive with the market, but we take into consideration an individual's background and experience in determining final salary. Base pay offered may vary depending on geographic location, job‑related knowledge, skills, and experience.

In addition to competitive compensation, FieldAI offers comprehensive benefits, equity participation, and the opportunity to contribute to cutting‑edge advancements in AI and robotics.

Our salary range is generous and we consider each individual’s background and experience when determining final compensation. Base pay may vary based on role scope, job‑related knowledge, skills, experience, and the Irvine, California market.

We value diverse perspectives and are committed to fostering an inclusive workplace. We evaluate candidates and employees based on merit, qualifications, and performance, and we do not discriminate on the basis of race, color, gender, national origin, ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, or any other legally protected statu

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer, Distributed Systems
Staff ML Systems Engineer, Distributed Systems

Medium • Seattle (WA)

On-site
USD 170,000 - 200,000
Comprehensive benefits
Equity participation
Opportunity to work in AI and robotics advancements
Staff ML Systems Engineer, Distributed Systems
Staff ML Systems Engineer, Distributed Systems

FieldAI • Seattle (WA)

On-site
USD 110,000 - 150,000
Comprehensive benefits
Equity participation
Opportunity for cutting-edge advancements
Staff ML Systems Engineer, Distributed Systems
Staff ML Systems Engineer, Distributed Systems

Field AI • Seattle (WA)

On-site
USD 110,000 - 160,000
Senior Machine Learning Platform Engineer
Senior Machine Learning Platform Engineer

FieldAI • Irvine (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Collaboration with experts from top organizations
Inclusive workplace culture
Senior Machine Learning Engineer
Senior Machine Learning Engineer

AI Chopping Block • Irvine (CA)

On-site
USD 180,000 - 215,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Field AI • Irvine (CA)

On-site
USD 90,000 - 130,000
Software Engineer, Developer Infrastructure
Software Engineer, Developer Infrastructure

Field AI • Irvine (CA)

On-site
USD 90,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

FieldAI • Seattle (WA)

On-site
USD 180,000 - 215,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Medium • Irvine (CA)

On-site
USD 180,000 - 215,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Medium • Seattle (WA)

On-site
USD 180,000 - 215,000