GenAI ML Systems Engineer - Distributed Training & Serving

Meta

New York (NY)

On-site

USD 122,000 - 181,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Meta is expanding its Systems ML Engineering team to design and optimize scalable ML training and inference pipelines. You will contribute to data processing, feature handling, and model serving infrastructure across large-scale distributed environments.

You will work with researchers and product engineers to translate model requirements into robust, high-throughput systems, participating in staged rollouts and ensuring code quality with reviews and documentation.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
  • 2+ years of software engineering focusing on ML systems or distributed data infrastructure.
  • Experience writing production-quality code in Python and a compiled language (C++/Java).
  • Experience designing ML training or inference pipelines (data preprocessing, model execution, serving).
  • Experience with distributed computing concepts like data/model parallelism in ML workloads.
  • Experience writing automated tests, logging, alerting, and production incident response.

Responsibilities

  • Design scalable systems for distributed ML training and inference, including data ingestion pipelines and model serving infrastructure.
  • Develop and optimize ML platform components like training orchestration and gradient communication.
  • Profile performance and diagnose bottlenecks across the ML stack; optimize latency and throughput.
  • Write automated tests and build monitoring/alerting for production anomalies.
  • Collaborate with ML researchers and engineers to translate model requirements into reliable system designs.
  • Own technical design for ML infrastructure features, balancing throughput, latency, and cost.

Skills

Python
C++
Java
Distributed systems
ML systems
Automated tests

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

PyTorch
JAX
Apache Spark
Flink
Kafka

Job description

Meta is expanding its Systems ML Engineering team to design and optimize scalable ML training and inference pipelines. You will contribute to data processing, feature handling, and model serving infrastructure across large-scale distributed environments.

You will work with researchers and product engineers to translate model requirements into robust, high-throughput systems, participating in staged rollouts and ensuring code quality with reviews and documentation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Software Engineer, GenAI Frameworks
Software Engineer, GenAI Frameworks

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Lead AI Systems Engineer - Production ML & Platforms
Lead AI Systems Engineer - Production ML & Platforms

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
GenAI Systems Architect | Scalable ML Infra & Prototyping
GenAI Systems Architect | Scalable ML Infra & Prototyping

Socket.dev • Washington

On-site
USD 237,000 - 309,000
Health and wellbeing resources
Volunteer days
Perks & benefits program
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000