GenAI ML Systems Engineer: Scalable Training & Inference

Meta

Menlo Park (CA)

On-site

USD 180,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Meta in Menlo Park seeks a Systems ML Engineer to design scalable ML training and inference systems, build data pipelines, and optimize serving latency. You will collaborate with researchers and product engineers to translate model requirements into high-throughput infrastructure.

The role involves writing production-grade Python and C++/Java, profiling performance, building tests, and participating in on-call rotations to ensure reliability across billions of ML-powered experiences.

Qualifications

  • Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
  • 2+ years of experience in software engineering with a focus on machine learning systems, distributed systems, or high-performance data infrastructure
  • Experience writing production-quality code in Python and at least one compiled language such as C++ or Java, including performance-sensitive systems code
  • Experience designing and implementing components of ML training or inference pipelines, including data preprocessing, model execution, or serving systems
  • Experience with distributed computing concepts such as data parallelism, model parallelism, or parameter server architectures as applied to ML workloads
  • Experience writing automated tests, building logging and alerting, and participating in production incident response for large-scale systems
  • Experience optimizing ML workloads for throughput and latency, including profiling GPU or CPU utilization, memory bandwidth, and communication overhead in distributed training or inference settings
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Experience with ML framework internals such as PyTorch or JAX, including custom operator development, execution graph optimization, or compiler integration
  • Experience with stream or batch data processing systems such as Apache Spark, Flink, or Kafka in the context of ML feature pipelines
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • 1+ years production experience in GenAI post-training, RLHF, RLVR, comms/collectives, parallelism, accuracy evaluation, and performance profiling
  • Familiarity with model quantization, mixed-precision training, or other techniques for reducing compute and memory costs in production ML systems

Responsibilities

  • Design and implement scalable systems for distributed ML training and inference, including data ingestion pipelines, feature processing, and model serving infrastructure
  • Develop and optimize ML platform components such as training orchestration, gradient communication, and checkpoint management across large-scale distributed environments
  • Profile and diagnose performance bottlenecks across the ML stack, including data loading, preprocessing, forward and backward passes, and serving latency
  • Write automated tests covering expected behaviors, failure modes, and error paths for ML systems components, and build monitoring and alerting for production anomalies
  • Collaborate with ML researchers and product engineers to translate model requirements into reliable, high-throughput system designs
  • Own technical design for features and components within ML infrastructure, evaluating trade-offs between throughput, latency, cost, and engineering maintainability
  • Participate in staged rollouts of ML system changes using feature flagging and A/B testing frameworks, monitoring key metrics and responding to regressions
  • Contribute to code quality through code reviews, clear technical documentation, and consolidation of duplicative implementations across the ML systems codebase
  • Support on-call rotations for owned ML infrastructure, investigating production incidents and contributing detailed retrospectives to prevent recurrence

Skills

Python
C++
Java
Distributed systems
ML pipelines
Testing & monitoring
Performance optimization

Education

Bachelor's degree

Tools

PyTorch
JAX
Apache Spark
Flink
Kafka
CUDA

Job description

Meta in Menlo Park seeks a Systems ML Engineer to design scalable ML training and inference systems, build data pipelines, and optimize serving latency. You will collaborate with researchers and product engineers to translate model requirements into high-throughput infrastructure.

The role involves writing production-grade Python and C++/Java, profiling performance, building tests, and participating in on-call rotations to ensure reliability across billions of ML-powered experiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Software Engineer, GenAI Frameworks
Software Engineer, GenAI Frameworks

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
ML Systems Engineer — On-Site in Palo Alto, High-Impact
ML Systems Engineer — On-Site in Palo Alto, High-Impact

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
ML Software Engineer: Build Scalable AI Systems
ML Software Engineer: Build Scalable AI Systems

Meta • San Francisco (CA)

On-site
USD 183,000 - 257,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Lead AI Systems Engineer - Production ML & Platforms
Lead AI Systems Engineer - Production ML & Platforms

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Enterprise ML Systems Research Engineer - GenAI & Agent
Enterprise ML Systems Research Engineer - GenAI & Agent

Scale AI • San Francisco (CA)

On-site
USD 265,000 - 331,000
Health, dental and vision coverage
Equity compensation
Learning and development stipend
+2