ML Systems Engineer - Large-Scale, Low-Latency Infra

Meta Careers

Menlo Park, Northern (CA, KY)

Hybrid

USD 180,000 - 300,000

Full time

43 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Meta is hiring a Software Engineer for the Systems ML team in Menlo Park, CA. You will design and optimize large-scale ML training and inference systems, spanning the full stack from pipelines to hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate workloads and improve AI infrastructure for billions of users, ensuring reliability and low-latency performance.

Qualifications

  • Experience designing and building large-scale ML training and inference systems.
  • Strong knowledge of distributed systems, low-latency pipelines, and memory management.
  • Proficiency in C++ and Python for production ML infrastructure.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python.
  • Identify and resolve performance bottlenecks across the ML stack using profiling and benchmarking.
  • Architect trade-offs in ML system design, including memory bandwidth and throughput.
  • Partner with research and product teams to translate model requirements into scalable infrastructure.
  • Write automated tests and monitoring for production ML components.
  • Contribute to staged rollout strategies with feature flags and experiments.

Skills

C++
Python
Distributed systems
Profiling
ML infrastructure

Education

BS in CS

Tools

CUDA
Profiling tools

Job description

Meta is hiring a Software Engineer for the Systems ML team in Menlo Park, CA. You will design and optimize large-scale ML training and inference systems, spanning the full stack from pipelines to hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate workloads and improve AI infrastructure for billions of users, ensuring reliability and low-latency performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Systems ML
Software Engineer, Systems ML

Meta Careers • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
ML Systems Architect — Scalable AI & Production Pipelines
ML Systems Architect — Scalable AI & Production Pipelines

Meta Careers • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Principal Systems ML Engineer — Architecture & Strategy
Principal Systems ML Engineer — Architecture & Strategy

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Principal ML Systems Architect & Tech Leader
Principal ML Systems Architect & Tech Leader

Meta • New York (NY)

On-site
USD 219,000 - 301,000
Bonus
Equity
Benefits
ML Systems Engineer: Low-Latency RL Training Infra
ML Systems Engineer: Low-Latency RL Training Infra

Periodic Labs • Menlo Park (CA)

On-site
USD 300,000 - 400,000
ML Inference Systems Engineer – Remote, Scalable, Low-Latency
ML Inference Systems Engineer – Remote, Scalable, Low-Latency

Atlassian Corp. • Seattle (WA), Northern (KY)

Hybrid
USD 178,000 - 233,000
Health and wellbeing resources
Volunteer days
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Software Engineer: Build Scalable AI Systems
ML Software Engineer: Build Scalable AI Systems

Meta • San Francisco (CA)

On-site
USD 183,997 - 257,000