Senior Systems ML Engineer - High-Performance AI Infra

Meta

Jefferson City (MO)

On-site

USD 154,000 - 217,000

Full time

26 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing ML infrastructure at massive scale. You will design and develop high-performance ML systems across the full stack, from training pipelines to hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads and improve AI infrastructure serving billions of users.

Qualifications

  • 6+ years in software engineering focusing on ML systems or high-performance computing.
  • Experience building ML training or inference pipelines with PyTorch or TensorFlow.
  • Experience with distributed architectures for ML workloads.
  • Proficiency in C++ and Python for performance-critical systems.
  • Experience using profiling tools to identify bottlenecks.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python.
  • Identify performance bottlenecks using profiling and benchmarking tools.
  • Architect trade-offs in ML system design including memory bandwidth, compute utilization, and I/O throughput.
  • Collaborate with research and product teams to translate model requirements into infrastructure solutions.
  • Define system-level metrics and SLIs to maintain production reliability.

Skills

ML systems
C++
Python
Distributed systems
Profiling tools
PyTorch
TensorFlow

Tools

CUDA
ROCm
MLIR
LLVM
TVM
XLA
IREE

Job description

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing ML infrastructure at massive scale. You will design and develop high-performance ML systems across the full stack, from training pipelines to hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads and improve AI infrastructure serving billions of users.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer: High-Performance AI Infra
Senior Systems ML Engineer: High-Performance AI Infra

Meta • Tallahassee (FL)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer - High-Performance Infra
Senior Systems ML Engineer - High-Performance Infra

Meta • Baton Rouge (LA)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Indianapolis (IN)

On-site
USD 154,000 - 217,000
Systems ML Engineer - High-Scale AI Infra (Equity)
Systems ML Engineer - High-Scale AI Infra (Equity)

Meta • Concord (NH)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer - Scalable AI Infra (Equity)
Senior Systems ML Engineer - Scalable AI Infra (Equity)

Meta • Charleston (WV)

On-site
USD 154,000 - 217,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Systems ML Engineer: AI Infrastructure & Hardware Acceleration
Systems ML Engineer: AI Infrastructure & Hardware Acceleration

Meta • Oklahoma City (OK)

On-site
USD 184,000 - 257,000