Staff ML Systems Engineer: Scalable AI Infrastructure

Meta

Sunnyvale (CA)

On-site

USD 183,997 - 257,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focusing on building and scaling the infrastructure and software systems that power large-scale machine learning workloads across Meta's production fleet.

In this role, you will architect and own critical components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling.

Qualifications

  • 8+ years of software engineering experience focusing on systems or ML infrastructure.
  • Experience designing and implementing large-scale distributed systems (training orchestration, serving, or data pipelines).
  • Experience with performance analysis and optimization of compute-intensive workloads (profiling, benchmarking, bottleneck ID).
  • Experience leading end-to-end delivery of complex technical projects across teams.
  • Proficiency in C++ or Python for production ML or infrastructure systems.

Responsibilities

  • Design scalable ML infrastructure components (training frameworks, model serving, ML platform tooling).
  • Lead architectural decisions for major ML infrastructure initiatives with trade-off analysis.
  • Identify and resolve performance bottlenecks in distributed ML systems through profiling and optimization.
  • Define SLAs for ML services, build dashboards, alerts, and runbooks to reduce MTTR.
  • Collaborate with researchers, engineers, and infra teams to translate model requirements into production systems.

Skills

Distributed systems design
Performance analysis
End-to-end project leadership
C++
Python

Job description

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focusing on building and scaling the infrastructure and software systems that power large-scale machine learning workloads across Meta's production fleet.

In this role, you will architect and own critical components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Principal Systems ML Engineer - Scalable AI Infra
Principal Systems ML Engineer - Scalable AI Infra

Meta • Sunnyvale (CA)

On-site
USD 219,000 - 301,000
Senior ML Engineer: Scalable AI for Global Impact
Senior ML Engineer: Scalable AI for Global Impact

Jobzhr • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Lead ML Software Engineer — Build Scalable AI Systems
Lead ML Software Engineer — Build Scalable AI Systems

Meta • Columbus (OH)

On-site
USD 183,997 - 257,000
Senior ML Infrastructure Architect — Scalable AI Platform
Senior ML Infrastructure Architect — Scalable AI Platform

Meta • Sunnyvale (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Lead ML Engineer – Scalable AI Systems
Lead ML Engineer – Scalable AI Systems

Meta • Bellevue (WA)

On-site
USD 183,997 - 257,000
ML Systems Infrastructure Engineer
ML Systems Infrastructure Engineer

Meta • San Francisco (CA)

On-site
USD 184,000 - 257,000
Lead AI Systems Engineer - Production ML & Platforms
Lead AI Systems Engineer - Production ML & Platforms

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Senior ML Software Engineer | Lead Scalable AI Systems
Senior ML Software Engineer | Lead Scalable AI Systems

Meta • Sunnyvale (CA)

On-site
USD 184,000 - 257,000