Senior Systems ML Engineer: Scalable AI Infrastructure

Meta

Providence (RI)

On-site

USD 154,000 - 217,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer for the Systems ML Engineering team to build and optimize ML infrastructure at massive scale. You will design high-performance ML systems, from training to inference pipelines, and work across the full stack for hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads, define metrics, and lead reviews while mentoring others in ML infrastructure best practices.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • 6+ years of software engineering focusing on ML systems, AI infrastructure, or HPC.
  • Experience developing and optimizing ML training or inference pipelines (PyTorch, TensorFlow or equivalent).
  • Experience with distributed computing architectures and large-scale ML systems design.
  • C++ and Python for performance-critical systems.
  • Experience using profiling and performance analysis tools to identify bottlenecks.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed frameworks and hardware-accelerated pipelines.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python.
  • Identify and resolve performance bottlenecks across the ML stack using profiling and benchmarking tools.
  • Architect and evaluate trade-offs in ML system design including memory bandwidth and I/O throughput.
  • Partner with research and product teams to translate ML model requirements into infrastructure solutions.
  • Define and track system-level metrics and SLAs to maintain production reliability of ML serving systems.
  • Lead technical design reviews and contribute to ML engineering standards.
  • Mentor engineers on ML infrastructure best practices and performance optimization.
  • Drive adoption of AI-augmented development workflows to expand engineering productivity.
  • Contribute to staged rollout strategies using feature flags and experimentation frameworks.

Skills

ML systems design
C++/Python infra
Performance profiling
System design trade-offs
Cross-functional collaboration
SLA metrics
Technical leadership
Mentorship
AI-powered workflows
Feature flags/ rollout
ML pipelines
Distributed systems
Profiling tools

Education

Bachelor's degree or equivalent

Tools

CUDA/ROCm

Job description

Meta is seeking a Software Engineer for the Systems ML Engineering team to build and optimize ML infrastructure at massive scale. You will design high-performance ML systems, from training to inference pipelines, and work across the full stack for hardware-aware optimizations.

You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads, define metrics, and lead reviews while mentoring others in ML infrastructure best practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer — Scalable AI Infra
Senior Systems ML Engineer — Scalable AI Infra

Meta • Nashville (TN)

On-site
USD 154,000 - 217,000
Lead Systems ML Engineer — Scalable AI Infrastructure
Lead Systems ML Engineer — Scalable AI Infrastructure

Meta • Frankfort (KY)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Bismarck (ND)

On-site
USD 154,000 - 217,000
Systems ML Engineer: Scalable AI Infrastructure
Systems ML Engineer: Scalable AI Infrastructure

Meta • Montgomery (AL)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems ML Engineer - Scalable AI Infrastructure
Systems ML Engineer - Scalable AI Infrastructure

Meta • Pierre (SD)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — High-Performance AI Infra
Senior Systems ML Engineer — High-Performance AI Infra

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Lead Systems ML Engineer – High-Performance AI Infra
Lead Systems ML Engineer – High-Performance AI Infra

Meta • Sacramento (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems ML Engineer: High-Performance AI Infra
Systems ML Engineer: High-Performance AI Infra

Meta • Saint Paul (MN)

On-site
USD 154,000 - 217,000
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000