Senior Systems ML Engineer — Scalable AI Infra

Meta

Nashville (TN)

On-site

USD 154,000 - 217,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize ML infrastructure at massive scale. You will work across the full stack from model training to inference pipelines, optimizing hardware-aware performance and reliability.

You will collaborate with researchers and product teams to accelerate ML workloads, define metrics, and lead design reviews while mentoring peers in ML infrastructure best practices.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 6+ years of software engineering focusing on ML systems, AI infrastructure, or high-performance computing.
  • Experience with ML training/inference pipelines using PyTorch, TensorFlow, or equivalent.
  • Experience with distributed computing architectures for ML workloads.
  • Experience programming in C++ and Python for performance-critical systems.
  • Experience using profiling and performance analysis tools to optimize bottlenecks.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python for reliability and low latency.
  • Identify and resolve performance bottlenecks across the ML stack using profiling and benchmarking tools.
  • Architect and evaluate trade-offs in ML system design considering memory bandwidth, compute, and I/O throughput.
  • Collaborate with research and product teams to translate ML model requirements into efficient infrastructure solutions.

Skills

ML systems design
C++
Python
Performance profiling
Distributed computing
System metrics
Leadership & mentoring
Experimentation workflows
Feature flags
CUDA

Education

Bachelor's degree in CS/Engineering

Tools

CUDA
ROCm
MLIR/LLVM

Job description

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize ML infrastructure at massive scale. You will work across the full stack from model training to inference pipelines, optimizing hardware-aware performance and reliability.

You will collaborate with researchers and product teams to accelerate ML workloads, define metrics, and lead design reviews while mentoring peers in ML infrastructure best practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer: Scalable AI Infrastructure
Senior Systems ML Engineer: Scalable AI Infrastructure

Meta • Providence (RI)

On-site
USD 154,000 - 217,000
Lead Systems ML Engineer — Scalable AI Infrastructure
Lead Systems ML Engineer — Scalable AI Infrastructure

Meta • Frankfort (KY)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Bismarck (ND)

On-site
USD 154,000 - 217,000
Systems ML Engineer: Scalable AI Infrastructure
Systems ML Engineer: Scalable AI Infrastructure

Meta • Montgomery (AL)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems ML Engineer - Scalable AI Infrastructure
Systems ML Engineer - Scalable AI Infrastructure

Meta • Pierre (SD)

On-site
USD 154,000 - 217,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Senior Systems ML Engineer — High-Performance AI Infra
Senior Systems ML Engineer — High-Performance AI Infra

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Lead Systems ML Engineer – High-Performance AI Infra
Lead Systems ML Engineer – High-Performance AI Infra

Meta • Sacramento (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems ML Engineer: High-Performance AI Infra
Systems ML Engineer: High-Performance AI Infra

Meta • Saint Paul (MN)

On-site
USD 154,000 - 217,000
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000