Senior Systems ML Engineer — Scalable AI Infrastructure

Meta

Bismarck (ND)

On-site

USD 154,000 - 217,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize the machine learning infrastructure powering Meta’s products at scale. You will build high-performance ML training and inference pipelines, work across the full stack, and collaborate with researchers and product teams to accelerate workloads.

Responsibilities include profiling, optimizing, and validating ML systems, leading design reviews, and mentoring engineers while driving AI-augmented development

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 6+ years of experience in software engineering with a focus on machine learning systems, AI infrastructure, or high-performance computing.
  • Experience developing and optimizing ML training or inference pipelines using frameworks such as PyTorch, TensorFlow, or equivalent.
  • Experience with distributed computing architectures and large-scale systems design for ML workloads.
  • Experience programming in C++ and Python for performance-critical systems.
  • Experience using profiling and performance analysis tools to identify and resolve bottlenecks in ML or compute-intensive systems.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency execution.
  • Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools.
  • Architect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughput.
  • Partner with research and product teams to translate ML model requirements into efficient infrastructure solutions.
  • Define and track system-level metrics and service level objectives to maintain production reliability of ML serving systems.
  • Lead technical design reviews and contribute to engineering standards for ML systems across the organization.
  • Mentor other engineers on ML infrastructure best practices, debugging methodologies, and performance optimization techniques.
  • Drive adoption of AI-augmented development workflows to expand engineering productivity and broaden the scope of deliverables.
  • Contribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes.

Skills

C++
Python
Distributed systems
Profiling & benchmarking
Performance optimization
Mentoring engineers
ML infrastructure concepts

Education

Bachelor's degree in CS / related field

Tools

CUDA
ROCm
MLIR
LLVM
TVM
XLA
IREE

Job description

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize the machine learning infrastructure powering Meta’s products at scale. You will build high-performance ML training and inference pipelines, work across the full stack, and collaborate with researchers and product teams to accelerate workloads.

Responsibilities include profiling, optimizing, and validating ML systems, leading design reviews, and mentoring engineers while driving AI-augmented development

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer: Scalable AI Infrastructure
Senior Systems ML Engineer: Scalable AI Infrastructure

Meta • Providence (RI)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infra
Senior Systems ML Engineer — Scalable AI Infra

Meta • Nashville (TN)

On-site
USD 154,000 - 217,000
Lead Systems ML Engineer — Scalable AI Infrastructure
Lead Systems ML Engineer — Scalable AI Infrastructure

Meta • Frankfort (KY)

On-site
USD 154,000 - 217,000
Systems ML Engineer - Scalable AI Infrastructure
Systems ML Engineer - Scalable AI Infrastructure

Meta • Pierre (SD)

On-site
USD 154,000 - 217,000
Systems ML Engineer: Scalable AI Infrastructure
Systems ML Engineer: Scalable AI Infrastructure

Meta • Montgomery (AL)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Senior Systems ML Engineer — High-Performance AI Infra
Senior Systems ML Engineer — High-Performance AI Infra

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Lead Systems ML Engineer – High-Performance AI Infra
Lead Systems ML Engineer – High-Performance AI Infra

Meta • Sacramento (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
Systems ML Engineer: High-Performance AI Infra
Systems ML Engineer: High-Performance AI Infra

Meta • Saint Paul (MN)

On-site
USD 154,000 - 217,000