Senior Systems ML Engineer - Scalable AI Infra

Meta

Montpelier (VT)

On-site

USD 154,000 - 217,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing the machine learning infrastructure that powers Meta's products at massive scale.

In this role you will design and develop high-performance ML systems, spanning from model training and inference pipelines to hardware-aware optimizations, collaborating across researchers, platform engineers, and product teams.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 6+ years of software engineering experience with ML systems or high-performance computing.
  • Experience with ML training/inference pipelines (PyTorch, TensorFlow).
  • Experience with distributed computing architectures and large-scale ML workloads.
  • Experience in C++ and Python for performance-critical systems.
  • Experience with profiling/performance analysis tools.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines
  • Develop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency execution
  • Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools
  • Architect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughput
  • Partner with research and product teams to translate ML model requirements into efficient infrastructure solutions
  • Define and track system‑level metrics and service level objectives to maintain production reliability of ML serving systems
  • Lead technical design reviews and contribute to engineering standards for ML systems across the organization
  • Mentor other engineers on ML infrastructure best practices, debugging methodologies, and performance optimization techniques
  • Drive adoption of AI‑augmented development workflows to expand engineering productivity and broaden the scope of deliverables
  • Contribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes

Skills

C++
Python
ML infrastructure
Distributed systems
Performance optimization
Profiling tools
Team mentorship

Education

Bachelor's degree in CS/Engineering

Tools

PyTorch
TensorFlow

Job description

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing the machine learning infrastructure that powers Meta's products at massive scale.

In this role you will design and develop high-performance ML systems, spanning from model training and inference pipelines to hardware-aware optimizations, collaborating across researchers, platform engineers, and product teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer: Scalable AI Infra
Senior Systems ML Engineer: Scalable AI Infra

Meta • Cheyenne (WY)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer: Build High-Performance AI Infra
Senior Systems ML Engineer: Build High-Performance AI Infra

Meta • Honolulu (HI)

On-site
USD 154,000 - 217,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
ML Systems Architect — Scalable AI & Production Pipelines
ML Systems Architect — Scalable AI & Production Pipelines

Meta Careers • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Senior Machine Learning Engineer — Scalable AI Systems
Senior Machine Learning Engineer — Scalable AI Systems

Meta • Honolulu (HI)

On-site
USD 184,000 - 257,000
Principal ML Systems Architect & Tech Leader
Principal ML Systems Architect & Tech Leader

Meta • New York (NY)

On-site
USD 219,000 - 301,000
Bonus
Equity
Benefits
Senior ML Engineer - Scalable AI & Product Impact
Senior ML Engineer - Scalable AI & Product Impact

SupportFinity™ • Washington

On-site
USD 154,000 - 217,000
Equity
Bonus
Benefits