Senior Systems ML Engineer - Scalable AI Infra

Meta

Raleigh (NC)

On-site

USD 154,000 - 217,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize machine learning infrastructure at massive scale. You will work across model training, inference pipelines, and hardware-aware optimizations, collaborating with researchers and product teams to accelerate ML workloads.

Responsibilities include building high-performance ML infrastructure in C++/Python, profiling for bottlenecks, and guiding ML system design trade-offs.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 6+ years of experience in software engineering with a focus on machine learning systems or AI infrastructure.
  • Experience developing and optimizing ML training or inference pipelines using PyTorch, TensorFlow, or equivalent.
  • Experience with distributed computing architectures and large-scale systems design for ML workloads.
  • Experience programming in C++ and Python for performance-critical systems.
  • Experience using profiling and performance analysis tools to identify and resolve bottlenecks in ML or compute-intensive systems.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency execution.
  • Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools.
  • Architect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughput.
  • Partner with research and product teams to translate ML model requirements into efficient infrastructure solutions.
  • Define and track system-level metrics and service level objectives to maintain production reliability of ML serving systems.
  • Lead technical design reviews and contribute to engineering standards for ML systems across the organization.
  • Mentor other engineers on ML infrastructure best practices, debugging methodologies, and performance optimization techniques.
  • Drive adoption of AI-augmented development workflows to expand engineering productivity and broaden the scope of deliverables.
  • Contribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes.

Skills

C++
Python
Profiling tools
Distributed computing

Education

Bachelor's degree in CS/Engineering

Tools

PyTorch
TensorFlow
CUDA

Job description

Meta is seeking a Software Engineer for the Systems ML Engineering team to design and optimize machine learning infrastructure at massive scale. You will work across model training, inference pipelines, and hardware-aware optimizations, collaborating with researchers and product teams to accelerate ML workloads.

Responsibilities include building high-performance ML infrastructure in C++/Python, profiling for bottlenecks, and guiding ML system design trade-offs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems ML Engineer: Scalable AI Infrastructure
Senior Systems ML Engineer: Scalable AI Infrastructure

Meta • Providence (RI)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Bismarck (ND)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer - Scalable AI Infra (Equity)
Senior Systems ML Engineer - Scalable AI Infra (Equity)

Meta • Charleston (WV)

On-site
USD 154,000 - 217,000
Systems ML Engineer: Scalable AI Infrastructure
Systems ML Engineer: Scalable AI Infrastructure

Meta • Montgomery (AL)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems ML Engineer - Scalable AI Infrastructure
Systems ML Engineer - Scalable AI Infrastructure

Meta • Pierre (SD)

On-site
USD 154,000 - 217,000
Lead Systems ML Engineer – High-Performance AI Infra
Lead Systems ML Engineer – High-Performance AI Infra

Meta • Sacramento (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000