Systems ML Engineer: Scalable AI Infrastructure

Meta

Montgomery (AL)

On-site

USD 154,000 - 217,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing the ML infrastructure that powers Meta's products at massive scale. You will design high-performance ML systems spanning training, inference, and hardware-aware optimizations.

You will collaborate with researchers and product teams to accelerate ML workloads, implement production-ready pipelines, and drive reliability with system-level metrics and SLIs.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 6+ years of software engineering focusing on ML systems, AI infrastructure, or high-performance computing.
  • Experience developing and optimizing ML training or inference pipelines with PyTorch or TensorFlow.
  • Experience with distributed computing architectures and large-scale ML systems design.
  • C++ and Python experience for performance-critical systems.
  • Experience using profiling and performance analysis tools to resolve bottlenecks.

Responsibilities

  • Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines.
  • Develop and maintain high-performance ML infrastructure components in C++ and Python with reliability and low latency.
  • Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools.
  • Architect and evaluate trade-offs in ML system design focusing on memory bandwidth and I/O throughput.
  • Collaborate with research and product teams to translate ML model requirements into efficient infrastructure solutions.
  • Define and track system-level metrics and SLIs to maintain production reliability of ML serving systems.
  • Lead technical design reviews and establish engineering standards for ML systems.
  • Mentor other engineers on ML infra best practices and optimization techniques.
  • Drive adoption of AI-augmented development workflows to boost productivity and expand deliverables.
  • Contribute to staged rollout strategies using feature flags and experimentation frameworks.

Skills

System design
Performance optimization
C++
Python
Profiling
Benchmarking
Cross-functional collaboration
Mentorship
Experimentation
ML infrastructure

Education

Bachelor's degree in CS/CE or equivalent

Tools

PyTorch
TensorFlow
CUDA

Job description

Meta is seeking a Software Engineer to join the Systems ML Engineering team, building and optimizing the ML infrastructure that powers Meta's products at massive scale. You will design high-performance ML systems spanning training, inference, and hardware-aware optimizations.

You will collaborate with researchers and product teams to accelerate ML workloads, implement production-ready pipelines, and drive reliability with system-level metrics and SLIs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems ML Engineer - Scalable AI Infrastructure
Systems ML Engineer - Scalable AI Infrastructure

Meta • Pierre (SD)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer: Scalable AI Infrastructure
Senior Systems ML Engineer: Scalable AI Infrastructure

Meta • Providence (RI)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infra
Senior Systems ML Engineer — Scalable AI Infra

Meta • Nashville (TN)

On-site
USD 154,000 - 217,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Bismarck (ND)

On-site
USD 154,000 - 217,000
Lead Systems ML Engineer — Scalable AI Infrastructure
Lead Systems ML Engineer — Scalable AI Infrastructure

Meta • Frankfort (KY)

On-site
USD 154,000 - 217,000
Systems ML Engineer: High-Performance AI Infra
Systems ML Engineer: High-Performance AI Infra

Meta • Saint Paul (MN)

On-site
USD 154,000 - 217,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Staff Systems ML Engineer — Scale AI Infrastructure
Staff Systems ML Engineer — Scale AI Infrastructure

Meta • Sunnyvale (CA), Menlo Park (CA), Bellevue (WA), Seattle (WA)

On-site
USD 260,000 - 360,000
Senior Systems ML Engineer — High-Performance AI Infra
Senior Systems ML Engineer — High-Performance AI Infra

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Lead Systems ML Engineer – High-Performance AI Infra
Lead Systems ML Engineer – High-Performance AI Infra

Meta • Sacramento (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits