Software Engineer, Systems ML Engineering

Meta

Sunnyvale, Menlo Park, Bellevue, Seattle (CA, CA, WA, WA)

On-site

USD 260,000 - 360,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focusing on building and scaling infrastructure and software systems that power large-scale ML workloads across Meta's production fleet. You will architect and own components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling.

You will work at the intersection of systems engineering and machine learning to drive reliability,

Qualifications

  • Bachelor's degree or equivalent practical experience in a relevant field.
  • 8+ years of software engineering experience in systems software, distributed computing, or ML infrastructure.
  • Experience designing large-scale distributed systems, including training orchestration, model serving, or data pipelines.
  • Experience with PyTorch and distributed training.
  • Experience with performance profiling and bottleneck identification.
  • Experience contributing to or maintaining open-source ML systems or distributed computing projects.

Responsibilities

  • Architect and own components of the ML systems stack.
  • Scale infrastructure and software systems for large-scale ML workloads across the production fleet.
  • Drive reliability, performance, and efficiency for AI workloads including LLMs and generative AI systems.
  • Collaborate across teams to deliver end-to-end projects with milestone planning and risk mitigation.

Skills

C++
Python
Distributed systems
ML infrastructure
Performance optimization
Cross-team coordination

Education

Bachelor's degree in CS/CE or related field

Tools

PyTorch
CUDA

Job description

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focused on building and scaling the infrastructure and software systems that power large-scale machine learning workloads across Meta's production fleet. In this role, you will architect and own critical components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling. You will work at the intersection of systems engineering and machine learning to drive reliability, performance, and efficiency for some of the world's most demanding AI workloads, including large language models and generative AI systems.

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience8+ years of experience in software engineering with a focus on systems software, distributed computing, or ML infrastructure
  • Experience designing and implementing large-scale distributed systems, including components such as training orchestration, model serving, or data pipeline infrastructure
  • Experience with performance analysis and optimization of compute-intensive or distributed workloads, including profiling, benchmarking, and bottleneck identification
  • Experience leading end-to-end delivery of complex technical projects, including cross-team coordination, milestone planning, and risk mitigation
  • Experience with C++, Python, or equivalent systems programming languages applied to production ML or infrastructure systems
  • Experience contributing to or maintaining open-source ML systems or distributed computing projects
  • Experience building or operating ML platform services including experiment tracking, model registries, feature stores, or inference serving infrastructure
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience with ML frameworks such as PyTorch, including distributed training paradigms such as data parallelism, model parallelism, or pipeline parallelism
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience with GPU computing, CUDA programming, or accelerator-aware systems optimization for large-scale AI workloads
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, GenAI Frameworks
Software Engineer, GenAI Frameworks

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Software Engineer, Recommendation Systems
Software Engineer, Recommendation Systems

Meta • Menlo Park (CA)

On-site
USD 230,000 - 310,000
Senior ML Engineer - Technical Leader, Scale & Impact
Senior ML Engineer - Technical Leader, Scale & Impact

Meta • New York (NY)

On-site
USD 180,000 - 280,000
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • Nashville (TN)

On-site
USD 154,000 - 217,000
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Software Engineer, Data Infrastructure (Technical Leadership)
Software Engineer, Data Infrastructure (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 250,000 - 320,000
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • Frankfort (KY)

On-site
USD 154,000 - 217,000
Machine Learning Engineer (Technical Leadership)
Machine Learning Engineer (Technical Leadership)

Meta • New York (NY)

On-site
USD 180,000 - 280,000
Machine Learning Engineer (Technical Leadership)
Machine Learning Engineer (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 210,000 - 320,000
Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations
Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations

Meta • Menlo Park (CA), Northern (KY)

On-site
USD 180,000 - 240,000