Software Engineer, Systems ML Engineering

Meta

Menlo Park (CA)

On-site

USD 183,997 - 257,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta seeks a Staff Software Engineer for the Systems ML Engineering team to architect and own ML infrastructure that powers Meta's production AI workloads. You will shape distributed training, model serving, and platform tooling across a demanding production fleet.

You will lead design initiatives, optimize performance, and collaborate with ML researchers and engineers to translate requirements into robust, scalable systems.

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/engineering.
  • 8+ years of software eng experience focusing on systems software, distributed computing, or ML infrastructure.
  • Experience designing and implementing large-scale distributed systems.
  • Experience with performance analysis and optimization of compute-intensive workloads.
  • Leadership of end-to-end tech projects across teams.
  • Experience with C++, Python, or equivalent languages.

Responsibilities

  • Design and implement scalable ML systems infrastructure components (distributed training frameworks, model serving, ML platform tooling).
  • Lead technical design and architecture for major ML infrastructure initiatives.
  • Identify and resolve performance bottlenecks in distributed training and inference systems.
  • Define and drive service level objectives, dashboards, and incident runbooks for ML services.
  • Mentor engineers on ML systems best practices and distributed computing patterns.
  • Drive adoption of engineering standards, testing, feature flagging, and monitoring.

Skills

Distributed systems
C++
Python
ML infrastructure
Performance optimization
Leadership/mentorship

Education

Bachelor's degree in Computer Science/Engineering or related field

Tools

Model serving systems
Training orchestration

Job description

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focused on building and scaling the infrastructure and software systems that power large-scale machine learning workloads across Meta's production fleet. In this role, you will architect and own critical components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling. You will work at the intersection of systems engineering and machine learning to drive reliability, performance, and efficiency for some of the world's most demanding AI workloads, including large language models and generative AI systems.

Software Engineer, Systems ML Engineering Responsibilities:
  • Design and implement scalable ML systems infrastructure components, including distributed training frameworks, model serving pipelines, and ML platform tooling used across Meta's production AI workloads
  • Lead technical design and architecture for major initiatives in the ML systems stack, evaluating trade-offs across performance, reliability, and engineering complexity
  • Identify and resolve performance bottlenecks in distributed ML training and inference systems through instrumentation, profiling, and targeted optimization
  • Define and drive service level objectives for ML infrastructure services, building dashboards, alerting, and runbooks to reduce mean time to mitigation during incidents
  • Collaborate with machine learning researchers, product engineers, and infrastructure teams to translate model development requirements into robust, production-grade systems
  • Leverage AI-assisted development workflows to accelerate implementation, code review, and system analysis, applying sound judgment on when to rely on AI tooling versus deep domain expertise
  • Mentor other engineers on ML systems best practices, distributed computing patterns, and engineering craft, including AI-native development workflows
  • Drive adoption of engineering standards across the team, including testing strategies, staged rollout practices using feature flagging and experimentation frameworks, and proactive monitoring
  • Contribute to roadmap definition and stakeholder alignment for multi-quarter ML infrastructure investments, communicating technical options and trade-offs to both engineering and cross-functional audiences
  • Conduct thorough code reviews and establish coding standards that improve maintainability and scalability of the ML systems codebase
Minimum Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 8+ years of experience in software engineering with a focus on systems software, distributed computing, or ML infrastructure
  • Experience designing and implementing large-scale distributed systems, including components such as training orchestration, model serving, or data pipeline infrastructure
  • Experience with performance analysis and optimization of compute-intensive or distributed workloads, including profiling, benchmarking, and bottleneck identification
  • Experience leading end-to-end delivery of complex technical projects, including cross-team coordination, milestone planning, and risk mitigation
  • Experience with C++, Python, or equivalent systems programming languages applied to production ML or infrastructure systems
Preferred Qualifications:
  • Experience contributing to or maintaining open-source ML systems or distributed computing projects
  • Experience building or operating ML platform services including experiment tracking, model registries, feature stores, or inference serving infrastructure
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience with ML frameworks such as PyTorch, including distributed training paradigms such as data parallelism, model parallelism, or pipeline parallelism
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience with GPU computing, CUDA programming,或 accelerator-aware systems optimization for large-scale AI workloads
About Meta:

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$183,997/year to $257,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • Seattle (WA)

On-site
USD 184,000 - 257,000
Software Engineer, Systems ML (Technical Leadership)
Software Engineer, Systems ML (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • Bellevue (WA)

On-site
USD 183,997 - 257,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • Sunnyvale (CA)

On-site
USD 183,997 - 257,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • New York (NY)

On-site
USD 184,000 - 257,000
Software Engineer, AI Specialist - Monetization (Technical Leadership)
Software Engineer, AI Specialist - Monetization (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineer, Core Machine Learning
Software Engineer, Core Machine Learning

Meta • Bellevue (WA)

On-site
USD 184,000 - 257,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Jobzhr • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

SupportFinity™ • Washington

On-site
USD 154,000 - 217,000
Equity
Bonus
Benefits