Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations

Meta

Menlo Park, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking Software Engineers to join the Frontier Evals Research team within Meta Superintelligence Labs. You will build platforms and services for evaluating advanced AI models across data modalities, enabling researchers to iterate quickly.

This highly technical role focuses on distributed systems, developer infrastructure, and production-grade ML platforms, collaborating with researchers and engineers to deliver reliable, scalable systems.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent.
  • 3+ years building backend, distributed, data, or ML infrastructure.
  • Proficiency in Python, C++, or similar systems language.
  • Experience designing reliable services and data processing systems.
  • Experience delivering medium to large technical projects from design to production.

Responsibilities

  • Design and own systems for scheduling and executing evaluation workloads.
  • Manage datasets and model artifacts for reproducible research.
  • Monitor correctness, performance, and reliability of platforms.
  • Build scalable infrastructure supporting frontier AI research.

Skills

Python
C++
Distributed systems
ML infrastructure
Data processing
Systems design
Observability
Production operations
Research collaboration
Cloud infrastructure

Education

Bachelor's degree in CS/CE

Tools

Containers
Kubernetes
Workflow orchestration

Job description

Meta is seeking Software Engineers to join the Frontier Evals Research team within Meta Superintelligence Labs. Evaluations are a critical part of AI progress at Meta Superintelligence Labs, determining what capabilities get built, which features get prioritized, and how quickly our models improve. As a Systems and ML Infrastructure Engineer on this team, you will build the platforms and services that enable reliable evaluation of our most advanced AI models across text, vision, audio, and beyond. You'll work alongside researchers and engineers to turn rapidly evolving research workflows into scalable, dependable infrastructure.

This is a highly technical software engineering role focused on distributed systems, developer infrastructure, and production-grade ML platforms.

You will design and own systems for scheduling and executing evaluation workloads, managing datasets and model artifacts, monitoring correctness and performance, and making results reproducible and easy to consume. The infrastructure you build will directly support research decisions and major model lines within MSL, making reliability, scalability, operational excellence, and engineering rigor paramount.

You will succeed by moving quickly in an open-ended research environment while building durable systems, reducing operational toil, and creating abstractions that help researchers iterate faster. If you are passionate about building the technical foundation for frontier AI development and thrive in fast-paced, high-impact environments.

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 3+ years of software engineering experience in building backend, distributed, data, or machine learning infrastructure
  • Proficiency in Python, C++, or another systems programming language
  • Experience designing, implementing, and operating reliable services, platforms, or data-processing systems
  • Experience independently delivering medium- to large-scale technical projects from design through production operation
  • Demonstrated knowledge of software engineering practices, including testing, code review, observability, incident response, and performance analysis
  • Ability to work effectively with researchers and engineers and to adapt to rapidly changing requirements
  • Experience building infrastructure for large-scale machine learning training, inference, evaluation, or data processing
  • Experience with distributed compute systems, workflow orchestration, containers, cluster schedulers, or cloud infrastructure
  • Experience with performance profiling, resource efficiency, reliability engineering, and production observability
  • Familiarity with language model post-training workflows, including supervised fine-tuning, reinforcement learning, evaluation, and inference, and the infrastructure needed to support them at scale
  • Experience building internal platforms or developer tools used by multiple teams in fast-moving technical environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems & ML Infra Engineer for Frontier AI Evaluations
Systems & ML Infra Engineer for Frontier AI Evaluations

Meta • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Research Engineer - Meta Superintelligence Labs (Technical Leadership)
Research Engineer - Meta Superintelligence Labs (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Research Engineer, Pre-training Data - MSL FAIR
Research Engineer, Pre-training Data - MSL FAIR

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Senior ML Engineer - Technical Leader, Scale & Impact
Senior ML Engineer - Technical Leader, Scale & Impact

Meta • New York (NY)

On-site
USD 180,000 - 280,000
Software Engineer Manager - ML Infra
Software Engineer Manager - ML Infra

Meta • New York (NY)

On-site
USD 240,000 - 340,000
Machine Learning Engineer (Technical Leadership)
Machine Learning Engineer (Technical Leadership)

Meta • New York (NY)

On-site
USD 180,000 - 280,000
Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Machine Learning Engineer (Technical Leadership)
Machine Learning Engineer (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 210,000 - 320,000
Software Engineer - Backend Infrastructure, Standalone Apps Team
Software Engineer - Backend Infrastructure, Standalone Apps Team

Meta • Menlo Park (CA)

On-site
USD 190,000 - 260,000
Research Engineer (Technical Leadership), FAIR Data - Meta Superintelligence Labs
Research Engineer (Technical Leadership), FAIR Data - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000