Software Engineer, GenAI Frameworks

Meta

New York (NY)

On-site

USD 122,000 - 181,000

Full time

19 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is expanding its Systems ML Engineering team to design and optimize scalable ML training and inference pipelines. You will contribute to data processing, feature handling, and model serving infrastructure across large-scale distributed environments.

You will work with researchers and product engineers to translate model requirements into robust, high-throughput systems, participating in staged rollouts and ensuring code quality with reviews and documentation.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
  • 2+ years of software engineering focusing on ML systems or distributed data infrastructure.
  • Experience writing production-quality code in Python and a compiled language (C++/Java).
  • Experience designing ML training or inference pipelines (data preprocessing, model execution, serving).
  • Experience with distributed computing concepts like data/model parallelism in ML workloads.
  • Experience writing automated tests, logging, alerting, and production incident response.

Responsibilities

  • Design scalable systems for distributed ML training and inference, including data ingestion pipelines and model serving infrastructure.
  • Develop and optimize ML platform components like training orchestration and gradient communication.
  • Profile performance and diagnose bottlenecks across the ML stack; optimize latency and throughput.
  • Write automated tests and build monitoring/alerting for production anomalies.
  • Collaborate with ML researchers and engineers to translate model requirements into reliable system designs.
  • Own technical design for ML infrastructure features, balancing throughput, latency, and cost.

Skills

Python
C++
Java
Distributed systems
ML systems
Automated tests

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

PyTorch
JAX
Apache Spark
Flink
Kafka

Job description

Meta is building the infrastructure and systems that power machine learning at scale across its family of products, including Feed, Reels, Ads ranking, and generative AI services. The Systems ML Engineering team sits at the intersection of ML and systems software, designing and optimizing the training and inference pipelines, distributed execution frameworks, and data processing systems that enable researchers and product teams to iterate quickly and deploy reliably. In this role, you will contribute to the full lifecycle of ML systems software — from designing scalable data pipelines and distributed training infrastructure to optimizing model serving latency and throughput — directly impacting the quality and speed of ML-powered experiences for billions of people.

Software Engineer, GenAI Frameworks Responsibilities:
  • Design and implement scalable systems for distributed ML training and inference, including data ingestion pipelines, feature processing, and model serving infrastructure
  • Develop and optimize ML platform components such as training orchestration, gradient communication, and checkpoint management across large-scale distributed environments
  • Profile and diagnose performance bottlenecks across the ML stack, including data loading, preprocessing, forward and backward passes, and serving latency
  • Write automated tests covering expected behaviors, failure modes, and error paths for ML systems components, and build monitoring and alerting for production anomalies
  • Collaborate with ML researchers and product engineers to translate model requirements into reliable, high-throughput system designs
  • Own technical design for features and components within ML infrastructure, evaluating trade-offs between throughput, latency, cost, and engineering maintainability
  • Participate in staged rollouts of ML system changes using feature flagging and A/B testing frameworks, monitoring key metrics and responding to regressions
  • Contribute to code quality through code reviews, clear technical documentation, and consolidation of duplicative implementations across the ML systems codebase
  • Support on-call rotations for owned ML infrastructure, investigating production incidents and contributing detailed retrospectives to prevent recurrence
Minimum Qualifications:
  • Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
  • 2+ years of experience in software engineering with a focus on machine learning systems, distributed systems, or high-performance data infrastructure
  • Experience writing production-quality code in Python and at least one compiled language such as C++ or Java, including performance-sensitive systems code
  • Experience designing and implementing components of ML training or inference pipelines, including data preprocessing, model execution, or serving systems
  • Experience with distributed computing concepts such as data parallelism, model parallelism, or parameter server architectures as applied to ML workloads
  • Experience writing automated tests, building logging and alerting, and participating in production incident response for large-scale systems
Preferred Qualifications:
  • Experience optimizing ML workloads for throughput and latency, including profiling GPU or CPU utilization, memory bandwidth, and communication overhead in distributed training or inference settings
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Experience with ML framework internals such as PyTorch or JAX, including custom operator development, execution graph optimization, or compiler integration
  • Experience with stream or batch data processing systems such as Apache Spark, Flink, or Kafka in the context of ML feature pipelines
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • 1+ years production experience in GenAI post-training, RLHF, RLVR, comms/collectives, parallelism, accuracy evaluation, and performance profiling
  • Familiarity with model quantization, mixed-precision training, or other techniques for reducing compute and memory costs in production ML systems
About Meta:

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and ______? We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$121,992/year to $181,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, GenAI Frameworks
Software Engineer, GenAI Frameworks

Meta • Bellevue (WA)

On-site
USD 122,000 - 181,000
Software Engineer, Systems ML (Technical Leadership)
Software Engineer, Systems ML (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Software Engineer, AI Specialist - Monetization (Technical Leadership)
Software Engineer, AI Specialist - Monetization (Technical Leadership)

Meta • New York (NY)

On-site
USD 219,000 - 301,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Burlingame (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • San Francisco (CA)

On-site
USD 183,000 - 257,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

SupportFinity™ • Washington

On-site
USD 154,000 - 217,000
Equity
Bonus
Benefits
Software Engineering Manager - Neural Interface ML Infra
Software Engineering Manager - Neural Interface ML Infra

Meta • Redmond (WA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Software Engineer, Infrastructure
Software Engineer, Infrastructure

SupportFinity™ • Burlingame (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer Leadership, Machine Learning RecSys
Software Engineer Leadership, Machine Learning RecSys

Meta • Burlingame (CA)

On-site
USD 219,000 - 301,000