Research Scientist, Artificial Intelligence

Meta

Menlo Park (CA)

On-site

USD 184,000 - 257,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Bonus
Equity

Job summary

Meta AI Research is seeking a Staff Research Scientist to lead TPU performance optimization and large-scale model training within Meta's native PyTorch stack in Menlo Park. You will drive systems-level ML research, collaborate across research and engineering teams, and translate findings into production-ready optimizations at scale.

The role requires deep expertise in TPU kernels, memory management, and distributed training strategies, with a track record of impactful publications and mentorship.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 8+ years of experience in ML systems, model optimization, or HPC research.
  • Experience with TPU architecture and performance optimization, including profiling, kernel development, and memory management.
  • Experience with XLA compilation, graph optimization, and low-level performance tuning for accelerator hardware.
  • Experience developing and optimizing large-scale distributed training systems, including parallelism strategies.
  • Experience with PyTorch and its integration with accelerator backends.
  • Experience communicating complex technical findings in writing, including technical reports and publications.

Responsibilities

  • Lead the design and execution of TPU performance optimization research, including kernel development, memory optimization, and compute efficiency improvements.
  • Develop and optimize Pallas kernels for large-scale model training and inference on TPU architectures.
  • Drive model optimization techniques including MoE, tensor, and pipeline parallelism and other distributed training strategies.
  • Optimize first party models within Meta's native PyTorch stack, integrating with XLA and TPU execution.
  • Identify and resolve complex challenges in training efficiency, latency, and reliability with novel approaches.
  • Define and drive multi-quarter research roadmaps for TPU optimization and align with organizational goals.
  • Establish experimentation frameworks for benchmarking, profiling, and data-driven decisions.
  • Translate research findings into production-ready optimizations with deployment and reliability at scale.
  • Communicate findings and trade-offs through publications, design docs, and presentations to technical and non-technical audiences.
  • Mentor researchers and engineers on TPU optimization techniques and experimental rigor.

Skills

TPU performance optimization
Large-scale model training
Systems ML research
PyTorch integration
Distributed training

Education

Bachelor's degree in CS/CE or equivalent

Tools

XLA compilation
Kernel development
Memory management
Profiling

Job description

Meta AI Research is at the forefront of advancing foundational and applied artificial intelligence, developing breakthroughs that power products used by billions of people and shape the future of human-computer interaction. We are seeking a Research Scientist at the Staff level (IC6) with deep expertise in TPU performance optimization, large-scale model training, and systems-level machine learning. In this role, you will lead high-impact research on model efficiency and optimization for first party models within Meta's native PyTorch stack, collaborating across research and engineering teams to drive AI capabilities that define Meta's next generation of products and platforms.

Research Scientist, Artificial Intelligence Responsibilities:
  • Lead the design and execution of TPU performance optimization research, including kernel development, memory optimization, and compute efficiency improvements
  • Develop and optimize Pallas kernels for large-scale model training and inference on TPU architectures
  • Drive model optimization techniques including Mixture of Experts (MoE), tensor parallelism, pipeline parallelism, and other distributed training strategies
  • Optimize first party models within Meta's native PyTorch stack, ensuring efficient integration with XLA compilation and TPU execution
  • Identify and resolve complex technical challenges in model training efficiency, inference latency, and system reliability that require novel approaches
  • Define and drive multi-quarter research roadmaps for TPU optimization, aligning project milestones with broader organizational goals
  • Establish rigorous experimentation frameworks for performance benchmarking, including metric selection, profiling methodology, and data-driven optimization decisions
  • Translate research findings into production-ready optimizations by collaborating with engineering teams on deployment pipelines and reliability at scale
  • Communicate research findings and technical trade-offs clearly through publications, design documents, and presentations to both technical and non-technical audiences
  • Mentor other researchers and engineers on TPU optimization techniques, providing structured feedback on technical direction and experimental rigor
Minimum Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 8+ years of experience in machine learning systems, model optimization, or high-performance computing research
  • Experience with TPU architecture and performance optimization, including profiling, kernel development, and memory management
  • Experience with XLA compilation, graph optimization, and low-level performance tuning for accelerator hardware
  • Experience developing and optimizing large-scale distributed training systems, including parallelism strategies such as data, tensor, and pipeline parallelism
  • Experience with PyTorch and its integration with accelerator backends
  • Experience communicating complex technical findings in writing, including technical reports, design documents, or peer-reviewed publications
Preferred Qualifications:
  • Experience developing custom kernels using Pallas or similar kernel authoring frameworks for TPU or GPU
  • Demonstrated track record of transitioning performance research into deployed systems used at significant scale
  • Publication record in systems for ML venues such as MLSys, OSDI, SOSP, or related AI conferences such as NeurIPS, ICML, or ICLR
  • Experience with Mixture of Experts (MoE) architectures and their optimization for efficient training and inference
  • Experience optimizing production-scale models with billions of parameters
About Meta:

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and Aff… . We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$184,000/year to $257,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Machine Learning
Research Scientist, Machine Learning

Meta • Bellevue (WA)

On-site
USD 122,000 - 181,000
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
AI Research Scientist, Computer Vision
AI Research Scientist, Computer Vision

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
AI Research Engineer
AI Research Engineer

Meta • Seattle (WA)

On-site
USD 184,000 - 257,000
AI Research Scientist, Robotics - Meta Superintelligence Labs
AI Research Scientist, Robotics - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer, Machine Learning
Software Engineer, Machine Learning

Meta • New York (NY)

On-site
USD 183,997 - 257,000
Bonus
Equity
Benefits
Research Scientist, Machine Learning for Monetization (PhD)
Research Scientist, Machine Learning for Monetization (PhD)

Meta • New York (NY)

On-site
USD 122,000 - 181,000
Equity
Bonus
Benefits
AI Research Scientist - Meta Superintelligence Labs (Technical Leadership)
AI Research Scientist - Meta Superintelligence Labs (Technical Leadership)

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Bonus
Equity
Benefits
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • Honolulu (HI)

On-site
USD 154,000 - 217,000
Research Scientist, Machine Learning for Monetization (PhD)
Research Scientist, Machine Learning for Monetization (PhD)

Meta • Bellevue (WA)

On-site
USD 122,000 - 181,000
Bonus
Equity
Benefits