Staff AI Research Scientist: TPU & Large-Scale ML

Meta

Menlo Park (CA)

On-site

USD 184,000 - 257,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus
Equity

Job summary

Meta AI Research is seeking a Staff Research Scientist to lead TPU performance optimization and large-scale model training within Meta's native PyTorch stack in Menlo Park. You will drive systems-level ML research, collaborate across research and engineering teams, and translate findings into production-ready optimizations at scale.

The role requires deep expertise in TPU kernels, memory management, and distributed training strategies, with a track record of impactful publications and mentorship.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • 8+ years of experience in ML systems, model optimization, or HPC research.
  • Experience with TPU architecture and performance optimization, including profiling, kernel development, and memory management.
  • Experience with XLA compilation, graph optimization, and low-level performance tuning for accelerator hardware.
  • Experience developing and optimizing large-scale distributed training systems, including parallelism strategies.
  • Experience with PyTorch and its integration with accelerator backends.
  • Experience communicating complex technical findings in writing, including technical reports and publications.

Responsibilities

  • Lead the design and execution of TPU performance optimization research, including kernel development, memory optimization, and compute efficiency improvements.
  • Develop and optimize Pallas kernels for large-scale model training and inference on TPU architectures.
  • Drive model optimization techniques including MoE, tensor, and pipeline parallelism and other distributed training strategies.
  • Optimize first party models within Meta's native PyTorch stack, integrating with XLA and TPU execution.
  • Identify and resolve complex challenges in training efficiency, latency, and reliability with novel approaches.
  • Define and drive multi-quarter research roadmaps for TPU optimization and align with organizational goals.
  • Establish experimentation frameworks for benchmarking, profiling, and data-driven decisions.
  • Translate research findings into production-ready optimizations with deployment and reliability at scale.
  • Communicate findings and trade-offs through publications, design docs, and presentations to technical and non-technical audiences.
  • Mentor researchers and engineers on TPU optimization techniques and experimental rigor.

Skills

TPU performance optimization
Large-scale model training
Systems ML research
PyTorch integration
Distributed training

Education

Bachelor's degree in CS/CE or equivalent

Tools

XLA compilation
Kernel development
Memory management
Profiling

Job description

Meta AI Research is seeking a Staff Research Scientist to lead TPU performance optimization and large-scale model training within Meta's native PyTorch stack in Menlo Park. You will drive systems-level ML research, collaborate across research and engineering teams, and translate findings into production-ready optimizations at scale.

The role requires deep expertise in TPU kernels, memory management, and distributed training strategies, with a track record of impactful publications and mentorship.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Artificial Intelligence
Research Scientist, Artificial Intelligence

Meta Careers • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Research Scientist, Artificial Intelligence
Research Scientist, Artificial Intelligence

Meta • Menlo Park (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Research Scientist, Machine Learning
Research Scientist, Machine Learning

Meta • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Research Scientist, Systems ML - HW/SW Co-Design
Research Scientist, Systems ML - HW/SW Co-Design

Meta • Menlo Park (CA)

On-site
USD 220,000 - 300,000
Senior ML Engineer — Scale Training Infra & AI Deployments
Senior ML Engineer — Scale Training Infra & AI Deployments

Best AI Tools Wiki • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Equity package
Health insurance
Unlimited PTO
+3
Staff ML Compute & TPU Infrastructure Engineer
Staff ML Compute & TPU Infrastructure Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
Staff ML Systems Architect - Co-Design & TPU Performance
Staff ML Systems Architect - Co-Design & TPU Performance

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff ML Research Engineer — From Prototype to Production
Staff ML Research Engineer — From Prototype to Production

Autonomous Technologies Group • New York (NY)

On-site
USD 110,000 - 150,000
ML Performance Engineering Manager (TPU & Optimization)
ML Performance Engineering Manager (TPU & Optimization)

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU Performance & ML Infrastructure
Staff Software Engineer, TPU Performance & ML Infrastructure

Google • New York (NY)

On-site
USD 207,000 - 300,000