Staff ML Compute & TPU Infrastructure Engineer

Apple Inc.

San Francisco (CA)

On-site

USD 210,000 - 300,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking an AIML Staff ML Infrastructure Engineer to advance pre-training infrastructure in the SF Bay Area. You will work on TPUs, JAX/XLA, and distributed training to push performance and scalability across large foundation models.

You will lead kernel development for attention/MoE, optimize workloads, and mentor engineers while collaborating with cross-functional teams to achieve engineering excellence.

Qualifications

  • 6+ years of experience building or optimizing high-performance ML or distributed systems.
  • Proficient in Python or other relevant programming languages.
  • Strong understanding of distributed systems, parallel computing, and performance optimization.
  • Ability to clearly communicate complex technical problems and collaborate with partners to develop solutions.

Responsibilities

  • Drive performance optimization for large-scale foundation model training on TPUs, focusing on efficiency, throughput, and scalability.
  • Profile and optimize JAX/XLA workloads across compute, memory, communication, and compilation.
  • Develop and optimize high-performance TPU kernels for critical ML operations such as attention and MoE.
  • Optimize distributed training techniques, sharding strategies, and collective communication over TPU interconnects (ICI/Fabric).
  • Research and implement new techniques across the JAX, XLA, and TPU stack to improve end-to-end training performance.
  • Develop performance profiling, benchmarking, and automated tuning capabilities for large-scale training workloads.
  • Collaborate with cross-functional engineers to solve large-scale ML training challenges.
  • Lead complex technical projects and mentor engineers in areas of your expertise.
  • Cultivate a team centered on collaboration, technical excellence, and innovation

Skills

Python programming
Distributed systems
Performance optimization

Education

Bachelor's degree in Computer Science, Engineering, or related field
Advanced degree in Computer Science, Engineering, or related field

Tools

JAX
XLA
PyTorch

Job description

Apple is seeking an AIML Staff ML Infrastructure Engineer to advance pre-training infrastructure in the SF Bay Area. You will work on TPUs, JAX/XLA, and distributed training to push performance and scalability across large foundation models.

You will lead kernel development for attention/MoE, optimize workloads, and mentor engineers while collaborating with cross-functional teams to achieve engineering excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

Socket.dev • New York (NY), California (MO)

On-site
USD 180,000 - 240,000
Senior ML Infrastructure Engineer — TPU & JAX Performance
Senior ML Infrastructure Engineer — TPU & JAX Performance

Socket.dev • New York (NY), California (MO)

On-site
USD 180,000 - 240,000
AIML Engineer: Build Scalable ML Pipelines & GPU Systems
AIML Engineer: Build Scalable ML Pipelines & GPU Systems

Apple Inc. • Cupertino (CA)

On-site
USD 184,700 - 324,800
Medical & dental
Retirement benefits
Discounted products
+1
Senior ML Foundation Model Compute Infra Engineer
Senior ML Foundation Model Compute Infra Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,700 - 324,800
Stock programs
Relocation
Tuition reimbursement
+1
Senior ML Infrastructure & Deployment Engineer
Senior ML Infrastructure & Deployment Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior ML Production Automation Engineer – On-Device AI
Senior ML Production Automation Engineer – On-Device AI

Apple • Cupertino (CA)

On-site
USD 185,000 - 325,000
Stock programs
Bonuses/relocation
Medical and dental coverage
+3
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000