Staff/Sr. ML Compute Efficiency Engineer

Socket.dev

Santa Clara (CA)

On-site

USD 130,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple’s Machine Learning Platform Technologies organization seeks a performance engineer to tackle challenges across thousands of GPUs/TPUs, optimize accelerator utilization, and reduce idle capacity while shortening recovery periods. This role focuses on improving efficiency across the ML compute fleet.

You will analyze accelerator performance, explore parallelism techniques, and refine scheduling and orchestration in collaboration with ML research and infrastructure teams.

Qualifications

  • Experience with large-scale distributed AI/ML workloads on GPUs/TPUs.
  • Strong software engineering skills with experience developing and optimizing training frameworks using C/C++ or Python.
  • Experience working on cross-functional projects with ML research and infrastructure teams.
  • Familiarity with model architectures and various training techniques.

Responsibilities

  • Analyze accelerator performance and identify inefficiencies to maximize utilization.
  • Explore parallelism techniques and refine workload scheduling and orchestration across the compute fleet.
  • Collaborate with ML research and infrastructure teams to implement performance improvements.

Skills

Distributed systems
C/C++
Python
PyTorch
JAX
Training frameworks
Cross-functional collaboration

Education

Bachelor’s degree in Computer Science or equivalent experience

Job description

Scaling machine learning workloads across thousands of GPUs and TPUs creates challenges that few engineers ever encounter. In Apple’s Machine Learning Platform Technologies organization, we build the infrastructure that powers large-scale ML training and inference workloads, bringing together expertise in distributed systems, machine learning infrastructure, and high-performance computing.

Description

As a performance engineer in the ML Compute Efficiency team, you’ll tackle ambiguous systems challenges, identify inefficiencies and build solutions that maximize accelerator utilization, reduce idle and fragmented capacity, and minimize recovery periods. This includes analyzing accelerator performance, digging into various parallelism techniques, and refining workload scheduling and orchestration across the compute fleet.

Minimum Qualifications

Experience with large-scale distributed systems for AI/ML workloads running on GPUs or TPUs. Strong software engineering skills with experience developing and optimizing training frameworks (e.g. PyTorch, JAX) using C/C++ or Python. Experience working on cross-functional projects with ML research and infrastructure teams. Familiarity with model architectures and various training techniques. Bachelor’s degree in Computer Science or equivalent experience, with 7+ years of industry experience.

Preferred Qualifications

Have a track record of delivering transformative performance improvements on large scale infrastructure. Ability to analyze ambiguous, distributed systems problems and articulate both high-level strategic metrics and underlying technical complexity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Compute Efficiency Engineer (GPU/TPU Performance)
Senior ML Compute Efficiency Engineer (GPU/TPU Performance)

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
SW Optimization Engineer AI/ML
SW Optimization Engineer AI/ML

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
SW ML Optimization Engineer
SW ML Optimization Engineer

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning
On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 260,000
ML Compute Efficiency Automation Engineer, Infrastructure & Planning
ML Compute Efficiency Automation Engineer, Infrastructure & Planning

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical coverage
Employee stock purchase program
Educational reimbursement
ML Framework (MetalLM) Engineer, Graphics, Game and ML
ML Framework (MetalLM) Engineer, Graphics, Game and ML

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000
AI Inference Platform Engineer
AI Inference Platform Engineer

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
AIML Senior Capacity Engineer - Apple Services Engineering
AIML Senior Capacity Engineer - Apple Services Engineering

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Educational expense reimbursement
ML Framework (MetalLM) Engineer
ML Framework (MetalLM) Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 226,000
Medical and dental coverage
Retirement benefits
Stock programs / RSUs
+2
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000