Senior ML Compute Efficiency Engineer (GPU/TPU Performance)

Socket.dev

Santa Clara (CA)

On-site

USD 130,000 - 180,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple’s Machine Learning Platform Technologies organization seeks a performance engineer to tackle challenges across thousands of GPUs/TPUs, optimize accelerator utilization, and reduce idle capacity while shortening recovery periods. This role focuses on improving efficiency across the ML compute fleet.

You will analyze accelerator performance, explore parallelism techniques, and refine scheduling and orchestration in collaboration with ML research and infrastructure teams.

Qualifications

  • Experience with large-scale distributed AI/ML workloads on GPUs/TPUs.
  • Strong software engineering skills with experience developing and optimizing training frameworks using C/C++ or Python.
  • Experience working on cross-functional projects with ML research and infrastructure teams.
  • Familiarity with model architectures and various training techniques.

Responsibilities

  • Analyze accelerator performance and identify inefficiencies to maximize utilization.
  • Explore parallelism techniques and refine workload scheduling and orchestration across the compute fleet.
  • Collaborate with ML research and infrastructure teams to implement performance improvements.

Skills

Distributed systems
C/C++
Python
PyTorch
JAX
Training frameworks
Cross-functional collaboration

Education

Bachelor’s degree in Computer Science or equivalent experience

Job description

Apple’s Machine Learning Platform Technologies organization seeks a performance engineer to tackle challenges across thousands of GPUs/TPUs, optimize accelerator utilization, and reduce idle capacity while shortening recovery periods. This role focuses on improving efficiency across the ML compute fleet.

You will analyze accelerator performance, explore parallelism techniques, and refine scheduling and orchestration in collaboration with ML research and infrastructure teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff/Sr. ML Compute Efficiency Engineer
Staff/Sr. ML Compute Efficiency Engineer

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
AI/ML Performance Engineer — Apple Silicon
AI/ML Performance Engineer — Apple Silicon

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
AI/ML SoC Performance Engineer
AI/ML SoC Performance Engineer

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
GPU ML Engineer - High-Performance ML on Silicon
GPU ML Engineer - High-Performance ML on Silicon

Apple Inc. • Cupertino (CA)

On-site
USD 150,400 - 277,600
Medical & dental
Retirement benefits
Employee stock programs
+2
SW Optimization Engineer AI/ML
SW Optimization Engineer AI/ML

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
SW ML Optimization Engineer
SW ML Optimization Engineer

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
TPU Kernel Engineer for High-Performance ML Systems
TPU Kernel Engineer for High-Performance ML Systems

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000
ML Performance Optimization Engineer for SoC
ML Performance Optimization Engineer for SoC

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
GPU ML Engineer: Accelerate ML on Modern SoCs
GPU ML Engineer: Accelerate ML on Modern SoCs

PVH (Tommy Hilfiger/Calvin Klein) • Cupertino (CA)

On-site
USD 150,400 - 277,600
Medical and dental coverage
Employee stock programs
Relocation support
+1
Automation Engineer, ML Compute Efficiency & Scale
Automation Engineer, ML Compute Efficiency & Scale

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical coverage
Employee stock purchase program
Educational reimbursement