ML Framework (MetalLM) Engineer, Graphics, Game and ML

Apple

Cupertino (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple’s Server ML Frameworks team in Cupertino seeks engineers to optimize high-performance GenAI applications. This role focuses on ML inference framework projects and GPU programming while collaborating with cross-functional teams to enhance performance.

The ideal candidates will have over 3 years of experience in C/C++ and proficiency in GPU kernel development using Metal or CUDA. Join us to contribute to the future of Apple Silicon in the data center.

Qualifications

  • 3+ years of programming experience in C/C++/ObjC.
  • Experience with GPU kernel development and optimizations.
  • Strong understanding of system-level programming and computer architecture.

Responsibilities

  • Optimize ML inference using distributed compute strategies.
  • Develop kernel and compiler level optimizations.
  • Analyze and improve performance metrics.

Skills

C/C++/ObjC programming
GPU kernel development
Distributed training or inference
System-level programming

Tools

Metal
CUDA
Graph compilers (CuTE, CuTile, Triton)

Job description

Summary

Apple’s Server ML Frameworks team in GPU, Graphics and Machine Learning works on enabling Apple Intelligence through high-performance, distributed inference of GenAI applications (such as LLMs) on Private Cloud Compute. You will get to work on custom-built server hardware that brings the power and security of Apple silicon to the data center. We are looking for engineers with systems background who are deeply passionate about building scalable, efficient, and production-grade solutions tailored for high-throughput GPU execution.

Description

Our team is seeking extraordinary machine learning and GPU programming engineers who are passionate about providing robust compute solutions for accelerating Machine learning libraries on Apple Silicon. Role has the opportunity to influence the design of compute and programming models in next generation GPU architectures.

Responsibilities
  • Work on cutting-edge ML inference framework project and optimize code for efficient and scalable ML inference using distributed compute strategies such as data, tensor, pipeline and expert parallelism.
  • Develop kernel and compiler level optimizations and perform in-depth analysis to ensure the best possible performance across Server hardware families.
  • Apply advanced model optimization techniques including speculation, quantization, compression, and others to maximize throughput and minimize latency.
  • Collaborate closely with hardware, compiler, and systems teams to align software performance with hardware capabilities.
  • Analyze and improve performance metrics such as end-to-end latency, TTFT, TBOT, memory footprint, and compute efficiency.
Minimum Qualifications
  • 3+ years of programming and problem-solving experience with C/C++/ObjC
  • Experience with GPU kernel development & optimizations using compute programming models such as Metal, CUDA etc.
  • Experience with Distributed training or inference techniques
  • Experience with system level programming and computer architecture
Preferred Qualifications
  • Experience with graph compilers such as CuTE, CuTile, Triton, OpenXLA or LLVM is a plus
  • Good understanding of LLM and Diffusion based model architectures
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Framework (MetalLM) Engineer, Graphics, Game and ML
ML Framework (MetalLM) Engineer, Graphics, Game and ML

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Employee stock purchase plan
Comprehensive medical and dental coverage
Educational expense reimbursement
+1
ML Framework Engineer (MetalLM) for GPU Inference
ML Framework Engineer (MetalLM) for GPU Inference

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Employee stock purchase plan
Comprehensive medical and dental coverage
Educational expense reimbursement
+1
MetalLM ML Framework Engineer — GPU & Server Inference
MetalLM ML Framework Engineer — GPU & Server Inference

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000
Pre-silicon ML and Compute Framework Engineer, Graphics, Game and ML
Pre-silicon ML and Compute Framework Engineer, Graphics, Game and ML

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000
GPU ML Engineer - High-Performance ML on Silicon
GPU ML Engineer - High-Performance ML on Silicon

Apple Inc. • Cupertino (CA)

On-site
USD 150,400 - 277,600
Medical & dental
Retirement benefits
Employee stock programs
+2
Pre-Silicon ML & GPU Compute Framework Architect
Pre-Silicon ML & GPU Compute Framework Architect

Apple • Cupertino (CA)

On-site
USD 120,000 - 160,000
SW ML Optimization Engineer
SW ML Optimization Engineer

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
GPU ML Engineer: Accelerate ML on Modern SoCs
GPU ML Engineer: Accelerate ML on Modern SoCs

PVH (Tommy Hilfiger/Calvin Klein) • Cupertino (CA)

On-site
USD 150,400 - 277,600
Medical and dental coverage
Employee stock programs
Relocation support
+1
On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning
On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 260,000