Sr. AI Inference Platform Engineer

Apple

Seattle (WA)

On-site

USD 175,000 - 309,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
Retirement benefits
Discounted products
Tuition reimbursement
Discretionary bonuses

Job summary

Apple seeks an engineer to build tooling, automation, and analysis capabilities for the AI inference platform, focusing on benchmarking, capacity projection models, and data pipelines that inform infrastructure teams and capacity planners.

You will design real-time performance dashboards, collaborate with AI infrastructure engineers, hardware teams, and capacity planners to forecast needs months in advance, and help accelerate data-driven decisions for scaling AI workloads.

Qualifications

  • BS or MS in Computer Science or related technical field.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • Experience with performance and infrastructure engineering in distributed systems.

Responsibilities

  • Design and build automations to evaluate AI inference performance across hardware generations and configurations.
  • Develop tooling to surface performance trends, regressions, and insights to infrastructure and planning teams.
  • Build projection and forecasting models to support long-term capacity planning decisions.
  • Analyze performance and utilization data to identify bottlenecks, trends, and optimization opportunities.
  • Partner with AI infrastructure engineers, hardware teams, and capacity planners to deliver critical data and tooling.
  • Create and enhance performance analysis workflows to increase team velocity and data reliability.
  • Continuously improve the accuracy, coverage, and usability of performance measurement and analysis systems.

Skills

Python
Go
C++

Education

BS or MS in Computer Science

Tools

Nsight
Prometheus
Grafana

Job description

We are looking for an engineer to build tooling, automation, and analysis capabilities that strengthen our AI inference platform. This role will focus on developing sophisticated performance benchmarking systems, capacity projection models, and data analysis pipelines that directly inform our AI infrastructure teams and capacity planners. You'll work at the intersection of AI systems performance, distributed infrastructure, and software engineering to help the team make data-driven decisions about scaling and optimizing our inference platform.

Description

At Apple, we believe the future of AI is defined not just by models, but by the infrastructure that powers them. Our AI inference platform sits at the heart of products and experiences used by hundreds of millions of people worldwide, and we are building the systems that ensure it scales reliably, efficiently, and intelligently.

As part of our next-generation datacenter engineering team, you will play a critical role in shaping how we understand, measure, and grow our AI infrastructure. You will design and build the tooling and analysis systems that give our engineers and capacity planners a clear, real-time picture of performance across our fleet. Your work will directly influence how we invest in hardware, how we detect regressions before they reach production, and how we forecast capacity needs months in advance.

This is a high-impact, cross-functional role for an engineer who is energized by complexity, thrives on turning raw data into actionable insight, and wants to work on problems that matter at massive scale.

Responsibilities
  • Design and build automations to evaluate AI inference performance across hardware generations and configurations.
  • Develop tooling to surface performance trends, regressions, and insights to infrastructure and planning teams.
  • Build projection and forecasting models to support long-term capacity planning decisions.
  • Analyze performance and utilization data to identify bottlenecks, trends, and optimization opportunities.
  • Partner with AI infrastructure engineers, hardware teams, and capacity planners to deliver critical data and tooling.
  • Create and enhance performance analysis workflows to increase team velocity and data reliability.
  • Continuously improve the accuracy, coverage, and usability of performance measurement and analysis systems.
Preferred Qualifications
  • Experience with performance benchmarking and methodologies for AI/ML inference systems.
  • Familiarity with capacity planning and forecasting/projection models for large-scale infrastructure.
  • Experience with GPU profiling and observability tools (e.g., Nsight, other vendor-specific profilers).
  • Experience with data visualization and reporting tools/frameworks for surfacing performance trends to stakeholders.
  • Familiarity with ML serving frameworks and runtimes (e.g., Triton, TensorRT-LLM, vLLM, or similar).
  • Experience with CI/CD and workflow orchestration tools for building automated performance analysis pipelines.
  • Knowledge of cluster schedulers and orchestration platforms (e.g., Kubernetes).
  • Experience with metrics and logging tools (e.g., Prometheus, Grafana, Splunk).
Minimum Qualifications
  • BS or MS in Computer Science or related technical field.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • Experience with performance and infrastructure engineering in distributed systems.
  • Proficiency in Python, Go, C++, or other programming languages.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Excellent communication skills and ability to turn data into clear guidance for infrastructure teams and capacity planners.
Pay & Benefits

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $175,000 and $308,500, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relo

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Platform Engineer
AI Inference Platform Engineer

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
System Performance Engineer - AI/ML, Platform Architecture
System Performance Engineer - AI/ML, Platform Architecture

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Retirement benefits
Discounted products and services
+1
Full Stack Software Engineer - ML Compute Capacity
Full Stack Software Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
Apple benefits
Stock programs
Relocation assistance
+1
AIML - Power & Performance Engineer, Siri and Information Intelligence
AIML - Power & Performance Engineer, Siri and Information Intelligence

Apple Inc. • Bridge Creek (OK)

On-site
USD 171,000 - 300,000
Stock programs
Medical and dental coverage
Retirement benefits
+2
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock programs
SW Optimization Engineer AI/ML
SW Optimization Engineer AI/ML

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Software Engineer - Data Solutions, AI & Data Platform (AiDP)
Software Engineer - Data Solutions, AI & Data Platform (AiDP)

Apple Inc. • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Comprehensive medical and dental
Retirement benefits
Employee stock purchase plan
+1
Sr. Engineering Program Manager, ML Compute Infrastructure, Apple Services Engineering
Sr. Engineering Program Manager, ML Compute Infrastructure, Apple Services Engineering

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 312,000
AI Inference Platform Performance Engineer
AI Inference Platform Performance Engineer

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms
Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 181,100 - 318,400
Comprehensive medical and dental coverage
Employee stock purchase plan
Tuition reimbursement