Sr. AI Inference Platform Engineer

Apple Inc.

Seattle (WA)

On-site

USD 175,000 - 309,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical and dental coverage
Retirement benefits
Discounted products and free services
Tuition reimbursement

Job summary

Apple Inc. in Seattle seeks a senior engineer to build tooling, automation, and analysis capabilities that strengthen our AI inference platform.

You will develop performance benchmarking systems, capacity projection models, and data pipelines that inform infrastructure teams and capacity planners. This cross‑functional role sits at the intersection of AI systems performance, distributed infrastructure, and software engineering, translating complex data into actionable guidance for scaling and

Qualifications

  • BS or MS in Computer Science or related technical field.
  • Strong knowledge of AI/ML inference architecture and the performance characteristics of serving systems.
  • 7+ years of experience in performance and infrastructure engineering in distributed systems.
  • 7 years of coding experience in Python, Go, C++, or other programming languages.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Excellent communication skills and ability to turn data into clear guidance for infrastructure teams and capacity planners.

Responsibilities

  • Design and build automations to evaluate AI inference performance across hardware generations and configurations.
  • Develop tooling to surface performance trends, regressions, and insights to infrastructure and planning teams.
  • Build projection and forecasting models to support long-term capacity planning decisions.
  • Analyze performance and utilization data to identify bottlenecks, trends, and optimization opportunities.
  • Partner with AI infrastructure engineers, hardware teams, and capacity planners to deliver critical data and tooling.
  • Create and enhance performance analysis workflows to increase team velocity and data reliability.
  • Continuously improve the accuracy, coverage, and usability of performance measurement and analysis systems.

Skills

Performance engineering
Distributed systems
Python
Go
C++
Automation engineering
Data pipelines
GPU architecture
Statistics
Communication

Education

BS or MS in Computer Science or related technical field

Tools

Nsight
Prometheus
Grafana
Kubernetes
TensorRT-LLM

Job description

Seattle, Washington, United States Software and Services

We are looking for senior engineer to build tooling, automation, and analysis capabilities that strengthen our AI inference platform. This role will focus on developing sophisticated performance benchmarking systems, capacity projection models, and data analysis pipelines that directly inform our AI infrastructure teams and capacity planners. You'll work at the intersection of AI systems performance, distributed infrastructure, and software engineering to help the team make data-driven decisions about scaling and optimizing our inference platform.

Description

At Apple, we believe the future of AI is defined not just by models, but by the infrastructure that powers them. Our AI inference platform sits at the heart of products and experiences used by hundreds of millions of people worldwide, and we are building the systems that ensure it scales reliably, efficiently, and intelligently.As part of our next-generation datacenter engineering team, you will play a critical role in shaping how we understand, measure, and grow our AI infrastructure. You will design and build the tooling and analysis systems that give our engineers and capacity planners a clear, real-time picture of performance across our fleet. Your work will directly influence how we invest in hardware, how we detect regressions before they reach production, and how we forecast capacity needs months in advance.This is a high-impact, cross-functional role for an engineer who is energized by complexity, thrives on turning raw data into actionable insight, and wants to work on problems that matter at massive scale.

Responsibilities
  • Design and build automations to evaluate AI inference performance across hardware generations and configurations.
  • Develop tooling to surface performance trends, regressions, and insights to infrastructure and planning teams.
  • Build projection and forecasting models to support long-term capacity planning decisions.
  • Analyze performance and utilization data to identify bottlenecks, trends, and optimization opportunities.
  • Partner with AI infrastructure engineers, hardware teams, and capacity planners to deliver critical data and tooling.
  • Create and enhance performance analysis workflows to increase team velocity and data reliability.
  • Continuously improve the accuracy, coverage, and usability of performance measurement and analysis systems.
Minimum Qualifications
  • BS or MS in Computer Science or related technical field.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • 7 or more years of experience with performance and infrastructure engineering in distributed systems.
  • 7 years of experience coding in Python, Go, C++, or other programming languages.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Excellent communication skills and ability to turn data into clear guidance for infrastructure teams and capacity planners.
Preferred Qualifications
  • Experience with performance benchmarking and methodologies for AI/ML inference systems.
  • Familiarity with capacity planning and forecasting/projection models for large-scale infrastructure.
  • Experience with GPU profiling and observability tools (e.g., Nsight, other vendor-specific profilers).
  • Experience with data visualization and reporting tools/frameworks for surfacing performance trends to stakeholders.
  • Familiarity with ML serving frameworks and runtimes (e.g., Triton, TensorRT-LLM, vLLM, or similar).
  • Experience with CI/CD and workflow orchestration tools for building automated performance analysis pipelines.
  • Knowledge of cluster schedulers and orchestration platforms (e.g., Kubernetes).
  • Experience with metrics and logging tools (e.g., Prometheus, Grafana, Splunk).
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $175,000 and $308,500, and your base pay will depend on your skills, qualifications, experience, and location.Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple BenefitsNote: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. AI Inference Platform Engineer
Sr. AI Inference Platform Engineer

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
Senior Cloud Infrastructure and AI Efficiency Engineer
Senior Cloud Infrastructure and AI Efficiency Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Medical and dental coverage
Employee stock programs
Relocation assistance
+1
Software Engineer, Applied AI
Software Engineer, Applied AI

Apple Inc. • San Diego (CA)

On-site
USD 123,000 - 214,000
System Performance Engineer - AI/ML, Platform Architecture
System Performance Engineer - AI/ML, Platform Architecture

Apple Inc. • Santa Clara (CA), Northern (KY)

On-site
USD 150,000 - 225,000
Medical and dental coverage
Retirement benefits
Discounted products and services
+1
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc. • New York (NY)

On-site
USD 185,000 - 325,000
Employee stock programs
Bonus opportunities
Relocation assistance
+1
ML Infrastructure Engineer - ML Compute Capacity
ML Infrastructure Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA), Northern (KY)

On-site
USD 185,000 - 325,000
AIML - Distinguished Engineer, Foundation Model
AIML - Distinguished Engineer, Foundation Model

Apple Inc. • Cupertino (CA)

On-site
USD 311,000 - 497,000
Software Engineer, Reliability Engineering, AiDP
Software Engineer, Reliability Engineering, AiDP

Apple Inc. • Sunnyvale (CA), Northern (KY)

On-site
USD 150,000 - 225,000
Stock programs
Relocation assistance
Education reimbursement
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 175,000 - 309,000
Sr Manager
Sr Manager

Apple Inc. • Cupertino (CA)

On-site
USD 238,000 - 402,000