AI Inference Platform Engineer

Socket.dev

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking an engineer to build tooling, automation, and analysis capabilities for its AI inference platform. You will develop advanced performance benchmarking systems, capacity projection models, and data pipelines to inform capacity-planning and infrastructure decisions.

You will work at the intersection of AI systems performance, distributed infrastructure, and software engineering to turn raw data into actionable insight for scaling and optimizing the inference platform.

Qualifications

  • BS or MS in Computer Science or related field.
  • Strong understanding of AI/ML inference architecture and performance characteristics of serving systems.
  • Experience with performance and infrastructure engineering in distributed systems.
  • Proficiency in Python, Go, C++, or other languages.
  • Experience with automation, tooling and data pipelines for engineering workflows.
  • Strong knowledge of GPU/accelerator architecture for AI workloads.
  • Practical statistics applicable to performance analysis and forecasting.
  • Excellent communication skills to guide infrastructure teams.

Responsibilities

  • Build tooling and analysis systems to monitor AI inference performance at scale.
  • Develop benchmarking frameworks and capacity projection models for forecasting needs.
  • Create data pipelines to support engineering workflows and capacity planning.
  • Collaborate with AI infrastructure teams to inform hardware investments and scaling decisions.

Skills

AI/ML inference architecture
Distributed systems performance
Python/Go/C++ programming
Automation tooling
GPU/accelerator architecture
Statistical analysis
Clear communication

Education

BS or MS in Computer Science

Tools

Nsight
Prometheus
Grafana
Kubernetes

Job description

We are looking for an engineer to build tooling, automation, and analysis capabilities that strengthen our AI inference platform. This role will focus on developing sophisticated performance benchmarking systems, capacity projection models, and data analysis pipelines that directly inform our AI infrastructure teams and capacity planners. You'll work at the intersection of AI systems performance, distributed infrastructure, and software engineering to help the team make data-driven decisions about scaling and optimizing our inference platform.

Description

At Apple, we believe the future of AI is defined not just by models, but by the infrastructure that powers them. Our AI inference platform sits at the heart of products and experiences used by hundreds of millions of people worldwide, and we are building the systems that ensure it scales reliably, efficiently, and intelligently. As part of our next-generation datacenter engineering team, you will play a critical role in shaping how we understand, measure, and grow our AI infrastructure. You will design and build the tooling and analysis systems that give our engineers and capacity planners a clear, real-time picture of performance across our fleet. Your work will directly influence how we invest in hardware, how we detect regressions before they reach production, and how we forecast capacity needs months in advance. This is a high-impact, cross-functional role for an engineer who is energized by complexity, thrives on turning raw data into actionable insight, and wants to work on problems that matter at massive scale.

Minimum Qualifications
  • BS or MS in Computer Science or related technical field.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • Experience with performance and infrastructure engineering in distributed systems.
  • Proficiency in Python, Go, C++, or other programming languages.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Excellent communication skills and ability to turn data into clear guidance for infrastructure teams and capacity planners.
Preferred Qualifications
  • Experience with performance benchmarking and methodologies for AI/ML inference systems.
  • Familiarity with capacity planning and forecasting/projection models for large-scale infrastructure.
  • Experience with GPU profiling and observability tools (e.g., Nsight, other vendor-specific profilers).
  • Experience with data visualization and reporting tools/frameworks for surfacing performance trends to stakeholders.
  • Familiarity with ML serving frameworks and runtimes (e.g., Triton, TensorRT-LLM, vLLM, or similar).
  • Experience with CI/CD and workflow orchestration tools for building automated performance analysis pipelines.
  • Knowledge of cluster schedulers and orchestration platforms (e.g., Kubernetes).
  • Experience with metrics and logging tools (e.g., Prometheus, Grafana, Splunk).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Platform Performance Engineer
AI Inference Platform Performance Engineer

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Sr. AI Inference Platform Engineer
Sr. AI Inference Platform Engineer

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
Senior AI Inference Platform Engineer — Benchmarking
Senior AI Inference Platform Engineer — Benchmarking

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
System Performance Engineer - AI/ML, Platform Architecture
System Performance Engineer - AI/ML, Platform Architecture

Socket.dev • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms
Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms

Socket.dev • Santa Clara (CA)

On-site
USD 180,000 - 240,000
SW Optimization Engineer AI/ML
SW Optimization Engineer AI/ML

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock programs
Software Engineer (AML), AI & Data Platforms (AiDP)
Software Engineer (AML), AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX)

On-site
USD 120,000 - 150,000
AI Data Platform Engineer
AI Data Platform Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Applied AI Engineer - iCloud Data
Applied AI Engineer - iCloud Data

Socket.dev • Cupertino (CA)

On-site
USD 210,000 - 320,000