Site Reliability Engineer - ML, Apple Ads

Socket.dev

New York (NY)

On-site

USD 140,000 - 210,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Apple Ads seeks a ML Platform Infrastructure Engineer to scale ML training, inference, and serving workloads. You’ll own health, performance, and automation of large-scale AWS-based infrastructure powering Apple Ads applications, building platform solutions rather than just pipelines.

Ideal candidates have 3+ years in SRE/ML Ops, strong coding in Python/Java/Rust/Go, and IaC with Terraform, plus expertise in Linux and observability concepts.

Qualifications

  • 3+ years in internet-facing backend production systems, SRE or ML Operations on large-scale cloud infra.
  • Experience with AWS-managed infrastructure and scalable services.
  • Familiarity with ML lifecycle tools such as NVIDIA Triton, AnyScale Ray, Apache Airflow.
  • Strong programming skills in Python, Java, Rust or Go.
  • Hands-on Linux and deep knowledge of its internals.
  • IaC experience, especially Terraform.
  • Solid grounding in SRE: monitoring, alerting, observability, incident response, error budgets, SLAs/SLOs.

Responsibilities

  • Own health, performance, and scalability of ML training, inference, serving workloads and platform tooling.
  • Build automation to reduce toil, improve resilience, and speed up delivery.
  • Develop platform solutions beyond basic CI/CD pipelines.
  • Collaborate across engineering, infrastructure, and data science teams.

Skills

AWS
ML Ops
Python
Java
Rust
Go
Terraform
Linux
Monitoring
Incident response

Tools

NVIDIA Triton
Apache Airflow
AnyScale Ray
Kubernetes

Job description

At Apple, we focus deeply on our customers’ experience. Apple Ads brings this same approach to advertising, helping people find exactly what they’re looking for and helping advertisers grow their businesses. Our technology powers ads and sponsorships across Apple Services, including the App Store, Apple News, and MLS Season Pass. Everything we do is designed for trust, connection, and impact: We respect user privacy, integrate advertising thoughtfully into the experience, and deliver value for advertisers of all sizes—from small app developers to big, global brands. Because when advertising is done right, it benefits everyone. The Site Reliability Engineering team within Apple Ads ensures the reliability, performance, and availability of ML Platform and Services at scale. The team partners closely with Ads engineering, data science and ML platform teams to enable product delivery through design, configuration, and automation of machine learning infrastructure powering Apple Ads applications. We are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform — enabling fast, reliable, and scalable operations across AWS-based environments supporting transactional and analytical workloads.

Description

As a site reliability engineer in Apple Ads focused on machine learning, you will own the health, performance, and scalability of large scale infrastructure powering ML training, inference, serving workloads and associated platform tooling. Your focus will be on building automation that eliminates manual processes, improves platform resilience, and enables teams to move faster with confidence. This is not a DevOps-only or CI/CD-focused role. We are looking for engineers who build platform solutions, not just configure pipelines.

Minimum Qualifications

3+ years of experience in internet-facing backend production systems, SRE or ML Operations focused roles on large scale distributed cloud infrastructure Proven expertise with AWS-managed infrastructure Familiarity with ML lifecycle and associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc. Strong programming skills in at least one of: Python, Java, Rust, Go or similar languages Hands-on experience with Linux systems and deep knowledge of its internals. Demonstrated experience with Infrastructure as Code, especially Terraform. Strong foundation in SRE concepts: Monitoring, alerting, observability, Incident response and root cause analysis, Error budgets, SLAs/SLOs, and system reliability

Preferred Qualifications

Built tools or services that automate platform operations, reduce toil, or improve cost efficiency. Experience managing Kubernetes clusters at scale in production environments. Hands-on experience troubleshooting distributed systems under real-world load. Clear communication skills and comfort collaborating across engineering, infrastructure, and product teams. AWS certifications or broad experience across multiple AWS services is a plus. Understanding of modern GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia ), associated driversUnderstanding of high-performance fabrics and network architecture, power, and thermal limits

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - ML, Apple Ads
Site Reliability Engineer - ML, Apple Ads

Apple Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 225,000
Relocation assistance
Site Reliability Engineer - ML, Apple Ads
Site Reliability Engineer - ML, Apple Ads

Apple Inc. • New York

On-site
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Relocation assistance
+2
ML Platform SRE: Scale, Reliability & Automation
ML Platform SRE: Scale, Reliability & Automation

Apple Inc. • New York

On-site
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Relocation assistance
+2
ML Platform SRE — Scale Apple Ads Infra
ML Platform SRE — Scale Apple Ads Infra

Apple Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 225,000
Relocation assistance
ML Platform SRE: Scale-Ready Infra for Ads
ML Platform SRE: Scale-Ready Infra for Ads

Socket.dev • New York (NY)

On-site
USD 140,000 - 210,000
Site Reliability Engineer, Apple Ads
Site Reliability Engineer, Apple Ads

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 190,000
Staff ML Engineer - Ads ML Infrastructure
Staff ML Engineer - Ads ML Infrastructure

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
SRE Manager, ML Operations
SRE Manager, ML Operations

Apple • New York (NY)

On-site
USD 238,000 - 356,000
Stock options
Relocation assistance
Comprehensive benefits
Site Reliability Engineer – Ad Platforms – job_id_000043 Job ID- 77
Site Reliability Engineer – Ad Platforms – job_id_000043 Job ID- 77

Apple • Pasadena (CA)

On-site
USD 100,000 - 150,000
SRE Manager, ML Operations
SRE Manager, ML Operations

Apple Inc. • New York (NY)

On-site
USD 228,000 - 343,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
+1