AIML - Senior ML/RL Training Infrastructure Engineer, AFM

Apple

Zürich

Vor Ort

CHF 120.000 - 150.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Flexible working hours
Health and wellness benefits
Diversity and inclusion programs

Zusammenfassung

Apple is looking for a skilled engineer to join its Core Foundation Models team in Zürich. The role involves designing and optimizing large-scale ML training infrastructure, particularly focusing on reinforcement learning with TPU-based systems using JAX. Candidates should have a PhD or MSc in Computer Science, solid experience with GPUs/TPUs, and a strong proficiency in PyTorch or JAX. This position is ideal for those passionate about high-performance computing and distributed systems.

Qualifikationen

  • Hands-on experience designing and maintaining large-scale ML training infrastructure.
  • Solid understanding of distributed systems concepts.
  • Experience in developing or optimizing training loops and ML frameworks.

Aufgaben

  • Design, build, and scale systems for large-scale reinforcement learning.
  • Focus on TPU-based training with JAX.
  • Ensure efficiency, reliability, and observability of training pipelines.

Kenntnisse

Designing ML training infrastructure
Optimizing distributed systems
Proficiency with PyTorch or JAX
High-performance computing
Experience with GPUs/TPUs
Software engineering in Python

Ausbildung

PhD or MSc in Computer Science or related field

Jobbeschreibung

Summary

Ready to transform how billions of people interact with technology? Apple’s Core Foundation Models team is driving the intelligence that powers experiences across billions of devices worldwide—and we’re looking for exceptional talent to join us! Join our Europe-based applied ML team building the next generation of large‑scale ML and RL training infrastructure for Apple’s foundation models. We develop high-performance, distributed systems that power cutting‑edge foundation model research on a massive scale. We are seeking an engineer who is passionate about designing, optimizing, and scaling the infrastructure that enables state-of-the-art machine learning and reinforcement learning workloads.

As a senior member of the team, you will work closely with researchers and systems engineers to build robust training frameworks, accelerate experimentation, and push the boundaries of performance and efficiency. You will collaborate with teams across Apple’s engineering hubs—including New York, Seattle, and Cupertino—to advance the tooling and systems that make large-scale model training possible. If you thrive at the intersection of distributed systems, ML frameworks, and high-performance computing, this is the role for you.

Description

As a core member of our ML infrastructure team, you will design, build, and scale the systems that enable large-scale reinforcement learning for Apple’s foundation models. You will focus on TPU-based training with JAX, developing robust, high-performance RL pipelines that support distributed actor/learner architectures, efficient experience replay, and large-scale environment execution.

In this role, you will work across the full stack of RL training systems—from low-level performance tuning and compiler optimization to cluster-level orchestration and resource management. You will ensure that training pipelines are efficient, reliable, reproducible, and observable, enabling research teams to iterate quickly and explore more complex RL environments and models.

Your work will directly impact the scalability, throughput, and stability of RL experiments, helping to unlock new capabilities in agentic reasoning, decision-making, and policy learning for Apple’s foundation models. This position is ideal for engineers who enjoy distributed systems, high-performance ML frameworks, and building the infrastructure that makes large-scale RL research possible.

Minimum Qualifications
  • PhD or MSc in Computer Science, Computer Engineering or a closely related field.
  • Hands‑on experience designing, building, or maintaining large‑scale ML training infrastructure.
  • Strong proficiency with PyTorch or JAX and experience running training workloads on GPUs/TPUs.
  • Solid understanding of distributed systems concepts (parallelism strategies, fault tolerance, synchronization).
Preferred Qualifications
  • Practical experience developing or optimizing training loops, RL pipelines, or large-scale model-training frameworks.
  • Strong software engineering skills in Python, with emphasis on reliability, debuggability, and high-performance execution.
  • Deep experience with PyTorch/JAX internals, XLA, debugging and performance profiling on GPU/TPU architectures.
  • Expertise in distributed RL training patterns, including actor/learner architectures, experience replay, and parallel environment execution.
  • Experience building training services, orchestration tools, or automated pipelines for large-scale experiments.
  • Proven success diagnosing bottlenecks in large-scale ML jobs (I/O, input pipelines, kernel performance, memory, compilation).
  • Familiarity with RL-specific infrastructure requirements (e.g., actor/learner architectures, experience replay systems, large-scale environment execution).
  • Strong software engineering practices: code quality, design reviews, testing, observability, CI/CD.
  • Experience working with cloud-scale clusters or specialized accelerators (TPU v5/v6, GPU, custom hardware).
  • Contributions to ML frameworks, distributed training libraries, or high-performance computing systems.
  • Excellent communication and collaboration skills for working with research and engineering partners.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Apple is committed to treating all applicants fairly and equally. We will work with applicants to make any reasonable accommodations.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AIML Researcher - Foundation Model, Post-Training
AIML Researcher - Foundation Model, Post-Training

Apple Inc. • Zürich

Vor Ort
CHF 140.000 - 210.000
AIML - Machine Learning Researcher, MLR
AIML - Machine Learning Researcher, MLR

Apple • Lausanne

Vor Ort
CHF 180.000 - 240.000
AIML - Machine Learning Research, Multimodal Foundation Models
AIML - Machine Learning Research, Multimodal Foundation Models

Apple • Zürich

Vor Ort
CHF 100.000 - 130.000
Inclusive work environment
Focus on accessibility
Research Scientist - LLM Efficiency
Research Scientist - LLM Efficiency

Apple • Zürich

Vor Ort
CHF 180.000 - 240.000
Research Scientist - LLM Efficiency
Research Scientist - LLM Efficiency

Apple Inc. • Zürich

Vor Ort
CHF 140.000 - 210.000
Senior RL Training Infrastructure Engineer (TPU/JAX)
Senior RL Training Infrastructure Engineer (TPU/JAX)

Apple • Zürich

Vor Ort
CHF 120.000 - 150.000
Applied Machine Learning Engineer - Security
Applied Machine Learning Engineer - Security

Apple • Zürich

Vor Ort
CHF 120.000 - 150.000
Applied Machine Learning Engineer - Security
Applied Machine Learning Engineer - Security

StudySmarter • Zürich

Vor Ort
CHF 43.000 - 72.000
Research Scientist/ Engineer – Special Projects
Research Scientist/ Engineer – Special Projects

Apple Switzerland AG • Zürich

Vor Ort
CHF 140.000 - 210.000
Research Scientist/ Engineer - Special Projects
Research Scientist/ Engineer - Special Projects

Apple • Zürich

Vor Ort
CHF 120.000 - 180.000