ML Engineer - Infrastructure

Flexion

Zürich

On-site

CHF 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Enhanced pension plan
Relocation & permit sponsorship
Enhanced holiday & paid leave perks
Central Zürich office with top-tier R

Job summary

Flexion in Zurich is seeking a senior ML engineer to own GPU compute platforms, design and bring up scalable clusters, and collaborate with AI engineers to accelerate training. You will shape multi-cloud compute strategies, optimize performance, and contribute to new tools that improve throughput on large models.

You will also mentor junior engineers and contribute to reliability improvements across the platform.

Qualifications

  • Hands-on experience with training or inference of large models on distributed multi-node GPU hardware.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure) or large-scale on-premises GPU infrastructure.
  • Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.

Responsibilities

  • Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently.
  • Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.
  • Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them.
  • Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.
  • Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability.

Skills

Python
Distributed training
Collaboration with AI engineers
Cloud experience
Profilers familiarity

Education

Degree in Computer Science, Electrical Engineering or Software Engineering

Tools

PyTorch
Slurm
Kubernetes/KubeRay
AWS/GCP/Azure

Job description

About Flexion

At Flexion, we are building the autonomy stack for humanoid robots. Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids. We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms. In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning. Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.

The role

We are looking for an experienced ML engineer to join Flexion’s experienced infrastructure team and take ownership of Flexion’s GPU compute platforms. This is a senior, on-site role with significant scope.

At Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters. You will own the design, bring-up, operation and optimization of performant clusters. You will work with AI engineers to help them optimize their training speed and hardware utilization. You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently. This will put you at the heart of Flexion’s AI development and allow you to directly impact the execution of our ambitious roadmap. You will closely collaborate with the company’s leadership, engineers of the infrastructure team and AI engineers across the company.

Key responsibilities
  • Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems.
  • Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.
  • Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them.
  • Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.
  • Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability.
  • Degree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experience.
  • Hands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware. This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etc.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.
  • Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.

Nice-to-haves

  • Familiarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight).
  • Experience with high-performance or parallel file systems (e.g., Lustre).
  • Experience provisioning compute on multiple cloud providers.
  • Experience with infrastructure-as-code and configuration management (Terraform, Ansible).
  • Competitive Compensation
  • Joining a leading robotics team & exposure to never-done-before research
  • Energetic, collaborative culture with a bias for action and regular community events

Zurich

  • Enhanced pension plan
  • Relocation & permit sponsorship
  • Enhanced holiday & paid leave perks
  • Central Zürich office with top-tier robotics testing facilities and infrastructure

San Franciso

  • 401(k) with company contributions
  • Health, dental & vision coverage with the flexibility to choose your own plan
  • Open PTO policy & paid company holidays
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer - Infrastructure
ML Engineer - Infrastructure

Flexion Robotics • Zürich

On-site
CHF 140,000 - 190,000
Enhanced pension plan
Relocation sponsorship
Enhanced holiday & paid leave perks
+1
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Flexion Robotics • Zürich

Hybrid
CHF 140,000 - 210,000
Enhanced pension plan
Central Zürich office with robotics 시설
Health, dental & vision coverage
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Flexion • Zürich

On-site
CHF 120,000 - 180,000
Competitive compensation
Relocation sponsorship
Pension plan
+2
Forward Deployed Engineer - Robotics
Forward Deployed Engineer - Robotics

Flexion • Zürich

Hybrid
CHF 120,000 - 180,000
Competitive compensation
Enhanced pension plan
Relocation sponsorship
+1
Product Manager - AI Platform
Product Manager - AI Platform

Flexion • Zürich

On-site
CHF 140,000 - 190,000
Competitive pay
Pension plan
Paid leave
+4
Product Manager - AI Platform
Product Manager - AI Platform

Flexion Robotics • Zürich

On-site
CHF 120,000 - 190,000
Competitive compensation
Enhanced pension plan
Enhanced holiday & paid leave perks
+2
AI Research Engineer - Human Data
AI Research Engineer - Human Data

Flexion • Zürich

On-site
CHF 120,000 - 180,000
Competitive compensation
Enhanced pension plan
Relocation & permit sponsorship
+2
AI Research Engineer - Human Data
AI Research Engineer - Human Data

Flexion Robotics • Zürich

On-site
CHF 120,000 - 180,000
Relocation sponsorship
Pension plan
Paid leave
+1
Product Manager - Humanoid Platform
Product Manager - Humanoid Platform

Flexion • Zürich

On-site
CHF 120,000 - 190,000
Competitive compensation
Relocation & permit sponsorship
Exposure to Europe’s leading robotics
+1
Senior ML Engineer, GPU Compute & Infrastructure - On-site
Senior ML Engineer, GPU Compute & Infrastructure - On-site

Flexion • Zürich

On-site
CHF 150,000 - 210,000
Enhanced pension plan
Relocation & permit sponsorship
Enhanced holiday & paid leave perks
+1