Senior / Staff ML Ops Engineer

RiseMe

United States

On-site

USD 184,000 - 272,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity awards
Health and Wellness benefits
Unlimited Vacation
Work from Home support
Catered meals and snacks
Team events and off-site activities

Job summary

Waabi is hiring a versatile software/infrastructure engineer to scale training infrastructure on Kubernetes, design developer-facing CLIs and SDKs, and improve end-to-end model training and deployment workflows. You will work on GPU scheduling, autoscaling, and orchestration tools across a multi-node setup.

The role emphasizes building durable tooling, experiment tracking, and a robust model registry, with a strong focus on collaboration and autonomous work in a fast-paced AI company.

Qualifications

  • 5+ years of software or infrastructure engineering, including ML or data-intensive production systems.
  • Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm, networking fundamentals.
  • Excellent Python, with API/CLI design experience and AWS depth (S3, IAM, GPU, VPC).
  • Experience in Distributed training in PyTorch (DDP/FSDP) and experiment tracking tools.
  • Fluency with containers, CI/CD, and large monorepos.
  • Ability to influence without authority and drive adoption among senior engineers.

Responsibilities

  • Build and evolve training infrastructure on Kubernetes with GPU scheduling and autoscaling.
  • Shape the developer surface — CLIs, SDKs, templates to simplify use for teams.
  • Improve iteration speed: shorten time to first training and reduce latency.
  • evangelize tooling and frameworks with prototypes and migration paths.
  • Strengthen data and artifact layers: dataset versioning and high-throughput loading of data.
  • Turn one-off Python into durable tooling with well-documented libraries and services.

Skills

Kubernetes expert
Python
API design
CLI tooling
AWS / cloud
Distributed training
CI/CD
Collaboration
Communication
Autonomy in ambiguous territory

Tools

Terraform
Pulumi
Argo Workflows
Ray
Kubeflow
Flyte
Slurm

Job description

Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.

With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai

You will..
  • Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable.
  • Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible.
  • Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down.
  • Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost.
  • Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal sensor data, so jobs saturate GPUs instead of waiting on I/O.
  • Turn one-off Python into durable tooling — tested, documented, observable libraries, CLIs, and services with sane defaults, and deletions where they're overdue.
  • Make experiments legible, with the teams who live in them: experiment hygiene, dashboards researchers trust, a real model registry, and lineage from dataset to checkpoint to simulation result.
  • Ship CI/CD for models alongside autonomy and simulation, so a model change is validated the same way a code change is.
  • Build observability across the ML stack — utilization, throughput, failure modes, queue times, cost per experiment. When a job fails at 3am on node 47, the researcher should find out why without you.
  • Treat docs, onboarding, and support as product surface — golden-path guides, a new researcher productive on day two,office hours that turn repeat questions into shipped fixes.
  • Drive adoption, not just availability. Prototype with real users, watch them work, iterate. A tool nobody adopts didedn't ship.
  • Make the platform boringly reliable — fewer failures, faster recovery, and none of the manual steps that quietly cost a team days.
  • Build guardrails that don't feel like walls, with Security, IT, and Infrastructure: access controls, data handling, and cost governance that hold up in an IP-sensitive environment while staying self-serve.
Qualifications:
  • 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems.
  • Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and the ability to debug a cluster under load rather than restart it.
  • Excellent Python, and a track record of designing APIs and CLIs other people enjoy using.
    Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar).
  • Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling — from the perspective of someone who made them pleasant for others to use.
  • Fluency with containers, CI/CD, and modern build systems, including large monorepos.
  • The ability to influence without authority: evaluate a framework on its merits, pilot it credibly, and persuade skeptical senior engineers to change how they work.
  • A collaborative default — you'd rather co‑owned a system than draw a boundary.
  • User empathy: you'd rather fix the third‑most-interesting problem blocking ten people than the most interesting one blocking nobody.
  • Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory.
  • Passionate about self‑driving technologies … and about what we…??
Bonus/nice to have:
  • Internal developer platform, research platform, or DevEx work — with a story about a tool whose adoption you grew from zero.
  • Large-scale distributed GPU training: hundreds to thousands of accelerators, NCCL, high-performance cluster networking, collective communication tuning.
  • High-throughput loading of LiDAR or camera data, and formats such as Parquet or WebDataset.
  • Workflow and scheduling systems — Argo Workflows, Ray, Flyte, Kubeflow, or Slurm.
  • Build-system depth (Bazel or similar), including remote caching in a monorepo.
  • Simulation infrastructure or large-scale batch evaluation pipelines.
  • Background in ML, robotics, or autonomous systems infrastructure.
  • Security- and IP-sensitive production environments.
  • Open‑source contributions to ML infrastructure or developer tools.

The US yearly salary range for this role is: $184,000 - $272,000 USD in addition to competitive perks & benefits. Waabi (US) Inc.’s yearly salary ranges are determined based on several factors in accordance with the Company’s compensation practices. The salary base range is reflective of the minimum and maximum target for new hire salaries for the position across all US locations. Note: The Company provides additional compensation for employees in this role, including equity incentive awards and an annual performance bonus.

Perks/Benefits:
  • Competitive compensation and equity awards.
  • Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time employees only).
  • Unlimited Vacation.
  • Flexible hours and Work from Home support.
  • Daily drinks, snacks and catered meals (when in office).
  • Regularly scheduled team building activities and social events both on-site, off-site & virtually.
  • As we grow, this list continues to evolve!

Waabi is a technology start-up building technologies to transform the way the world moves. Join our talented team to be a part of the future and to make an impact!

Waabi is an equal opportunity employer. We celebrate diversity and are committed to creating a supportive, inclusive, and accessible workplace for all our employees. We seek applicants of all backgrounds and identities, across race, color, ethnicity, national origin or ancestry, age, citizenship, religion, sex, sexual orientation, gender identity or expression, military or veteran status, marital status, pregnancy or parental status, caregiver status, disability, or any other characteristic protected by law. We make workplace accommodations for qualified individuals with disabilities as required by applicable law. If reasonable accommodation is needed to participate in the job application or interview process please let our recruiting team know.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior / Staff ML Ops Engineer
Senior / Staff ML Ops Engineer

Lever, Inc. • Dallas (TX)

On-site
USD 184,000 - 272,000
Competitive compensation and equity
Health and Wellness benefits (medical,
Unlimited Vacation
+4
Senior / Staff Software Engineer, ML Datasets & Data Pipelines
Senior / Staff Software Engineer, ML Datasets & Data Pipelines

ProducePay • United States

On-site
USD 148,000 - 260,000
Competitive compensation and equity
Health and Wellness benefits (Medical,
Unlimited Vacation
+3
Software Engineer, Labelling, Data & Automation
Software Engineer, Labelling, Data & Automation

Waabi • San Francisco (CA)

On-site
USD 127,000 - 225,000
Competitive compensation
Health and Wellness benefits
Unlimited Vacation
+3
Software Engineer, Labelling, Data & Automation
Software Engineer, Labelling, Data & Automation

ProducePay • United States

Hybrid
USD 127,000 - 225,000
Competitive compensation
Equity awards
Health benefits
+4
Senior HW Systems Engineer
Senior HW Systems Engineer

Waabi • San Francisco (CA)

On-site
USD 151,000 - 230,000
Health and wellness benefits
Paid parental, medical and family care
Generous PTO
+4
Senior / Staff Research Engineer, Simulation Assets & Content Systems
Senior / Staff Research Engineer, Simulation Assets & Content Systems

ProducePay • United States

Hybrid
USD 155,000 - 269,000
Health and wellness benefits
Unlimited vacation
Flexible hours and work from home
+2
Senior / Staff Applied Scientist
Senior / Staff Applied Scientist

ProducePay • United States

Hybrid
USD 146,000 - 280,000
Competitive compensation
Equity awards
Health benefits
+5
Senior / Staff Software Engineer, AI Tooling
Senior / Staff Software Engineer, AI Tooling

Waabi • United States

Hybrid
USD 165,000 - 265,000
Competitive compensation and equity
Health, Dental and Vision coverage
Unlimited vacation
+2
Staff Systems Engineer, Embedded Safety & Architecture
Staff Systems Engineer, Embedded Safety & Architecture

Socket.dev • San Francisco (CA)

Hybrid
USD 160,000 - 260,000
Competitive compensation and equity
Health and Wellness benefits
Unlimited Vacation
+3
Senior IT Administrator
Senior IT Administrator

ProducePay • Dallas (TX)

On-site
USD 117,000 - 172,000
Health benefits
Unlimited vacation
Flexible hours
+2