ML Ops Engineer: Reproducible Runs & Scalable Tools

Veeda

California (MO)

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is building next-gen multimodal world models for Physical AI, a small, fast-moving team of engineers and researchers. We are seeking an engineer to own training runs, experiment tracking, and CI governance across multi-node jobs.

You will define run definitions, track artifacts across code and data, and ensure reproducible pipelines end-to-end. Strong Python/CI/CD skills and a passion for rigorous evaluation are essential.

Qualifications

  • Bachelor's degree or equivalent hands-on CS/engineering background.
  • Strong Python and software engineering skills with CI/CD experience.
  • Experience operating multi-node training jobs and managing runs.
  • Experience building reproducible end-to-end pipelines.
  • Strong evaluation methodology and statistical rigor.

Responsibilities

  • Define how a run is defined, launched, resumed, and killed with Hydra configs and pinned digests.
  • Bind every checkpoint to code, config, dataset version, and container digest in tracking systems (Weights & Biases/MLflow).
  • Own retention, GC policy, DCP resharding, and promotion from raw checkpoint to evaluated artifact.
  • Gate checkpoints on seeded rollout and policy tests in CI pipelines (Slurm/Argo workflows).
  • Monitor live runs, report throughput, data loader issues, and GPU-hour usage.

Skills

Python
CI/CD
Distributed training
Reproducible pipelines
Evaluation rigor

Education

Bachelor's degree in CS or related

Tools

Hydra
OmegaConf
Weights & Biases
MLflow
Slurm
Argo Workflows
GitHub Actions
Buildkite

Job description

Veeda AI is building next-gen multimodal world models for Physical AI, a small, fast-moving team of engineers and researchers. We are seeking an engineer to own training runs, experiment tracking, and CI governance across multi-node jobs.

You will define run definitions, track artifacts across code and data, and ensure reproducible pipelines end-to-end. Strong Python/CI/CD skills and a passion for rigorous evaluation are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda • California (MO)

On-site
USD 120,000 - 190,000
Staff Data Engineer, Multimodal AI Pipelines
Staff Data Engineer, Multimodal AI Pipelines

Veeda • California (MO)

On-site
USD 130,000 - 180,000
AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Staff Engineer, World Models for Embodied AI
Staff Engineer, World Models for Embodied AI

Veeda AI • Seattle (WA)

On-site
USD 150,000 - 230,000
Technical Staff Engineer - World Models & Multimodal AI
Technical Staff Engineer - World Models & Multimodal AI

Veeda • California (MO)

On-site
USD 130,000 - 200,000
AI Infrastructure Engineer: HPC GPU Clusters
AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Engineer, Physical AI Simulation & World Models
Staff Engineer, Physical AI Simulation & World Models

Veeda Innovation • California (MO), Northern (KY)

Hybrid
USD 140,000 - 190,000
Member of Technical Staff - Data
Member of Technical Staff - Data

Veeda • California (MO)

On-site
USD 130,000 - 180,000
Senior MLOps Engineer: Build Scalable AI for Logistics
Senior MLOps Engineer: Build Scalable AI for Logistics

Veho • United States

On-site
USD 150,000 - 210,000
ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000