Member of Technical Staff - ML Operations

Veeda AI

Toronto

On-site

CAD 120,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Veeda AI in Toronto is seeking a hands-on ML/software engineer to own end-to-end tooling for lifecycle management, experiment tracking, and checkpoint lineage on multi-node training jobs.

You will ensure reproducible pipelines, robust CI/CD practices (GitHub Actions, Buildkite), and responsive incident handling, with a strong Python background and proven software craftsmanship. Candidates should have experience shipping internal tooling that others adopt.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
  • Strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite), and you have shipped internal tooling that other engineers chose to keep using.
  • You have operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride.
  • You have built reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.
  • You are rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can tell a genuine regression from a flaky harness.

Responsibilities

  • Run Lifecycle & Launch Tooling: Define, launch, resume, and kill runs from typed configs to pinned container digests.
  • Experiment Tracking & Provenance: Bind every checkpoint to code commit, config hash, dataset version, and container digest in ML tools.
  • Checkpoint Registry & Lineage: Own retention, GC policy, and format conversions for artifacts used by simulation/robotics.
  • Evaluation in CI: Gate checkpoints on seeded rollout and policy-success suites in CI workflows.
  • Goodput & Incident Response: Be on call for live runs and report goodput against allocated GPU-hours.

Skills

Python
CI/CD
Multi-node training
Reproducible pipelines
Experiment tracking
Docker / container tooling
Telemetry / monitoring

Education

Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or related field

Tools

GitHub Actions
Buildkite
Weights & Biases
MLflow
Argo Workflows
Flyte
Ray
Prometheus
Grafana

Job description

ABOUT US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

RESPONSIBILITIES
  • Run Lifecycle & Launch Tooling: Own how a run is defined, launched, resumed, and killed, from typed configs (Hydra, OmegaConf) to pinned container digests to a relaunch that takes one command.

  • Experiment Tracking & Provenance: Bind every checkpoint to its code commit, config hash, dataset version, and container digest in Weights & Biases or MLflow, so an old run rebuilds from its manifest, not from memory.

  • Checkpoint Registry & Lineage: Own retention and garbage-collection policy, PyTorch DCP resharding and format conversion, and promotion from raw checkpoint to evaluated artifact that simulation and robotics can safely build on.

  • Evaluation in CI: Gate each checkpoint on seeded rollout and policy-success suites, run per-change and nightly as Slurm arrays under Argo Workflows, with confidence intervals wide enough to separate regression from eval noise.

  • Goodput & Incident Response: Carry the pager for live runs (loss spikes, throughput cliffs, data loader stalls), and report goodput against allocated GPU-hours as the number capacity decisions actually run on.

REQUIREMENTS
  • You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.

  • You have strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite), and you have shipped internal tooling that other engineers chose to keep using.

  • You have operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride.

  • You have built reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.

  • You are rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can tell a genuine regression from a flaky harness.

NICE TO HAVE
  • You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.

  • You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.

  • You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.

  • You are fluent with Prometheus, Grafana, and OpenTelemetry, and instrument a training job before its first outage.

  • You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.

  • You have written a postmortem that changed how a team ran jobs, not just what it logged.

  • You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda AI • Toronto

On-site
CAD 110,000 - 150,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
AI Engineer
AI Engineer

Valsoft Corporation • Canada

On-site
CAD 100,000 - 140,000
ML Ops Engineer - Reproducible AI Pipelines
ML Ops Engineer - Reproducible AI Pipelines

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
Full-Stack AI Developer
Full-Stack AI Developer

CoFoMo Inc. • Montreal (administrative region)

Hybrid
CAD 95,000 - 140,000
Senior Full-Stack Software Engineer
Senior Full-Stack Software Engineer

Jobtailor • Toronto

On-site
CAD 130,000 - 185,000
Full Stack Software Developer - Engineered Arts
Full Stack Software Developer - Engineered Arts

Opportunities with AppDirect's Advisor Partners (Recruitment as a Service) • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
ML Engineer New Vancouver
ML Engineer New Vancouver

Tigera, Inc. • Vancouver

On-site
CAD 160,000 - 180,000
Health benefits
Vision benefits
Dental benefits
+1
Senior Software Developer, Data & MLOps
Senior Software Developer, Data & MLOps

OSEDEA • Montreal (administrative region)

Hybrid
CAD 85,000 - 115,000
Competitive Salary
Pension plan contribution (RRSP)
Flexible work hours
+5
AI/ML Ops Engineer
AI/ML Ops Engineer

Blackpoint Cyber • Canada

Hybrid
CAD 110,000 - 145,000