Member of Technical Staff - ML Operations

Veeda AI

California

On-site

USD 120,000 - 190,000

Full time

16 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. A small, fast-moving team tackles the most challenging problems at the intersection of AI, robotics, and embodied intelligence.

The role focuses on designing and operating the experiment lifecycle, inference fleet, model evaluation in CI, data pipelines, visualization platforms, and end-to-end ownership of software projects for scalable AI systems.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
  • Proficiency in at least one scripting (Python or Bash) and one compiled (Java, Rust, or Go) languages.
  • Experience in shipping production-quality developer tools.
  • Proficient in CICD automations (pipelines, runners, deployment).
  • Experience in building reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.
  • One of the following: proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns; experience building large-scale inference control planes or distributed serving infrastructure; experience with multi-node workloads and Slurm; or in-production model evaluation frameworks.

Responsibilities

  • Experiment Lifecycle Tracking and Tooling: Design, build and deploy tools that own how a run is defined, launched, resumed, and killed.
  • Inference Fleet Orchestration: Design, build, and operate the serving control plane for large volumes of concurrent client requests across inference clusters; optimize throughput and tail latency.
  • Model Evaluation in CI: Design, build, and operate automatic model checkpoint evaluation systems on seeded rollout and policy-success suites.
  • Data Pipeline Operations: Design, build, and operate high-performance, fault-tolerant, distributed backend services for large-scale data processing.
  • Visualization Platforms: Design, build, and deploy experiment observability and dataset visualization platforms for interactive data visualization, progress tracking, search, and comparison.
  • End-to-End Ownership: Lead projects through the complete software lifecycle, including specs, CI/CD, on-call support, and production observability.

Skills

Python
Bash
Java
Rust
Go
CI/CD automations
end-to-end pipelines
event-driven architecture
concurrency models

Education

Bachelor's degree

Tools

Argo Workflows
Flyte
Ray
Slurm scheduling

Job description

About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Experiment Lifecycle Tracking and Tooling: Design, build and deploy tools that own how a run is defined, launched, resumed, and killed. Build and operate experiment databases for code version tracking, data version tracking, reproducibility, and checkpoint ancestry.
  • Inference Fleet Orchestration: Design, build, and operate the serving control plane that accepts large volumes of concurrent client requests and assigns them across inference clusters. Develop cache-aware admission, routing, batching, and scheduling policies that improve cache locality, balance workload, protect tail latency, and keep the fleet highly utilized and reliable. Partner with ML Performance on model runtime, kernel, and per-worker throughput optimization.
  • Model Evaluation in CI: Design, build, and operate automatic model checkpoint evaluation systems on seeded rollout and policy-success suites, run per-change and nightly.
  • Data Pipeline Operations: Design, build, and operate high-performance, fault-tolerant, distributed backend services and event-driven systems for our large scale data processing pipeline.
  • Visualization Platforms: Design, build, and deploy experiment observability and dataset visualization platform(s) that provides interactive data visualization, progress tracking, search, and comparison.
  • End-to-End Ownership: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability.
Requirements
  • Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
  • Proficiency in at least one scripting (Python or Bash) and one compiled (Java, Rust, or Go) languages.
  • Experience in shipping production-quality developer tools.
  • Proficient in CICD automations (pipelines, runners, deployment)
  • Experience in building reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.
  • One of the following:
    • Proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns.
    • Experience building or operating large-scale inference control planes or distributed serving infrastructure, with hands-on work in traffic management, admission control, request scheduling, routing, batching, or cache-aware load balancing; able to reason clearly about cache locality, queueing, tail latency, availability, and fleet utilization.
    • Experience working with multi-node workloads and building around slurm based scheduling systems.
    • Experience working on in-production model evaluation frameworks, in particular regression testing of large multimodal models.
NICE TO HAVE
  • You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.
  • You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.
  • Full stack development experience with web-based front-end.
  • You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.
  • You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.
  • You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda AI • California

On-site
USD 140,000 - 210,000
Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda AI • Seattle (WA)

On-site
USD 140,000 - 190,000
Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda • California (MO)

On-site
USD 140,000 - 180,000
Member of Technical Staff - Data
Member of Technical Staff - Data

Veeda Innovation • California (MO)

Hybrid
USD 130,000 - 180,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • California

On-site
USD 150,000 - 210,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Seattle (WA)

On-site
USD 140,000 - 220,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • California

On-site
USD 180,000 - 240,000
Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda AI • Seattle (WA)

On-site
USD 150,000 - 210,000