Member of Technical Staff - ML Operations

Veeda AI

Zürich

On-site

CHF 140,000 - 190,000

Full time

38 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We are a small, fast-moving team tackling challenging problems at the intersection of AI, robotics, and embodied intelligence.

The role focuses on experiment tooling, inference fleet orchestration, CI model evaluation, and scalable data pipelines, with ownership from design to production observability.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in CS, engineering, or a related technical field.
  • Proficiency in Python or Bash and one compiled language (Java, Rust, or Go).
  • Experience in shipping production-quality developer tools.
  • Proficient in CI/CD automations (pipelines, runners, deployment).
  • Experience building reproducible end-to-end pipelines and knowledge of bit-reproducibility.
  • Understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns.

Responsibilities

  • Experiment Lifecycle Tracking and Tooling: design, build and deploy tools for run definition, launch, resume and kill; manage experiment databases for version tracking and reproducibility.
  • Inference Fleet Orchestration: design and operate the serving control plane; develop cache-aware admission, routing, batching and scheduling policies.
  • Model Evaluation in CI: design and operate automatic model checkpoint evaluation systems.
  • Data Pipeline Operations: design and operate high-performance distributed backend services for large-scale data processing.
  • Visualization Platforms: design observability and dataset visualization platforms with interactive data viz, progress tracking and search.
  • End-to-End Ownership: lead projects through full software lifecycle including CI/CD and production observability.

Skills

Python
Bash
Java
Rust
Go
CI/CD automation
Event-driven architecture
Concurrency models
Distributed serving

Education

Bachelor's degree or equivalent

Tools

Argo Workflows
Flyte
Ray

Job description

About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Experiment Lifecycle Tracking and Tooling: Design, build and deploy tools that own how a run is defined, launched, resumed, and killed. Build and operate experiment databases for code version tracking, data version tracking, reproducibility, and checkpoint ancestry.
  • Inference Fleet Orchestration: Design, build, and operate the serving control plane that accepts large volumes of concurrent client requests and assigns them across inference clusters. Develop cache-aware admission, routing, batching, and scheduling policies that improve cache locality, balance workload, protect tail latency, and keep the fleet highly utilized and reliable. Partner with ML Performance on model runtime, kernel, and per-worker throughput optimization.
  • Model Evaluation in CI: Design, build, and operate automatic model checkpoint evaluation systems on seeded rollout and policy-success suites, run per-change and nightly.
  • Data Pipeline Operations: Design, build, and operate high-performance, fault-tolerant, distributed backend services and event-driven systems for our large scale data processing pipeline.
  • Visualization Platforms: Design, build, and deploy experiment observability and dataset visualization platform(s) that provides interactive data visualization, progress tracking, search, and comparison.
  • End-to-End Ownership: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability.
Requirements
  • Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
  • Proficiency in at least one scripting (Python or Bash) and one compiled (Java, Rust, or Go) languages.
  • Experience in shipping production-quality developer tools.
  • Proficient in CICD automations (pipelines, runners, deployment)
  • Experience in building reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.
  • One of the following:
    • Proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns.
    • Experience building or operating large-scale inference control planes or distributed serving infrastructure, with hands-on work in traffic management, admission control, request scheduling, routing, batching, or cache-aware load balancing; able to reason clearly about cache locality, queueing, tail latency, availability, and fleet utilization.
    • Experience working with multi-node workloads and building around slurm based scheduling systems.
    • Experience working on in-production model evaluation frameworks, in particular regression testing of large multimodal models.
NICE TO HAVE
  • You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.
  • You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.
  • Full stack development experience with web-based front-end.
  • You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.
  • You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.
  • You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Zürich

On-site
CHF 120,000 - 160,000
Member of Technical Staff - Simulation
Member of Technical Staff - Simulation

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Internship
Internship

Veeda AI • Zürich

Hybrid
CHF 17,000 - 28,000
Staff ML Engineer: Systems & Experimentation
Staff ML Engineer: Systems & Experimentation

Veeda AI • Zürich

On-site
CHF 140,000 - 190,000
Member of Technical Staff – AI Inference platform
Member of Technical Staff – AI Inference platform

Lyceum • Zürich

On-site
CHF 130,000 - 170,000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Switzerland

On-site
CHF 100,000 - 140,000
Research Engineer
Research Engineer

Nomagic • Zürich

On-site
CHF 90,000 - 120,000
Relocation package
Flexible working hours
English-speaking environment