ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI

San Francisco (CA)

On-site

USD 170,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

NomadicML, based in San Francisco, seeks a Machine Learning Engineer to advance motion understanding with vision‑language models and production‑grade ML pipelines. You will train, evaluate, and deploy multi‑modal systems that reason about complex video events and interactions across large datasets.

You will collaborate with founders to scale GPU‑accelerated workflows, build robust APIs, and publish research while delivering customer‑facing features.

Qualifications

  • Proficiency in Python, PyTorch, and large‑scale ML workflows.
  • Research experience in foundation models, VLMs, or multi‑modal learning (publications/patents a plus).
  • Ability to iterate quickly and autonomously, running experiments end‑to‑end.
  • Experience training or fine‑tuning models on video or sensor data.
  • Understanding of retrieval systems, embeddings, and GPU optimization.

Responsibilities

  • Train and evaluate VLMs specialized for motion understanding in autonomous‑driving and robotics datasets.
  • Design and scale GPU‑accelerated pipelines for training, fine‑tuning, and inference on multi‑modal data (video + language + sensor metadata).
  • Build agentic evaluation frameworks that benchmark spatiotemporal reasoning, localization accuracy, and narrative consistency.
  • Develop and productionize curation loops that use our own models to generate and refine datasets (“AI training AI”).
  • Publish high‑impact research (e.g., NeurIPS, CVPR) while shipping features that customers use immediately.

Skills

Python
PyTorch
large‑scale ML workflows
foundation models
VLMs / multi‑modal learning
video data
embeddings / retrieval
GPU optimization
experimental iteration

Tools

DeepSpeed
Hugging Face
Ray / Kubeflow / MLflow

Job description

NomadicML, based in San Francisco, seeks a Machine Learning Engineer to advance motion understanding with vision‑language models and production‑grade ML pipelines. You will train, evaluate, and deploy multi‑modal systems that reason about complex video events and interactions across large datasets.

You will collaborate with founders to scale GPU‑accelerated workflows, build robust APIs, and publish research while delivering customer‑facing features.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer: Vision-Language & Motion for Autonomy
ML Engineer: Vision-Language & Motion for Autonomy

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Engineer - Vision & Multimodal (Production)
Senior ML Engineer - Vision & Multimodal (Production)

Clearview AI • United States

On-site
USD 150,000 - 200,000
Medical, Dental, Vision
STD and LTD Plans
Senior Vision-Language ML Engineer, Autonomous Driving
Senior Vision-Language ML Engineer, Autonomous Driving

XPENG • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Massive real-world data
Industry-scale compute
Top-tier researchers and engineers
+2
Staff ML Engineer: Vision-Language-Action for Autonomous Driving
Staff ML Engineer: Vision-Language-Action for Autonomous Driving

XPENG • Santa Clara (CA)

On-site
USD 215,280 - 364,320
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Access to massive real-world data and industry-scale compute
Senior ML Infra Engineer — Train & Deploy Vision Models
Senior ML Infra Engineer — Train & Deploy Vision Models

Voxel Sa • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Equity via equity plan
Health benefits
Unlimited PTO
+2
Vision-Language ML Architect & Inventor of Novel VLMs
Vision-Language ML Architect & Inventor of Novel VLMs

Arcade • Presidio (TX)

On-site
USD 180,000 - 240,000
Daily catered lunch
Company events
Equal opportunity employer
Senior ML/CV Engineer – Autonomy Perception & Vision
Senior ML/CV Engineer – Autonomy Perception & Vision

Nuro • Mountain View (CA)

On-site
USD 184,000 - 276,000