ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI

San Francisco (CA)

On-site

USD 170,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NomadicML, based in San Francisco, seeks a Machine Learning Engineer to advance motion understanding with vision‑language models and production‑grade ML pipelines. You will train, evaluate, and deploy multi‑modal systems that reason about complex video events and interactions across large datasets.

You will collaborate with founders to scale GPU‑accelerated workflows, build robust APIs, and publish research while delivering customer‑facing features.

Qualifications

  • Proficiency in Python, PyTorch, and large‑scale ML workflows.
  • Research experience in foundation models, VLMs, or multi‑modal learning (publications/patents a plus).
  • Ability to iterate quickly and autonomously, running experiments end‑to‑end.
  • Experience training or fine‑tuning models on video or sensor data.
  • Understanding of retrieval systems, embeddings, and GPU optimization.

Responsibilities

  • Train and evaluate VLMs specialized for motion understanding in autonomous‑driving and robotics datasets.
  • Design and scale GPU‑accelerated pipelines for training, fine‑tuning, and inference on multi‑modal data (video + language + sensor metadata).
  • Build agentic evaluation frameworks that benchmark spatiotemporal reasoning, localization accuracy, and narrative consistency.
  • Develop and productionize curation loops that use our own models to generate and refine datasets (“AI training AI”).
  • Publish high‑impact research (e.g., NeurIPS, CVPR) while shipping features that customers use immediately.

Skills

Python
PyTorch
large‑scale ML workflows
foundation models
VLMs / multi‑modal learning
video data
embeddings / retrieval
GPU optimization
experimental iteration

Tools

DeepSpeed
Hugging Face
Ray / Kubeflow / MLflow

Job description

NomadicML, based in San Francisco, seeks a Machine Learning Engineer to advance motion understanding with vision‑language models and production‑grade ML pipelines. You will train, evaluate, and deploy multi‑modal systems that reason about complex video events and interactions across large datasets.

You will collaborate with founders to scale GPU‑accelerated workflows, build robust APIs, and publish research while delivering customer‑facing features.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer: Vision-Language & Motion for Autonomy
ML Engineer: Vision-Language & Motion for Autonomy

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Pear VC • Austin (TX), California (MO)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
ML Engineer: Vision‑Language Models for Motion
ML Engineer: Vision‑Language Models for Motion

Pear VC • Austin (TX), California (MO)

On-site
USD 120,000 - 150,000
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Engineer, Computer Vision & Graphics for VisionOS
ML Engineer, Computer Vision & Graphics for VisionOS

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
Production ML Engineer — Vision, LLMs & RAG Systems
Production ML Engineer — Vision, LLMs & RAG Systems

Electronic Transaction Consultants Corporation • Frisco (TX)

On-site
USD 120,000 - 190,000
Paid time off
Health and Dental plans
Retirement plans
+2
Senior ML Engineer – Vision & Multimodal Models
Senior ML Engineer – Vision & Multimodal Models

clearview • Washington

On-site
USD 150,000 - 210,000
Staff ML Engineer – Vision-Language-Action for Autonomous Driving
Staff ML Engineer – Vision-Language-Action for Autonomous Driving

Xpengmotors • Santa Clara (CA)

On-site
USD 215,000 - 365,000
Competitive compensation package
Meals provided
Collaborative work environment