ML Engineer: Vision-Language & Motion for Autonomy

Praxis, Inc.

San Francisco (CA)

On-site

USD 150,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NomadicML is seeking a Machine Learning Engineer to push the frontier of foundation-model research and production engineering. You will help define how machines learn from motion by training and fine-tuning large-scale Vision-Language Models on motion-rich video data.

You will build multi-modal architectures that perceive, localize, and describe motion events across millions of frames, turning breakthroughs into robust APIs and SDKs for enterprise customers.

Qualifications

  • Strong proficiency in Python, PyTorch, and large-scale ML workflows.
  • Research experience in foundation models, VLMs, or multi-modal learning (publications/patents a plus).
  • Ability to iterate quickly and autonomously, running experiments end-to-end.
  • Experience training or fine-tuning models on video or sensor data.
  • Understanding of retrieval systems, embeddings, and GPU optimization.

Responsibilities

  • Train and evaluate VLMs specialized for motion understanding in autonomous-driving and robotics datasets.
  • Design and scale GPU-accelerated pipelines for training, fine-tuning, and inference on multi-modal data (video + language + sensor metadata).
  • Build agentic evaluation frameworks that benchmark spatiotemporal reasoning, localization accuracy, and narrative consistency.
  • Develop and productionize curation loops that use our own models to generate and refine datasets (“AI training AI”).
  • Publish high-impact research while shipping features that customers use immediately.

Skills

Python
PyTorch
Foundation models
Vision-Language Models
Multi-modal learning
Experimentation

Tools

Hugging Face
DeepSpeed
Ray
Kubeflow
MLflow

Job description

NomadicML is seeking a Machine Learning Engineer to push the frontier of foundation-model research and production engineering. You will help define how machines learn from motion by training and fine-tuning large-scale Vision-Language Models on motion-rich video data.

You will build multi-modal architectures that perceive, localize, and describe motion events across millions of frames, turning breakthroughs into robust APIs and SDKs for enterprise customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer – Vision-Language for Motion & Autonomy
ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff ML Engineer: Vision-Language Foundations & Data-Flywheel
Staff ML Engineer: Vision-Language Foundations & Data-Flywheel

Waymo • Mountain View (CA)

On-site
USD 251,000 - 310,000
Discretionary annual bonus
Equity incentive plan
Generous Company benefits program
Staff ML Engineer: Vision-Language-Action for Autonomous Driving
Staff ML Engineer: Vision-Language-Action for Autonomous Driving

XPENG • Santa Clara (CA)

On-site
USD 215,280 - 364,320
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Access to massive real-world data and industry-scale compute
Staff ML Engineer – Vision-Language-Action for Autonomous Driving
Staff ML Engineer – Vision-Language-Action for Autonomous Driving

Xpengmotors • Santa Clara (CA)

On-site
USD 215,000 - 365,000
Competitive compensation package
Meals provided
Collaborative work environment
Senior ML Engineer – Vision & Multimodal Models
Senior ML Engineer – Vision & Multimodal Models

clearview • Washington

On-site
USD 150,000 - 210,000
Production ML Engineer — Vision, LLMs & RAG Systems
Production ML Engineer — Vision, LLMs & RAG Systems

Electronic Transaction Consultants Corporation • Frisco (TX)

On-site
USD 120,000 - 190,000
Paid time off
Health and Dental plans
Retirement plans
+2
Senior ML Engineer: Computer Vision & Vision-Language
Senior ML Engineer: Computer Vision & Vision-Language

Uber • Seattle (WA)

On-site
USD 202,000 - 224,000
Bonus program
Equity awards
401(k) plan