ML Engineer: Multimodal Vision-Language

Spur

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Stock options

Job summary

NomadicML is seeking a Machine Learning Engineer in San Francisco to advance foundation-model research and production engineering. You will train and fine-tune large-scale Vision-Language Models to reason about motion in real-world video, building multi-modal architectures that process video, language, and sensor data.

You will work closely with founders to deploy robust APIs and SDKs for enterprise use, publish research at top venues, and drive end-to-end experiments from data collection to

Qualifications

  • Proficient in Python and PyTorch with large-scale ML workflows.
  • Experience with foundation models, VLMs, or multi-modal learning.
  • Ability to run end-to-end experiments and iterate quickly.
  • Experience training or fine-tuning models on video or sensor data.
  • Understanding of retrieval systems, embeddings, and GPU optimization.

Responsibilities

  • Train and evaluate VLMs for motion understanding in autonomous-driving datasets.
  • Design and scale GPU-accelerated pipelines for training, fine-tuning, and inference on multi-modal data.
  • Build evaluation frameworks for spatiotemporal reasoning and narrative consistency.
  • Develop curation loops that generate and refine datasets using models.
  • Publish high-impact research while shipping customer-facing features.

Skills

Python
PyTorch
ML workflows
Experimentation

Tools

DeepSpeed
Hugging Face

Job description

NomadicML is seeking a Machine Learning Engineer in San Francisco to advance foundation-model research and production engineering. You will train and fine-tune large-scale Vision-Language Models to reason about motion in real-world video, building multi-modal architectures that process video, language, and sensor data.

You will work closely with founders to deploy robust APIs and SDKs for enterprise use, publish research at top venues, and drive end-to-end experiments from data collection to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer – Vision-Language for Motion & Autonomy
ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
ML Engineer: Vision-Language & Motion for Autonomy
ML Engineer: Vision-Language & Motion for Autonomy

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Edge Vision & Multimodal AI Engineer
Edge Vision & Multimodal AI Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 180,000
Junior AI/ML Engineer (Vision + Multimodal) On-site (SF)
Junior AI/ML Engineer (Vision + Multimodal) On-site (SF)

Blueprints AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
ML Engineer: Multimodal Data & Production Systems
ML Engineer: Multimodal Data & Production Systems

Sieve, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Multimodal AI Research Scientist
Multimodal AI Research Scientist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Multimodal AI Research Engineer (Vision & Language)
Multimodal AI Research Engineer (Vision & Language)

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Senior ML Engineer: Lead Multimodal AI Systems
Senior ML Engineer: Lead Multimodal AI Systems

Twenty80 LLC • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Multimodal ML Engineer - Vision & Foundation Models
Senior Multimodal ML Engineer - Vision & Foundation Models

Waymo • California (MO)

Hybrid
USD 213,000 - 263,000