Get more replies from employers
Send a job-specific resume in minutes.
NomadicML is seeking a Machine Learning Engineer in San Francisco to advance foundation-model research and production engineering. You will train and fine-tune large-scale Vision-Language Models to reason about motion in real-world video, building multi-modal architectures that process video, language, and sensor data.
You will work closely with founders to deploy robust APIs and SDKs for enterprise use, publish research at top venues, and drive end-to-end experiments from data collection to
About NomadicMLAmericans drive over 5 trillion miles a year, more than 500 billion of them recorded. Buried in that footage is the next frontier of machine intelligence. At NomadicML, we’re building the platform that unlocks it.Our Vision-Language Models (VLMs) act as the new “hydraulic mining” for video, transforming raw footage into structured intelligence that powers real-world autonomy and robotics. We partner with industry leaders across self-driving, robotics, and industrial automation to mine insights from petabytes of data that were once unusable.NomadicML was founded by Mustafa Bal and Varun Krishnan, who met at Harvard University while studying Computer Science.Mustafa is a core contributor to ONNX Runtime and DeepSpeed with deep expertise in distributed systems and large-scale model training infrastructureVarun is an INFORMS Wagner Prize Finalist for his research in large-scale driver navigation AI models and one of the top chess players in the US.Our team has built mission-critical AI systems at Snowflake, Lyft, Microsoft, Amazon, and IBM Research, holds top-tier publications in VLMS and AI at conferences like CVPR, and moves with the speed and clarity of a startup obsessed with impact.
We’re seeking a Machine Learning Engineer who thrives at the frontier of foundation-model research and production engineering. You’ll help define how machines learn from motion: training and fine-tuning large-scale Vision-Language Models to reason about complex, real-world video.
Your work will involve building multi-modal architectures that perceive, localize, and describe motion events (turns, lane changes, interactions, anomalies) across millions of frames, and turning those breakthroughs into robust APIs and SDKs used by enterprise customers.
You’ll work directly with the founders to: