Machine Learning: Multimodal Foundation Models

Pantera Capital

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Bot Company in San Francisco is seeking an ML engineer to work on multimodal foundation models and own the stack from data to deployment. You will develop systems that enable unified representations across text, image, video, and kinematics for robotic policies.

You will build scalable architectures, improve cross-modal reasoning, and ship real-world systems with edge inference in a small, fast-paced team.

Qualifications

  • Strong coding skills in Python, C++, or Rust.
  • Experience training and deploying large-scale multimodal models.
  • Experience with massive GPU clusters and distributed training.

Responsibilities

  • Build Native Multimodal Policies: Develop architectures where vision, language, and more modalities share a unified representation.
  • Improve Cross-Modal Reasoning: Research methods to ensure the model reasons across modalities, not just associates them.
  • Own the Training Loop End-to-End: Design, run, debug, and iterate on large-scale training experiments.
  • Ship and Iterate on Real Systems: Integrate models into robotic stacks and optimize for edge inference.

Skills

Python
C++
Rust
MLLM Experience

Job description

The Bot Company

We're building a helpful robot for every home.

We're a small team of engineers, designers, and operators based in San Francisco. Our team comes from Tesla, Cruise, OpenAI, Google, Pixar, and many other great companies. In the past we've shipped to hundreds of millions of users and know what it takes to build amazing products and experiences.

Our team is deliberately lean to promote rapid decision making and do away with bureaucracy and hierarchy. Everyone is an IC and is empowered with massive scope, radical ownership, and direct responsibility. We work across the stack with a culture built for rapid iteration and fast execution.

What we look for in all candidates

All roles at The Bot Company demand extreme sharpness and the ability to move fast in high-intensity environments. Throughout the process, we expect candidates to demonstrate:

  • Exceptional mental acuity: you think quickly, learn instantly, and reason across unfamiliar domains.

  • Engineering curiosity: you naturally dig into how systems work, even outside your specialty.

  • High performance mindset: you move fast, handle ambiguity, and excel when the environment is demanding.

Machine Learning: Multimodal Foundation Models

We are building unified foundation models that natively reason across text, image, video, and kinematics to drive intelligent robotic policies.

You will work on large multi-modal networks and own the entire stack from data to training and deploying models.

What You’ll Do
  • Build Native Multimodal Policies: Develop architectures where vision, language, and more modalities share a unified representation.

  • Improve Cross-Modal Reasoning: Research and implement methods to ensure the model doesn't just "associate" modalities but actually reasons through them (e.g., grounding visual physics in kinematic constraints).

  • Own the Training Loop End-to-End: Design, run, debug, and iterate on large-scale training experiments; diagnosing failure modes, improving data mixtures, and tightening evaluation to drive measurable gains.

  • Ship and Iterate on Real Systems: Integrate models into real robotic stacks, build on robot code to deploy your models, and optimize performance for edge inference.

Requirements
  • Very strong coding skills in Python, C++, or Rust.

  • Production MLLM Experience: Track record of training and deploying large-scale multimodal models.

  • Pretraining & RL Mastery: Deep intuition for LLM-style pretraining, post-training, and Reinforcement Learning at scale.

  • Infrastructure Fluency: Comfortable managing and optimizing large-scale experiments on massive GPU clusters.

Why Join

You’ll work with a small, elite team on challenges that require speed, intelligence, and deep engineering instinct. If you enjoy understanding systems at all levels, move fast, and think even faster, you’ll thrive here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Machine Learning: World Models
Machine Learning: World Models

The Bot Company • San Francisco (CA)

On-site
USD 150,000 - 230,000
Multimodal Foundation Models Engineer for Robotics
Multimodal Foundation Models Engineer for Robotics

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Lead Multimodal ML for Robotic Systems (Edge AI)
Lead Multimodal ML for Robotic Systems (Edge AI)

Pantera Capital • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning: Whole-Body Control
Machine Learning: Whole-Body Control

Pantera Capital • San Francisco (CA)

On-site
USD 200,000 - 350,000
Comprehensive benefits package including medical, dental, and vision coverage
Access to a 401(k) plan
Equity through the company's discretionary equity program
Research Vision Expertise
Research Vision Expertise

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research, Vision Expertise
Research, Vision Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff, Machine Learning - NomadicML
Member of Technical Staff, Machine Learning - NomadicML

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Member of Technical Staff — ML Research, Multimodal
Member of Technical Staff — ML Research, Multimodal

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000