Machine Learning: Multimodal Foundation Models

The Bot Company

San Francisco (CA)

On-site

USD 180,000 - 290,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

The Bot Company in San Francisco is building a helpful robot for every home. We are a lean team of engineers, designers, and operators, shipping to hundreds of millions of users.

Our culture promotes radical ownership, rapid iteration, and direct responsibility across the stack. We are seeking an experienced ML engineer to work on multimodal foundation models that natively reason across text, image, video, and kinematics.

Qualifications

  • Proficient coding in Python, C++, or Rust.
  • Experience training and deploying large-scale multimodal models.
  • Deep understanding of LLM-style pretraining and RL at scale.
  • Able to manage and optimize large GPU-based experiments.

Responsibilities

  • Build native multimodal policies where vision, language, and other modalities share a unified representation.
  • Improve cross-modal reasoning beyond simple associations; ground visual physics in kinematic constraints.
  • Own the training loop end-to-end: design, run, debug, and iterate large-scale training experiments.

Skills

Python
C++
Rust
Production MLLM Experience
Pretraining & RL Mastery
Infrastructure Fluency

Job description

The Bot Company

We're building a helpful robot for every home.

We're a small team of engineers, designers, and operators based in San Francisco. Our team comes from Tesla, Cruise, OpenAI, Google, Pixar, and many other great companies. In the past we've shipped to hundreds of millions of users and know what it takes to build amazing products and experiences.

Our team is deliberately lean to promote rapid decision making and do away with bureaucracy and hierarchy. Everyone is an IC and is empowered with massive scope, radical ownership, and direct responsibility. We work across the stack with a culture built for rapid iteration and fast execution.

What we look for in all candidates

All roles at The Bot Company demand extreme sharpness and the ability to move fast in high-intensity environments. Throughout the process, we expect candidates to demonstrate:

  • Exceptional mental acuity: you think quickly, learn instantly, and reason across unfamiliar domains.

  • Engineering curiosity: you naturally dig into how systems work, even outside your specialty.

  • High performance mindset: you move fast, handle ambiguity, and excel when the environment is demanding.

Machine Learning: Multimodal Foundation Models

We are building unified foundation models that natively reason across text, image, video, and kinematics to drive intelligent robotic policies.

You will work on large multi-modal networks and own the entire stack from data to training and deploying models.

What You'll Do
  • Build Native Multimodal Policies: Develop architectures where vision, language, and more modalities share a unified representation.

  • Improve Cross-Modal Reasoning: Research and implement methods to ensure the model doesn't just "associate" modalities but actually reasons through them (e.g., grounding visual physics in kinematic constraints).

  • Own the Training Loop End-to-End: Design, run, debug, and iterate on large-scale training experiments; diagnosing failure modes, improving data mixtures, and tightening evaluation to drive measurable gains.

  • Ship and Iterate on Real Systems: Integrate models into real robotic stacks, build on robot code to deploy your models, and optimize performance for edge inference.

Requirements
  • Very strong coding skills in Python, C++, or Rust.

  • Production MLLM Experience: Track record of training and deploying large-scale multimodal models.

  • Pretraining & RL Mastery: Deep intuition for LLM-style pretraining, post-training, and Reinforcement Learning at scale.

  • Infrastructure Fluency: Comfortable managing and optimizing large-scale experiments on massive GPU clusters.

Why Join

You’ll work with a small, elite team on challenges that require speed, intelligence, and deep engineering instinct. If you enjoy understanding systems at all levels, move fast, and think even faster, you’ll thrive here.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

Maven Ventures • San Francisco (CA)

On-site
USD 150,000 - 230,000
Lead Multimodal ML for Robotic Systems (Edge AI)
Lead Multimodal ML for Robotic Systems (Edge AI)

Maven Ventures • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning: World Models
Machine Learning: World Models

The Bot Company • San Francisco (CA)

On-site
USD 150,000 - 230,000
Multimodal Foundation Models Engineer for Robotics
Multimodal Foundation Models Engineer for Robotics

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Senior ML Engineer — Whole-Body Control & Simulation
Senior ML Engineer — Whole-Body Control & Simulation

Maven Ventures • San Francisco (CA)

On-site
USD 200,000 - 350,000
Comprehensive benefits package including medical, dental, and vision coverage
Access to a 401(k) plan
Equity through the company's discretionary equity program
Founding ML Engineer
Founding ML Engineer

a16z-speedrun • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer
Machine Learning Engineer

OpenMind • San Francisco (CA)

On-site
USD 130,000 - 210,000
ML Infra - Data Infrastructure
ML Infra - Data Infrastructure

Maven Ventures • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff (MTS) - Multimodal Foundation Models
Member of Technical Staff (MTS) - Multimodal Foundation Models

Deeproute Dot A I • Iowa (LA)

On-site
USD 150,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000