Multimodal Foundation Models Engineer for Robotics

The Bot Company

San Francisco (CA)

On-site

USD 180,000 - 290,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The Bot Company in San Francisco is building a helpful robot for every home. We are a lean team of engineers, designers, and operators, shipping to hundreds of millions of users.

Our culture promotes radical ownership, rapid iteration, and direct responsibility across the stack. We are seeking an experienced ML engineer to work on multimodal foundation models that natively reason across text, image, video, and kinematics.

Qualifications

  • Proficient coding in Python, C++, or Rust.
  • Experience training and deploying large-scale multimodal models.
  • Deep understanding of LLM-style pretraining and RL at scale.
  • Able to manage and optimize large GPU-based experiments.

Responsibilities

  • Build native multimodal policies where vision, language, and other modalities share a unified representation.
  • Improve cross-modal reasoning beyond simple associations; ground visual physics in kinematic constraints.
  • Own the training loop end-to-end: design, run, debug, and iterate large-scale training experiments.

Skills

Python
C++
Rust
Production MLLM Experience
Pretraining & RL Mastery
Infrastructure Fluency

Job description

The Bot Company in San Francisco is building a helpful robot for every home. We are a lean team of engineers, designers, and operators, shipping to hundreds of millions of users.

Our culture promotes radical ownership, rapid iteration, and direct responsibility across the stack. We are seeking an experienced ML engineer to work on multimodal foundation models that natively reason across text, image, video, and kinematics.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

Maven Ventures • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Lead Multimodal ML for Robotic Systems (Edge AI)
Lead Multimodal ML for Robotic Systems (Edge AI)

Maven Ventures • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Robotics ML Engineer: Multimodal Perception & Action
Robotics ML Engineer: Multimodal Perception & Action

Human Computer Lab • San Francisco (CA)

On-site
USD 120,000 - 150,000
Autonomy ML Engineer: Foundation Models & Pipelines
Autonomy ML Engineer: Foundation Models & Pipelines

Humble Robotics • San Francisco (CA)

On-site
USD 100,000 - 300,000
Multimodal Robotics Research Engineer
Multimodal Robotics Research Engineer

Human Archive • San Francisco (CA)

On-site
USD 100,000 - 140,000
Multimodal Perception & Authentication ML Engineer
Multimodal Perception & Authentication ML Engineer

OpenAI • United States

Hybrid
USD 140,000 - 210,000
Relocation assistance
Staff Robotics ML/Data Engineer — Multimodal Pipelines
Staff Robotics ML/Data Engineer — Multimodal Pipelines

Persona AI Inc • Houston (TX)

On-site
USD 140,000 - 180,000
Competitive compensation
Bonus program
Medical benefits (99% employer covered
+4
Agentic AI/ML Engineer - Multimodal Robotics
Agentic AI/ML Engineer - Multimodal Robotics

Field AI • Irvine (CA)

On-site
USD 100,000 - 150,000
Work with real robots and hardware
Collaborate with industry experts
Generous salary range based on experience