Member of Technical Staff, Vision / Language

XDOF

San Mateo (CA)

Hybrid

USD 140,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Health benefits
Flexible work
Collaborative culture
Impactful work

Job summary

XDOF, in San Mateo, is seeking a Research Engineer/Scientist to lead data-driven efforts at the intersection of vision-language models and robot learning. You will build systems that convert raw egocentric and teleoperation video into high-signal training data for VLMs and increasingly contribute to the models themselves.

You'll drive research into what makes robot data useful—discovering metadata and annotations to unlock cross-embodiment transfer, curriculum generation, and world models.

Qualifications

  • MS or PhD in Computer Science, Robotics, Machine Learning or related field from a top program
  • 3–7 years of research or applied research experience in vision-language models, video understanding, robot learning, or generative modeling
  • Deep fluency in PyTorch and experience with distributed training, mixed precision, and large batch workflows
  • Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or related areas
  • Strong engineering fundamentals and ability to design clean systems, not just run experiments

Responsibilities

  • Design and implement vision-language pipelines for egocentric and teleoperation video: structured captioning, temporal grounding, action-conditioned scene understanding, and semantic annotation at scale
  • Develop and evaluate representations bridging vision, language, and robot action
  • Build and improve data curation systems that assess quality, diversity, and coverage of large-scale robot demonstration datasets
  • Work hands-on with bimanual and high-DoF manipulation data, including real teleoperation footage and sim-generated rollouts
  • Collaborate with partner labs to define data requirements and optimize data quality for downstream policy performance
  • Stay current on research frontier and translate insights into production systems

Skills

Vision-language models
PyTorch
Robot learning
Video understanding
Distributed training
Data pipelines
Research experience

Education

MS or PhD in Computer Science, Robotics, ML

Tools

Large-scale training infra

Job description

About XDOF

Frontier labs are racing to build general‑purpose robots, and the bottleneck isn't compute. It's data. At XDOF, we’re building the foundation behind the foundation models: the data collection systems, annotation pipelines, exabyte‑scale data infrastructure, and software toolchain that enable our partners to push the field forward.

We’re hiring a Research Engineer / Scientist to help lead technical efforts at the intersection of vision‑language models and robot learning. You will build systems that turn raw egocentric and teleoperation video into high‑signal training data for VLA models, and increasingly, contribute to the models themselves.

Beyond pipelines, you will drive research into what makes robot data useful: discovering new metadata (contact events, affordance labels, implicit reward signals, dynamics priors from video) that unlock capabilities current approaches miss. You’ll explore how structured annotations can improve cross‑embodiment transfer, automatic curriculum generation, and world models that predict what actually matters for manipulation. The data layer isn’t downstream of the research. It is the research.

What You’ll Do
  • Design and implement vision‑language pipelines for egocentric and teleoperation video: structured captioning, temporal grounding, action‑conditioned scene understanding, and semantic annotation at scale

  • Develop and evaluate representations that bridge visual perception, language, and low‑level robot action — spanning VLAs, video prediction, and world models

  • Build and improve data curation systems that assess quality, diversity, and coverage of large‑scale robot demonstration datasets

  • Work hands‑on with bimanual and high‑DoF manipulation data, including real teleoperation footage and sim‑generated rollouts

  • Collaborate directly with partner labs to define data requirements and close the loop between data quality and downstream policy performance

  • Stay current on the research frontier (VLAs, video foundation models, flow matching, DiT architectures, egocentric pretraining) and translate insights into production systems

Required:
  • MS or PhD in Computer Science, Robotics, Machine Learning, or a related field from a top‑tier program

  • 3–7 years of research or applied research experience (industry or academic) in one or more of: vision‑language models, video understanding, robot learning, or generative modeling

  • Deep fluency in PyTorch; working knowledge of large‑scale training infrastructure (distributed training, mixed precision, large batch workflows)

  • Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or a closely related area

  • Strong engineering fundamentals — you can design clean systems, not just run experiments

Benefits
  • Competitive compensation and equity

  • Comprehensive health and wellness benefits

  • Flexible work arrangements

  • Collaborative and fast‑paced work environment

  • Opportunity to shape the future of robotics and AI alongside an ambitious, values‑driven team

Level: Mid Level to Senior Research Scientist (L4–L5 equivalent) Location: San Mateo

Note: Junior candidates will still be considered

If you’re excited to help build the infrastructure powering tomorrow’s intelligent machines, we’d love to hear from you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Vision / Language
Member of Technical Staff, Vision / Language

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment
Member of Technical Staff, Vision/Language
Member of Technical Staff, Vision/Language

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Staff Research Engineer - Vision-Language Robotics
Staff Research Engineer - Vision-Language Robotics

XDOF • San Mateo (CA)

Hybrid
USD 140,000 - 200,000
Equity
Health benefits
Flexible work
+2
Vision-Language Robotics Research Engineer
Vision-Language Robotics Research Engineer

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment
Vision-Language Robotics Data Systems Engineer
Vision-Language Robotics Data Systems Engineer

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Senior Research Scientist
Senior Research Scientist

Sereact • Massachusetts

On-site
USD 150,000 - 210,000
Health Insurance
401(k) with company match
20 days PTO
+5
Research Engineer/ Scientist Redwood City, CA Fulltime
Research Engineer/ Scientist Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Robotics
Member of Technical Staff, Robotics

xdof • San Mateo (CA)

Hybrid
USD 120,000 - 170,000
Senior Perception Engineer
Senior Perception Engineer

xdof, Inc. • San Mateo (CA)

On-site
USD 150,000 - 210,000
Direct impact on robotics data
Collaborative environment
Proprietary hardware platforms