Research, Vision Expertise

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. in San Francisco, CA is seeking a researcher to advance multimodal AI at scale, blending vision and language. You will design experiments, build datasets, and contribute to training large-scale models that ground concepts in the physical world.

The role blends research and engineering, requiring writing high-performance code, reading technical reports, and publishing results. A PhD or equivalent industry experience in ML or related fields is preferred.

Qualifications

  • Design, run, and analyze experiments with empirical rigor.
  • Strong understanding of ML fundamentals and distributed compute.
  • Proficiency in Python and at least one DL framework; scalable code.
  • Bachelor’s degree or higher in CS/ML/Physics/Math with strong grounding.
  • Excellent written communication.

Responsibilities

  • Own research projects on training and performance analysis of multimodal AI models.
  • Curate and build large-scale datasets and benchmarks.
  • Collaborate with data infra engineers, researchers, and product teams.
  • Publish research and share code, datasets, and insights with the community.

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

Thinking Machines builds multimodal-first.We’re looking for new team members to advance the science of visual perception and multimodal learning. We think about how vision and language interact at scale. We design architectures that fuse pixels and text, build datasets and evaluation methods that test real-world comprehension, and develop representations that let models ground abstract concepts in the physical world. Our goal is to create multimodal systems that support seamless integration into real-world environments.

You’ll work at the intersection of visual understanding, multimodal reasoning, and large-scale model training. You’ll help develop the architectures, data, and evaluation tools that teach AI to see, understand, and collaborate. The best candidate is curious about multimodal interfaces, has experience running large scale experiments and is comfortable contributing to complex engineering systems. While we are looking for a person with expertise in multimodality, Thinking Machines Lab operates in a unified fashion and expects new hires to work across modalities as one team.

This role blends fundamental research and practical engineering, as we do not distinguish between the two roles internally. You will be expected to write high-performance code and read technical reports. It’s an excellent fit for someone who enjoys both deep theoretical exploration and hands‑on experimentation, and who wants to shape the foundations of how AI learns.

What You’ll Do
  • Own research projects on training and performance analysis of multimodal AI models.

  • Curate and build large-scale datasets and evaluation benchmarks to advance vision capabilities.

  • Work with our data infrastructure engineers, pretraining researchers and engineers, and product team to create frontier multimodal models and the products that leverage them.

  • Publish and present research that moves the entire community forward. Share code, datasets, and insights that accelerate progress across industry and academia.

Skills and Qualifications

Minimum qualifications:

  • Ability to design, run, and analyze experiments thoughtfully, with demonstrated research judgment and empirical rigor.

  • Understanding of machine learning fundamentals, large-scale training, and distributed compute environments.

  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales.

  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.

  • Clarity in communication, an ability to explain complex technical concepts in writing.

Preferred qualifications — we encourage you to apply even if you don’t meet all preferred qualifications, but at least some:

  • Research or engineering contributions in visual reasoning, spatial understanding, or multimodal architecture design.

  • Experience developing evaluation frameworks for multimodal tasks.

  • Publications or open-source contributions in vision‑language modeling, video understanding, or multimodal AI.

  • A strong grasp of probability, statistics, and ML fundamentals. You can look at experimental data and distinguish between real effects, noise, and bugs.

  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.

Logistics
  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Pre-Training Data
Research, Pre-Training Data

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Research Tools
Software Engineer, Research Tools

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Member of Technical Staff - Multimodal Understanding
Member of Technical Staff - Multimodal Understanding

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical coverage
401(k) retirement plan
+2
Multimodal AI Research Scientist
Multimodal AI Research Scientist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research Engineer, Visual Knowledge Work
Research Engineer, Visual Knowledge Work

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Generous vacation
Parental leave
Flexible working hours
Member of Technical Staff, Multimodal Vision San Jose
Member of Technical Staff, Multimodal Vision San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Multimodal AI Researcher
Multimodal AI Researcher

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000