Machine Learning Engineer, Multimodal Perception and Authentication

OpenAI

United States

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Relocation assistance

Job summary

OpenAI's Future of Computing Research team in San Francisco is seeking a machine learning engineer to help shape how AI systems understand the physical world and the people in it.

You will focus on multimodal perception and authentication, bringing together signals from cameras, microphones, and other sensors, and working with hardware, firmware, software, and product teams to turn research into real-world systems.

Qualifications

  • Strong background in computer vision, audio or speech machine learning, multimodal learning, or sensing.
  • Experience developing specialized ML models, larger multimodal models, or both.
  • Experience bringing research ideas into practical systems, prototypes, or products.
  • Know how to design experiments, build evaluations, and investigate model behavior.

Responsibilities

  • Research and develop multimodal perception and authentication methods across visual, audio, and other sensing signals.
  • Explore how specialized perception models and larger multimodal models can work together.
  • Design data, training, and evaluation approaches that improve performance in real-world conditions.
  • Study model behavior, robustness, and failure modes across sensing, data, and deployment environments.
  • Integrate and validate new capabilities in real-time or resource-constrained systems.
  • Work with hardware, firmware, software, and product teams to turn research into working systems.

Skills

Computer vision
Audio or speech ML
Multimodal learning
Sensing

Job description

About the Team

The Future of Computing Research team is an applied research team within OpenAI’s Consumer Devices group. We study how AI systems perceive people and their surroundings, and we turn that research into capabilities for future products. Our work spans machine learning, sensing, and hardware, with a focus on building systems that work beyond controlled environments.

About the Role

We’re looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and authentication, bringing together signals from cameras, microphones, and other sensors. You’ll work with specialized perception models and larger multimodal models, and partner with hardware, firmware, software, and product teams to bring new research into real-world systems.

This role is based in San Francisco. We work in the office three days per week and offer relocation assistance.

In this role, you will:
  • Research and develop multimodal perception and authentication methods across visual, audio, and other sensing signals.
  • Explore how specialized perception models and larger multimodal models can work together.
  • Design data, training, and evaluation approaches that improve performance in real-world conditions.
  • Study model behavior, robustness, and failure modes across sensing, data, and deployment environments.
  • Integrate and validate new capabilities in real-time or resource-constrained systems.
  • Work with hardware, firmware, software, and product teams to turn research into working systems.
You might thrive in this role if you:
  • Have a strong background in computer vision, audio or speech machine learning, multimodal learning, or sensing.
  • Have experience developing specialized machine learning models, larger multimodal models, or both.
  • Have brought research ideas into practical systems, prototypes, or products.
  • Know how to design experiments, build evaluations, and investigate model behavior.
  • Have wo
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Multimodal Perception and Authentication
Machine Learning Engineer, Multimodal Perception and Authentication

Slope • San Francisco (CA)

On-site
USD 150,000 - 230,000
Multimodal Perception & Authentication ML Engineer
Multimodal Perception & Authentication ML Engineer

OpenAI • United States

Hybrid
USD 140,000 - 210,000
Relocation assistance
Machine Learning Engineer, Multimodal Perception and Authentication
Machine Learning Engineer, Multimodal Perception and Authentication

OpenAI • San Francisco (CA)

On-site
USD 342,000 - 399,000
Relocation assistance
Hybrid work model
ML Engineer: Multimodal Perception & Authentication
ML Engineer: Multimodal Perception & Authentication

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 399,000
Relocation assistance
Hybrid work model
Multimodal ML Engineer: Perception & Authentication
Multimodal ML Engineer: Perception & Authentication

Slope • San Francisco (CA)

On-site
USD 150,000 - 230,000
Research, Vision Expertise
Research, Vision Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Machine Learning Engineer
Machine Learning Engineer

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000
Applied Machine Learning Research Engineer
Applied Machine Learning Research Engineer

Apple • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Research Scientist - Multimodal Agent, Consumer Devices
Research Scientist - Multimodal Agent, Consumer Devices

OpenAI • United States

Remote
USD 150,000 - 230,000
AI Researcher — Perception & Multimodal Learning
AI Researcher — Perception & Multimodal Learning

Momi US • New York (NY)

On-site
USD 200,000 - 300,000