Senior Machine Learning Engineer - Perception 3D Segmentation

Zoox

Foster City (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoox is seeking a senior researcher/engineer to advance 3D occupancy and segmentation networks. You will design multi-modal fusion (Lidar, Camera, Radar) and build voxel-based or BEV models with temporal coherence for on-vehicle inference.

Collaboration with downstream teams will refine geometry outputs for complex urban driving scenarios. Requires 6+ years in 3D computer vision/ML, proficiency in PyTorch, and experience with TensorRT/CUDA and C++.

Qualifications

  • MS or PhD in Computer Science, Robotics, Machine Learning, or related field with 6+ years of industry experience.
  • Deep expertise in 3D Computer Vision and Deep Learning, specifically with voxel-based or BEV architectures.
  • Strong proficiency in Python and deep learning frameworks (PyTorch) for model training and design; some C++ experience.

Responsibilities

  • Design and implement state-of-the-art multi-modal sensor fusion architectures (Lidar, Camera, Radar) to predict 3D occupancy, semantic segmentation, and flow.
  • Develop vision-first fusion strategies to enhance geometric understanding and reduce dependency on sparse sensor modalities.
  • Engineer temporal processing modules to improve stability and consistency of predictions over time.
  • Optimize model architectures for real-time on-vehicle inference, balancing high-fidelity range extension with strict latency constraints.
  • Collaborate with downstream consumers (Tracking, Prediction, Planner) to refine geometric outputs for complex maneuvering.

Skills

3D computer vision
BEV architectures
Temporal data sequences
Real-time inference optimization

Education

MS or PhD in CS/Robotics/ML or related

Tools

PyTorch
C++
TensorRT/CUDA
Sparse convolutions
NeRF / Gaussian splats

Job description

The Perception team at Zoox is responsible for the robot’s understanding of the world, fusing data from Lidar, Radar, and Cameras to create a unified representation of the environment. In this role, you will contribute to the development of our next-generation 3D occupancy and segmentation networks. You will architect and optimize high-performance deep learning models that generate dense, temporally consistent voxel representations of the driving environment. This work is critical for enabling our vehicle to navigate complex urban scenarios, handle rare obstacles, and drive safely in tight spaces by providing precise geometry and motion estimates to downstream planners.

Responsibilities
  • Design and implement state-of-the-art multi-modal sensor fusion architectures (Lidar, Camera, Radar) to predict 3D occupancy, semantic segmentation, and flow.
  • Develop vision-first fusion strategies to enhance geometric understanding and reduce dependency on sparse sensor modalities.
  • Engineer temporal processing modules to improve the stability and consistency of predictions over time.
  • Optimize model architectures for real-time on-vehicle inference, balancing high-fidelity range extension with strict latency constraints.
  • Collaborate with downstream consumers (Tracking, Prediction, Planner) to refine geometric outputs, such as contours and free-space estimations, for complex maneuvering.
Qualifications
  • MS or PhD in Computer Science, Robotics, Machine Learning, or related field with 6+ years of industry experience.
  • Deep expertise in 3D Computer Vision and Deep Learning, specifically with voxel-based or BEV (Bird's Eye View) architectures.
  • Strong proficiency in Python and deep learning frameworks (PyTorch) for model training and design as well as some experience in C++ for model integration.
  • Experience with multi-sensor fusion (Lidar, Camera, Radar) and handling temporal data sequences.
  • Experience with occupancy networks, implicit representations (NeRF/Gaussian Splats), or scene flow estimation.
  • Experience optimizing models for TensorRT/CUDA to achieve low-latency inference.
  • Familiarity with sparse convolutions or query-based architectures for efficient 3D processing.
  • Experience with Vision Language Model, or multi-modal 3D foundation model, or World Model, or VLA.
Accommodations

If you need an accommodation to participate in the application or interview process please reach out to accommodations@zoox.com or your assigned recruiter.

Diversity and Inclusion

You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer - Perception
Senior Machine Learning Engineer - Perception

Zoox • Boston (MA)

On-site
USD 242,000 - 290,000
Health insurance
Paid time off
Stock appreciation rights
+1
Software Engineer - Perception & Sensing
Software Engineer - Perception & Sensing

Zoox • Foster City (CA)

On-site
USD 196,000 - 278,000
Paid time off
Health insurance
Long-term and short-term disability insurance
+2
Senior ML Engineer - 3D Perception & Multimodal Fusion
Senior ML Engineer - 3D Perception & Multimodal Fusion

Zoox • Foster City (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer - Semantic Reasoning (Highway) [Filled July 10, 2026]
Machine Learning Engineer - Semantic Reasoning (Highway) [Filled July 10, 2026]

Zoox • Foster City (CA)

On-site
USD 180,000 - 260,000
Health insurance
Zoox Stock Appreciation Rights
Amazon RSUs
+1
Machine Learning Engineer - 3D Sensor Simulation
Machine Learning Engineer - 3D Sensor Simulation

Jobzhr • Foster City (CA)

On-site
USD 176,000 - 257,000
Senior Staff Machine Learning Engineer - Perception
Senior Staff Machine Learning Engineer - Perception

Zoox • Boston (MA)

On-site
USD 277,000 - 389,000
Health Insurance
Paid Time Off
Long-term Care Insurance
+1
Machine Learning Engineer - Semantic Reasoning (Highway)
Machine Learning Engineer - Semantic Reasoning (Highway)

Zoox • Foster City (CA)

On-site
USD 189,000 - 258,000
Zoox Stock Appreciation Rights
Amazon RSUs
Health insurance
+1
Director, Perception Detection
Director, Perception Detection

Dormont Manufacturing Co • Foster City (CA)

On-site
USD 150,000 - 200,000
Data Scientist - Perception Verification and Validation
Data Scientist - Perception Verification and Validation

Zoox • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 228,000
Health insurance
Stock Appreciation Rights
Amazon RSUs
Technical Program Manager - Data Operations Lead
Technical Program Manager - Data Operations Lead

Zoox • Boston (MA)

On-site
USD 163,000 - 223,000
Paid time off
Health insurance
Stock Appreciation Rights
+1