Robotics Perception Engineer, Vision Models & Mapping

ndimensions labs

Boston (MA)

On-site

USD 130,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Ndimensions is building the infrastructure for next-generation robotics AI systems, and seeks a Robotics Perception Engineer to own the perception layer. The role sits between AI and navigation teams, tackling vision problems to verify object grounding, language grounding to pixels, and real-time on-robot inference on embedded GPUs.

It is a hands-on role on real hardware in real homes. You will work with Python and C++ in ROS2, calibrate cameras, build semantic maps, and support ground-truth

Qualifications

  • Strong background in computer vision for robotics.
  • Experience grounding language to pixels.
  • Experience with depth sensing and 3D scene representations.
  • Experience deploying detection/segmentation models on embedded GPUs.

Responsibilities

  • Bridge language and pixels to verify the grasped object matches the instruction.
  • Ground language in the scene to produce a pick target or navigation goal.
  • Build a persistent, object-level map that can be queried by language.
  • Classify dynamic vs static elements to indicate trustworthy map regions.
  • Evaluate and deploy perception models with real-time latency constraints on embedded GPUs.
  • Calibrate cameras and manage depth/localization accuracy.

Skills

Robotics CV
Vision-language grounding
Python
C++
ROS2
On-robot experimentation

Education

Master’s or PhD in Computer Vision / Robotics

Tools

ROS2
OpenCV
CUDA
PCL

Job description

At Ndimensions, we're inventing the infrastructure for next-generation robotics AI systems. Almost everything our robots know about the world starts as pixels from a camera, and so do most of the hard problems we have left. We're looking for a Robotics Perception Engineer to own that layer. This role sits between our AI and navigation teams, working on the vision problems both depend on: whether the robot is really picking up the object it was asked to, whether it can find an object described in plain language, whether perception is fast enough to run on the robot itself, and whether the cameras are calibrated well enough to trust. It is a hands‑on role on real hardware in real homes.

What You’ll Do
  • Close the loop between language and pixels: verify that the object the robot grasps is the object it was commanded to grasp, both in recorded training data and live at inference time.
  • Ground language in the scene, turning a text query and an RGB‑D frame into a pick target or a navigation goal.
  • Take our maps from geometric to semantic: build a persistent, object‑level representation of a home that the robot can query in language, so a named object resolves to a place it can drive to.
  • Give the map a sense of what moves. Classify people, pets, and movable furniture as dynamic or semi‑static, so the robot knows which parts of a map to trust and which to expect to have changed.
  • Evaluate, optimize, and deploy detection and segmentation models on‑robot within a real‑time latency budget on embedded GPUs.
  • Build vision‑based quality checks over our training data: pre‑verify episodes, catch mis‑routed or dropped camera streams, and redact faces and personal information from what we keep.
  • Own camera calibration and rectification across the robot, and the metrics that show depth and localization are good enough to trust.
What We’re Looking For
  • Strong background in computer vision for robotics, with real experience on physical systems rather than datasets and simulation alone.
  • Experience with vision‑language models: grounding text to pixels, and judging whether a model’s output actually matches the instruction it was given.
  • Experience with semantic or open‑vocabulary 3D scene representations: fusing 2D detections into a persistent map and keeping object identity stable across viewpoints and sessions.
  • Experience training, fine‑tuning, and deploying detection and segmentation models, and optimizing them for real‑time inference on embedded hardware.
  • Deep hands‑on calibration experience and command of the underlying geometry: intrinsics, stereo and multi‑camera extrinsics, rectification, hand‑eye calibration, projection and back‑projection, and frame conventions.
  • Practical experience with depth from cameras, classical or learned, and a feel for how depth error propagates downstream.
  • Strong Python and working C++, with ROS2 experience on real robots.
  • A measurement‑first instinct: you define the metric and build the rig before claiming an improvement, and you can design a credible evaluation when no ground truth exists.
  • Master’s or PhD in Computer Vision, Robotics, Computer Science, or a related field (or equivalent industry experience).
Bonus (Not Required)
  • Publications in vision or robotics venues (CVPR, ICRA, IROS, CoRL), or open‑source contributions to vision and calibration tooling.
  • Work on vision‑language‑action policies or other multimodal models for manipulation.
  • Experience with visual or visual‑inertial odometry and SLAM, or with scene graphs and other queryable map representations.
  • Experience building or evaluating multimodal perception for vision‑language‑action (VLA) policies, including grounding observations and instructions into objects, locations, or manipulation targets.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Robotics Perception Engineer, Vision Models & Mapping
Robotics Perception Engineer, Vision Models & Mapping

Ndimensions • Boston (MA), Northern (KY)

Hybrid
USD 130,000 - 190,000
Software Engineer, Perception
Software Engineer, Perception

Relling • San Francisco (CA)

On-site
USD 140,000 - 210,000
Robotics Vision & Mapping Engineer
Robotics Vision & Mapping Engineer

Ndimensions • Boston (MA), Northern (KY)

Hybrid
USD 130,000 - 190,000
Robotics Perception AI Engineer
Robotics Perception AI Engineer

Veriipro • Warren (MI)

On-site
USD 90,000 - 120,000
Robot Perception Expert (human)
Robot Perception Expert (human)

NEURA Robotics • Germany (OH)

On-site
USD 120,000 - 180,000
Computer Vision Engineer
Computer Vision Engineer

BrightAI • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Computer Vision Engineer Palo Alto, California
Computer Vision Engineer Palo Alto, California

BrightAI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Robotics Vision Engineer: Grounding Language & Perception
Robotics Vision Engineer: Grounding Language & Perception

ndimensions labs • Boston (MA)

On-site
USD 130,000 - 180,000
Perception Engineer — 3D Representation & Navigation
Perception Engineer — 3D Representation & Navigation

AI Chopping Block • Irvine (CA), Northern (KY)

Hybrid
USD 120,000 - 160,000
Mechanical Engineer, Robotics
Mechanical Engineer, Robotics

ndimensions labs • Boston (MA)

On-site
USD 120,000 - 180,000