Robotics Perception Engineer, Vision Models & Mapping

Ndimensions

Boston, Northern (MA, KY)

Hybrid

USD 130,000 - 190,000

Full time

9 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Ndimensions Labs is seeking a Robotics Perception Engineer to own the perception layer that links vision, language, and navigation for home robotics. The role sits between AI and navigation teams, working with real hardware in real homes, on a path toward robust, language-grounded perception.

The position emphasizes real-world experimentation, calibration, and deployment on embedded GPUs, with close collaboration across perception, mapping, and control teams in Boston, MA or Toronto, ON.

Qualifications

  • Experience on physical robotic systems, not just simulation.
  • Grounding text to pixels with vision-language models.
  • Real-time inference on embedded GPUs.
  • Deep calibration and geometry knowledge (intrinsics, stereo, extrinsics).
  • Strong Python and C++, with ROS2 on real robots.

Responsibilities

  • Close the loop between language and pixels: verify that the grasped object is the commanded one in data and live inference.
  • Ground language in the scene, turning a text query and an RGB-D frame into a pick target or navigation goal.
  • Transform 2D detections into a persistent, object-level semantic map usable by language queries.
  • Classify dynamic vs static map elements so the robot trusts changing parts of the environment.
  • Evaluate, optimize, and deploy detection/segmentation models with real-time constraints on embedded GPUs.
  • Build vision-based quality checks on training data, pre-verify episodes, and redact faces/personal info.
  • Own camera calibration and rectification across the robot, ensuring depth/localization accuracy.

Skills

Python
C++
ROS2
Vision-language models
Grounding text to pixels
Embedded hardware

Education

Master's or PhD in Computer Vision/Robotics/CS

Tools

ROS2

Job description

Robotics Perception Engineer, Vision Models & Mapping

Full time position • Boston, MA or Toronto, ON

About Ndimensions Labs

At Ndimensions, we're inventing the infrastructure for next-generation robotics AI systems. Almost everything our robots know about the world starts as pixels from a camera, and so do most of the hard problems we have left.

We're looking for a Robotics Perception Engineer to own that layer. This role sits between our AI and navigation teams, working on the vision problems both depend on: whether the robot is really picking up the object it was asked to, whether it can find an object described in plain language, whether perception is fast enough to run on the robot itself, and whether the cameras are calibrated well enough to trust. It is a hands-on role on real hardware in real homes.

What You'll Do
  • Close the loop between language and pixels: verify that the object the robot grasps is the object it was commanded to grasp, both in recorded training data and live at inference time.
  • Ground language in the scene, turning a text query and an RGB-D frame into a pick target or a navigation goal.
  • Take our maps from geometric to semantic: build a persistent, object-level representation of a home that the robot can query in language, so a named object resolves to a place it can drive to.
  • Give the map a sense of what moves. Classify people, pets, and movable furniture as dynamic or semi-static, so the robot knows which parts of a map to trust and which to expect to have changed.
  • Evaluate, optimize, and deploy detection and segmentation models on-robot within a real-time latency budget on embedded GPUs.
  • Build vision-based quality checks over our training data: pre-verify episodes, catch mis-routed or dropped camera streams, and redact faces and personal information from what we keep.
  • Own camera calibration and rectification across the robot, and the metrics that show depth and localization are good enough to trust.
What We're Looking For
  • Strong background in computer vision for robotics, with real experience on physical systems rather than datasets and simulation alone.
  • Experience with vision-language models: grounding text to pixels, and judging whether a model's output actually matches the instruction it was given.
  • Experience with semantic or open-vocabulary 3D scene representations: fusing 2D detections into a persistent map and keeping object identity stable across viewpoints and sessions.
  • Experience training, fine-tuning, and deploying detection and segmentation models, and optimizing them for real-time inference on embedded hardware.
  • Deep hands-on calibration experience and command of the underlying geometry: intrinsics, stereo and multi-camera extrinsics, rectification, hand-eye calibration, projection and back-projection, and frame conventions.
  • Practical experience with depth from cameras, classical or learned, and a feel for how depth error propagates downstream.
  • Strong Python and working C++, with ROS2 experience on real robots.
  • A measurement-first instinct: you define the metric and build the rig before claiming an improvement, and you can design a credible evaluation when no ground truth exists.
  • Master's or PhD in Computer Vision, Robotics, Computer Science, or a related field (or equivalent industry experience).
Bonus (Not Required)
  • Publications in vision or robotics venues (CVPR, ICRA, IROS, CoRL), or open-source contributions to vision and calibration tooling.
  • Work on vision-language-action policies or other multimodal models for manipulation.
  • Experience with visual or visual-inertial odometry and SLAM, or with scene graphs and other queryable map representations.
  • Experience building or evaluating multimodal perception for vision-language-action (VLA) policies, including grounding observations and instructions into objects, locations, or manipulation targets.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Robotics Perception Engineer, Vision Models & Mapping
Robotics Perception Engineer, Vision Models & Mapping

ndimensions labs • Boston (MA)

On-site
USD 130,000 - 180,000
Robotics Vision & Mapping Engineer
Robotics Vision & Mapping Engineer

Ndimensions • Boston (MA), Northern (KY)

Hybrid
USD 130,000 - 190,000
Software Engineer, Perception
Software Engineer, Perception

Relling • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff Perception Vision Engineer
Staff Perception Vision Engineer

Motion Recruitment • Waltham (MA)

On-site
USD 140,000 - 190,000
Health, Dental, and Vision Insurance
PTO
401k
Robot Perception Expert (human)
Robot Perception Expert (human)

NEURA Robotics • Germany (OH)

On-site
USD 120,000 - 180,000
Computer Vision Engineer Palo Alto, California
Computer Vision Engineer Palo Alto, California

BrightAI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Mechanical Engineer, Robotics
Mechanical Engineer, Robotics

ndimensions labs • Boston (MA)

On-site
USD 120,000 - 180,000
Computer Vision Engineer
Computer Vision Engineer

BrightAI • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Robotics Perception AI Engineer
Robotics Perception AI Engineer

Veriipro • Warren (MI)

On-site
USD 90,000 - 120,000
Robotics Perception Engineer
Robotics Perception Engineer

Fluency Digital, Inc. • Boston (MA)

On-site
USD 80,000 - 100,000