Research Engineer - Egocentric Perception Models

Humyn Labs

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Humyn Labs, a leader in physical-world AI, seeks a senior CV/ML engineer to own the model layer of our annotation and perception pipeline in Bengaluru. You will build in-house hand pose models, depth estimation, and ego-centric vision components that survive diverse real-world rigs.

This role requires 4+ years in vision/robotics, Python and PyTorch expertise, and production experience on AWS. You will design training data, evaluate metrics, and deploy scalable inference pipelines.

Qualifications

  • Must have trained a perception model that beat a public baseline and shipped it.
  • Data curation, loss design and eval decisions were yours.
  • 4+ years in this space.
  • Strong Python and PyTorch; you own training and inference code, not notebooks.
  • Production inference experience on AWS: containerization, batch pipelines, orchestration.

Responsibilities

  • Own the model layer of the pipeline for egocentric AI.
  • Develop in-house 3D hand pose model robust to world frame.
  • Improve ego-centric pose across devices with learned despiking and drift correction.
  • Establish real GT via fiducial markers (AprilTag/ArUco) and IMU grounding.
  • Train and deploy models on AWS with GPU orchestration.

Skills

Python & PyTorch
3D hand pose reconstruction
VIO / SLAM
Open‑vocabulary detection/segmentation
Multi‑view geometry
Calibration & GT reading

Education

MS / PhD in computer vision, robotics or ML

Job description

Humyn Labs builds the intelligence layer for physical-world AI — systems that perceive, reason, and act in real environments. Our work sits at the intersection of egocentric video understanding, embodied AI, robotics perception, and voice-driven interaction. We move fast, obsess over data quality, and ship at scale.

Humyn Labs converts human action - across sound, sight, movement, and touch - into high-quality multi-modal data signals for physical AI. Operating across 20+ countries in India, southeast Asia, Latin America, and the Middle East: the real-world environments where physical AI deploys, not the labs where it is built.

Our data isn't just collected; it's evaluated, defended, and production-ready. Because before AI can be trusted, its training data must be.

THE OPPORTUNITY

Our auto-annotation stack currently runs on public off-the-shelf models — for head pose, hand keypoints, depth, and object detection and segmentation. Every one of them was trained on data that does not look like ours, and every one of them has a documented failure mode on our captures: wide-FOV fisheye rigs, gloved hands, heavy occlusion, motion blur, ego-motion, real workshops instead of clean scenes.

We have measured those failures precisely. We are hiring you to remove them — by training our own models, on our own data, and putting them in production behind a gate. This is not a research role adjacent to the pipeline. You own the model layer of the pipeline.

WHAT YOU'LL OWN
  • Metric 3D hand pose. Build the in-house hand model that is stable in the world frame, and decide the representation (MANO topology, 6D vs axis-angle) rather than inheriting it.
  • Egocentric pose that survives the camera. We have narrowed the suspects (rectification residual, FOV, shutter type) and currently route between several public SLAM / VIO systems per clip. Replace routing‑by‑heuristic with a model that makes any capture device deliverable — including learned despiking and drift correction.
  • Ground truth where we have none. Today we validate pose without ground truth (loop‑closure drift, static‑window jitter). Stand up real GT — fiducial‑based (AprilTag / ArUco), gravity‑aligned via IMU — under a hard constraint: our subjects are real workers doing their own jobs, so any protocol must add zero operator burden.
  • Our own depth model. Public depth models do not survive our rigs. We need a depth model of our own: metric, stereo‑native, valid across the full frame on wide FOV, and cheap enough to run on every video hour we deliver. Own the training data, the architecture call, and the accuracy‑versus‑throughput trade.
  • Learned QC instead of thresholds. Our gates are hand‑tuned numbers: epipolar residual, depth QC metrics, MCAP QA checks. Train error‑detection models that flag a bad label before it reaches a customer, per clip, and quantify their catch rate against known‑bad deliveries.
  • Training and serving, both. Train on AWS and ship on AWS: batch GPU orchestration, throughput tuning, daily delivery SLAs. Cost per labeled video hour is one of your numbers.
WHAT WE'RE LOOKING FOR
  • MS / PhD in computer vision, robotics or ML — or equivalent published research output.
  • You have trained a perception model that beat an off‑the‑shelf baseline on a real domain and shipped it. Data curation, loss design and eval decisions were yours.
  • Deep hands‑on work in at least two of: 3D hand / human pose reconstruction, stereo or monocular depth, VIO / SLAM, open‑vocabulary detection and segmentation, video VLMs.
  • Multi‑view geometry is fluent, not looked up: rectification, epipolar residual, Q matrix and depth sign conventions, triangulation, world vs camera frame.
  • You can read a calibration report and say whether the problem is the camera or the model — and be right.
  • Strong Python and PyTorch; you own training and inference code, not notebooks.
  • Production inference experience on AWS: containerization, batch pipelines, orchestration.
  • 4+ years on problems in this space.
NICE TO HAVE
  • Publications in egocentric vision, 3D hand pose, VLA / VLM models, or robot learning.
  • IMU‑synced multimodal data, gravity alignment, camera–IMU extrinsics.
  • Worked with Ego4D, EgoExo4D, Open X‑Embodiment or comparable large egocentric corpora.
  • Open‑source contributions to vision or robotics projects.
FIRST 60 DAYS
  • Day 30. Reproduce our label QC numbers end to end and tell us which of the four label types is costing us the most, with evidence.
  • Day 60. One in‑house model beating its public baseline on our eval set, on that label type.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Engineer
Hardware Engineer

Humyn Labs • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Lead - Computer Vision & AI Architect
Lead - Computer Vision & AI Architect

Eonix Partners Llp • Bengaluru

Hybrid
INR 6,000,000 - 11,000,000
Platform Engineer - ML Data Pipelines
Platform Engineer - ML Data Pipelines

Humyn Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Computer Vision Engineer
Computer Vision Engineer

Ripik.AI • Dadri

On-site
INR 800,000 - 1,200,000
Head of Technology
Head of Technology

Humyn Labs • Bengaluru

On-site
INR 3,000,000 - 7,000,000
Senior Applied AI/ML Engineer – Computer Vision & Video
Senior Applied AI/ML Engineer – Computer Vision & Video

Objectways • Chennai District

On-site
INR 3,500,000 - 7,000,000
Machine Learning Engineer — Computer Vision
Machine Learning Engineer — Computer Vision

Lear Labs • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Real ownership of production systems
Direct founder access
Strong hardware resources
Machine Learning & Computer Vision Engineer
Machine Learning & Computer Vision Engineer

Sentiac • India

On-site
INR 3,000,000 - 6,000,000
Lead Perception Engineer
Lead Perception Engineer

Origin Control Solutions Ltd • Bengaluru

Hybrid
INR 1,800,000 - 3,000,000
AI Model Developer - Computer Vision
AI Model Developer - Computer Vision

JobItUs • Rajkot

On-site
INR 700,000 - 1,200,000
20 GB GPU server access for experimentation and training
Lead your own annotation team
Exposure to multimodal, real-world AI problems
+1