Research Scientist (End-to-End & Multimodal Models)

Black Sesame Technologies (Singapore) Pte Ltd

Singapore

On-site

SGD 150,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Black Sesame Technologies (Singapore) Pte Ltd is seeking a research-led engineer to design and develop autonomous driving frameworks that integrate perception, prediction, and planning into a unified system, leveraging vision-only and vision-language modeling to support urban and highway scenarios.

The role emphasizes end-to-end and hybrid architectures, exploring Vision-Language Models to enhance scene understanding, reasoning, and robustness, with hands-on work in model deployment on embedded

Qualifications

  • PhD in CS/AI/Robotics or related field.
  • Hands-on experience with end-to-end deep learning–based modeling.
  • Experience in planning, control, or decision-making modules using DL.
  • Interest in Vision-Language Models and multimodal learning.
  • Proficiency in C/C++ and Python with real-time deployment.
  • Experience with BEV representations and multi-task learning.
  • Experience with system integration and real-vehicle testing is a plus.

Responsibilities

  • Lead the design and implementation of end-to-end autonomous driving models.
  • Develop pure vision-based end-to-end systems with multi-task capabilities.
  • Explore Vision-Language Models to improve scene understanding and reasoning.
  • Optimize and deploy models on embedded platforms, including deployment and testing.
  • Deliver production-ready solutions for highway and urban driving with scalable deployment.

Skills

Deep learning modeling
C/C++
Python
Vision-Language Models
BEV representations
Multi-task learning
Real-time inference
System integration

Education

Ph.D. in CS/AI/Robotics

Job description

In this role, you will be responsible for the end-to-end design and development of autonomous driving frameworks. You will integrate mainstream perception, prediction, and planning technologies into a unified modeling system, leveraging both vision-only and vision-language modeling paradigms, to support autonomous driving tasks across urban and highway scenarios.

You will play a key role in advancing end-to-end and hybrid architectures, including the exploration of Vision-Language Models (VLMs) to enhance scene understanding, reasoning, and decision‑making robustness in complex driving environments.

Responsibilities:
  • Lead the design and implementation of end-to-end autonomous driving models, including one-stage (sensor-to-control) and two-stage (e.g., perception–planning decoupled) architectures. Define model structures, training pipelines, and optimization strategies for stable and explainable planning outputs.
  • Drive the development of pure vision-based end-to-end systems, integrating multi-task capabilities such as BEV perception, static and dynamic occupancy inference, trajectory prediction, and planning.
  • Explore and apply Vision-Language Models (VLMs) to improve high-level scene understanding, semantic reasoning, and cross-modal representation learning for autonomous driving tasks.
  • Optimize and deploy models on embedded platforms, including inference acceleration, post-processing, system-level integration, performance tuning, stability validation, and on-road testing.
  • Deliver production-ready solutions for elevated highways and urban driving scenarios, enabling scalable deployment and continuous progression toward higher levels of autonomy.
Qualification/ Requirements:
  • Ph.D. degree in Computer Science, Artificial Intelligence, Robotics, or a related field.
  • Strong foundation in autonomous driving systems, with hands-on experience in end-to-end deep learning–based modeling.
  • Practical experience in planning, control, or decision‑making modules using deep learning approaches.
  • Experience or strong interest in Vision-Language Models (VLMs), multimodal learning, or cross-modal representation learning, particularly in applications involving visual scene understanding and reasoning.
  • Proficiency in C/C++ and Python, with experience in real-time inference deployment and performance optimization.
  • Familiarity with BEV-based representations, occupancy prediction, and multi-task learning frameworks.
  • Experience with system integration and real‑vehicle testing is a strong plus.
  • Strong problem‑solving skills, adaptability to complex real-world scenarios, and a results-driven mindset.
  • Strong mathematical foundation in optimization techniques relevant to computer vision and deep learning.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Autonomous Driving: End-to-End Multimodal AI Scientist
Autonomous Driving: End-to-End Multimodal AI Scientist

Black Sesame Technologies (Singapore) Pte Ltd • Singapore

On-site
SGD 150,000 - 210,000
Vision Perception Algorithm Engineer
Vision Perception Algorithm Engineer

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 80,000 - 140,000
Research Engineer (Machine Learning)
Research Engineer (Machine Learning)

Nanyang Technological University Singapore • Singapore

On-site
SGD 60,000 - 90,000
Scientist/Researcher | Robotics | ML & Computer vision
Scientist/Researcher | Robotics | ML & Computer vision

Randstad Singapore • Singapore

On-site
SGD 90,000 - 150,000
Senior Software Engineer, Perception (Robotics)
Senior Software Engineer, Perception (Robotics)

Grab • Singapore

On-site
SGD 120,000 - 180,000
Autonomous Vision Perception Engineer
Autonomous Vision Perception Engineer

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 80,000 - 140,000
Research Engineer: AI-Driven Vehicle Routing & Optimization
Research Engineer: AI-Driven Vehicle Routing & Optimization

Singapore Management University • Singapore

On-site
SGD 48,000 - 66,000
Senior Planning & Prediction Engineer
Senior Planning & Prediction Engineer

Inceptio Technology • Singapore

On-site
SGD 120,000 - 180,000
Urgent! AI Researcher - Computer Vision & Physical Intelligence
Urgent! AI Researcher - Computer Vision & Physical Intelligence

TRUST RECRUIT PTE. LTD. • Singapore

On-site
SGD 150,000 - 190,000
Staff Research Scientist, Localization and Mapping
Staff Research Scientist, Localization and Mapping

Venti • Singapore

On-site
SGD 100,000 - 130,000