Research Scientist (AI Model Optimization), IPV, ARTC
A*STAR - Agency for Science, Technology and Research
Singapore
On-site
SGD 80,000 - 120,000
Full time
14 days+
Application generator
Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Job summary
A*STAR - Agency for Science, Technology and Research in Singapore is seeking an expert to lead research into optimization techniques for AI deployment. The ideal candidate will have a Ph.D. and extensive experience with AI inference engines and model compression techniques. Responsibilities include designing scalable architectures for high-throughput data streams and conducting co-design to optimize models for specific hardware. Join a cutting-edge team to advance high-performance AI solutions.
Qualifications
Ph.D. in a related field with a focus on High-Performance AI.
Deep understanding of AI Inference Engines like TensorRT and ONNX Runtime.
Mastery of Model Compression techniques including Pruning and Quantization.
Responsibilities
Lead research into optimization techniques to minimize latency.
Design scalable AI deployment architectures for high-resolution data streams.
Conduct hardware-software co-design for optimized model deployment.
Skills
C++
Python
Model Compression
AI Inference Engines
Parallel Computing
High-Performance AI
Education
Ph.D. in Computer Engineering, Computer Science, or Electrical Engineering
Job description
Responsibilities
Lead research into state-of-the-art optimization techniques, including Quantization-Aware Training (QAT), Pruning, Knowledge Distillation, and Neural Architecture Search (NAS) to minimize latency.
Design and implement scalable AI deployment architectures that can handle high-throughput data streams from multiple high-resolution cameras and process sensors simultaneously.
Conduct hardware-software co-design to optimize models for specific deployment targets (e.g., NVIDIA Jetson, TensorRT, FPGAs, or specialized AI accelerators).
Develop and manage asynchronous data pipelines that ensure zero-bottleneck performance from image acquisition to "final sentencing" decisions.
Establish rigorous performance profiling benchmarks to track model latency and memory footprint across various manufacturing environments.
Work with the System Integrator (SI) to ensure that optimized models are seamlessly integrated into the factory-level software stack.
Requirements
Ph.D. in Computer Engineering, Computer Science, Electrical Engineering, or a related field with a focus on High-Performance AI.
Deep understanding of AI Inference Engines (e.g., TensorRT, ONNX Runtime, OpenVINO).
Mastery of Model Compression techniques (Pruning, Quantization, Distillation).
Expertise in C++ and Python for high-performance implementation.
Hands‑on experience with Parallel Computing (CUDA, OpenCL).
Familiarity with Mixed-Precision Training and FP16/INT8 deployment.
Proven ability to architect end-to-end AI systems that balance the trade-off between throughput, latency, and model precision.