AI Computing Architecture Researcher

PERSOL SINGAPORE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PERSOL SINGAPORE PTE. LTD. is seeking an expert to design and optimize next-generation AI computing architectures and platforms for LLM, AIGC, and distributed AI workloads.

You will collaborate with researchers and engineers to improve performance across GPU, NPU, and heterogeneous environments. You will develop software components interfacing with hardware accelerators, improve performance, and contribute to research on agent capabilities and deployment.

Qualifications

  • Master’s or PhD in Computer Science or Electrical Engineering.
  • Strong experience in high-performance AI systems.
  • Proficiency in C/C++, Python, and AI-DSL for low-level programming.
  • Expertise in model optimization: quantization, pruning, sparsity.
  • Experience with GPU/NPU programming and AI frameworks.
  • Background in distributed systems and multi-agent coordination.
  • Familiarity with containerization and cloud infrastructure.

Responsibilities

  • Design and optimize AI computing architectures and platforms for LLMs, AIGC, distributed AI workloads, and intelligent agents.
  • Develop software components interfacing with GPUs, NPUs, and AI chips.
  • Improve performance and efficiency of distributed computing and GPU/NPU acceleration.
  • Collaborate with AI researchers to enhance platform performance and resolve system-level issues.
  • Design and implement tools for deployment, monitoring, scaling, and orchestration of AI workloads and agents.
  • Integrate new AI models and algorithms ensuring scalability and fault tolerance.
  • Conduct profiling and benchmarking to achieve high throughput and low latency.
  • Contribute to research on agent capabilities including skill abstraction and deployment.

Skills

C/C++
Python
AI-DSL
CUDA
Triton
Distributed systems
Multi-agent coordination
PyTorch
TensorFlow
MXNet
Docker
Kubernetes
AWS
GCP
Azure
LLVM
MLIR
TVM
Agent skill deployment

Education

Master’s or PhD in Computer Science or Electrical Engineering

Tools

PyTorch
TensorFlow
CUDA
Docker
Kubernetes
SGLang
vLLM
HuggingFace
AWS
GCP
Azure

Job description

In this role, you will focus on designing and optimizing next-generation AI computing architectures and platforms for LLM, AIGC, distributed AI workloads, and intelligent agent systems. You will collaborate with researchers and engineers to improve the performance, scalability, and efficiency of AI infrastructure across GPU, NPU, and heterogeneous computing environments.

Key Responsibilities:
  • Design and optimize AI computing architectures and platforms for LLM, AIGC, distributed AI, and intelligent agent workloads. Develop software components that interface with hardware accelerators such as GPUs, NPUs, and specialized AI chips.
  • Improve performance and efficiency of distributed computing, parallel processing, and GPU/NPU acceleration.
  • Collaborate with AI researchers to enhance platform performance for complex AI and agent-based applications and also to resolve system-level performance issues across architecture and platform components.
  • Design and implement tools and frameworks for deployment, monitoring, scaling, and orchestration of AI workloads and agent systems.
  • Integrate new AI models, agent controllers, and algorithms into the computing architecture and platform while ensuring scalability, fault tolerance, and efficiency.
  • Conduct profiling, benchmarking, and performance optimization to achieve high throughput and low latency.
  • Contribute to research on agent capabilities, including skill abstraction, composition, learning, and deployment.
  • Stay updated on advancements in AI hardware, system architecture, and agent technologies, and continuously improve platform capabilities.
Required Qualifications:
  • Master’s or PhD in Computer Science, Electrical Engineering, or related fields, with strong experience in high-performance AI systems or computing platforms.
  • Proficiency in C/C++, Python, AI-DSL with a focus on low-level programming for high-performance systems.
  • In-depth knowledge of AI model optimization techniques such as quantization, model graph pruning, and model parameter compression and sparsity algorithm.
  • In-depth knowledge of parallel programming, distributed systems, and multi-agent coordination strategies.
  • Experience with GPU/NPU programming (e.g., CUDA, Triton, cuTile or similar DSLs).
  • Familiarity with AI/machine learning frameworks such as PyTorch, TensorFlow, or MXNet.
  • Understanding of AI model deployment, orchestration, and optimization on large-scale platforms, including agent skill deployment and runtime management.
  • Experience with containerization (Docker, Kubernetes), LLM deployment platforms (SGLang, vLLM, HuggingFace), and cloud infrastructure (AWS, GCP, Azure).
  • Strong problem-solving skills and the ability to optimize software performance for both traditional AI and agent-based workloads.
  • Experience with AI/deep learning frameworks and tools for performance profiling and optimization, including optimization in agent systems.
  • Knowledge of low-level hardware optimization, including memory management and instruction-level tuning for CPU/GPU/NPU architectures.
  • Familiarity with AI compiler and AI computing architecture, such as LLVM, MLIR, TVM.
  • Background in designing and implementing high-performance distributed systems and storage solutions, including those supporting agent coordination and skill sharing.
  • Strong understanding of networking, I/O, and data management techniques for AI and agent workloads.
  • Exposure to multi-agent system design, including communication protocols, coordination mechanisms, skill composition frameworks techniques, harness engineering for AI workloads.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Computing Architecture Researcher
AI Computing Architecture Researcher

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI Computing Architect for LLM & Distributed AI
AI Computing Architect for LLM & Distributed AI

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
AI Compute Architect: High-Performance Systems Lead
AI Compute Architect: High-Performance Systems Lead

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI Engineer
AI Engineer

Accenture Southeast Asia • Singapore

On-site
SGD 180,000 - 240,000
Senior Scientist I , Computing & Intelligence, IAIC
Senior Scientist I , Computing & Intelligence, IAIC

A*STAR RESEARCH ENTITIES • Singapore

On-site
SGD 90,000 - 140,000
Research Scientist (AI Model Optimization), IPV, ARTC
Research Scientist (AI Model Optimization), IPV, ARTC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 80,000 - 120,000
Senior AI Engineer - AI Agents / RAG / Fine-Tuning
Senior AI Engineer - AI Agents / RAG / Fine-Tuning

magellan technology research institute (mtri) • Singapore

On-site
SGD 70,000 - 120,000
AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000