AI20P Library Engineer, Machine Learning Acceleration

Qpisemi

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Qpisemi in Bengaluru is seeking an AI20P Library Engineer to bridge AI research and hardware accelerators, designing high-performance libraries, kernels, and runtimes for distributed environments.

You will optimize operators, build scalable runtimes, and collaborate with compiler teams and hardware engineers to enable researchers to train and deploy frontier models at scale.

Strong C++ and Python, plus hands-on ML framework experience, are required for this role.

Qualifications

  • Education: B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a highly quantitative field.
  • Programming: Exceptional proficiency in C++ and Python.
  • ML Ecosystem: Hands-on experience with at least one major AI/ML framework (e.g., JAX, PyTorch).
  • Systems Knowledge: Strong understanding of concurrent computations, memory hierarchies (HBM, cache), and hardware accelerators.
  • Communication: Excellent cross-functional communication skills to translate complex research needs into production-ready software architecture.

Responsibilities

  • Kernel Development & Optimization: Design, implement, and tune high- performance custom operators and mathematical kernels specifically for AI20P architectures (assembly or intrinsic level).
  • Distributed Computing Runtimes: Build software abstractions and libraries that manage multi-host setups, sharding, and multi-dimensional parallelization (e.g., Megatron-LM style tensor/pipeline parallelism).
  • Communication Primitives: Design and optimize custom collective communication algorithms (AllReduce, AllGather) to minimize latency and maximize throughput over high-speed interconnects.
  • Performance Tuning: Profile distributed training workloads to identify and eliminate memory bottlenecks, network stalls, and suboptimal hardware utilization.
  • Cross-Functional Collaboration: Partner with compiler developers (XLA), hardware architects, and research scientists to co-design the future software- hardware ecosystem.

Skills

C++
Python
ML framework experience
Parallel computing

Education

CS/EE degree

Tools

Megatron-LM
Kubernetes
CUDA

Job description

AI20P Library Engineer, Machine Learning Acceleration

We are looking for an experienced AI20P Library Engineer to bridge the gap between cutting-edge AI research and our hardware accelerators. In this role, you will design, develop, and optimize high-performance software libraries, kernels, and parallel compute runtimes for distributed AI20P environments. You will empower ML researchers and cloud developers to train and deploy frontier AI models at unprecedented scales.

Key Responsibilities
  • Kernel Development & Optimization: Design, implement, and tune high- performance custom operators and mathematical kernels specifically for AI20P architectures (assembly or intrinsic level).
  • Distributed Computing Runtimes: Build software abstractions and libraries that manage multi-host setups, sharding, and multi-dimensional parallelization (e.g., Megatron-LM style tensor/pipeline parallelism).
  • Communication Primitives: Design and optimize custom collective communication algorithms (AllReduce, AllGather) to minimize latency and maximize throughput over high-speed interconnects.
  • Performance Tuning: Profile distributed training workloads to identify and eliminate memory bottlenecks, network stalls, and suboptimal hardware utilization.
  • Cross-Functional Collaboration: Partner with compiler developers (XLA), hardware architects, and research scientists to co-design the future software- hardware ecosystem.
Qualifications
  • Education: B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a highly quantitative field.
  • Programming: Exceptional proficiency in C++ and Python.
  • ML Ecosystem: Hands-on experience with at least one major AI/ML framework (e.g., JAX, PyTorch).
  • Systems Knowledge: Strong understanding of concurrent computations, memory hierarchies (HBM, cache), and hardware accelerators.
  • Communication: Excellent cross-functional communication skills to translate complex research needs into production-ready software architecture.
Preferred Qualifications (Nice-to-Have)
  • Experience working with compiler construction or optimizing code generation for hardware.
  • Familiarity with cloud-based cluster managers (e.g., Kubernetes, SLURM) for executing large-scale distributed ML workloads.
  • Prior contributions to optimizing Large Language Models (LLMs) or multimodal architectures for low-precision inference and training.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI20P Compiler Engineer
AI20P Compiler Engineer

Qpisemi • Bengaluru

On-site
INR 2,000,000 - 3,500,000
Senior AI Software Performance Engineer
Senior AI Software Performance Engineer

BigStep Technologies • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

Mulya Technologies • India

On-site
INR 4,500,000 - 9,000,000
Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Chennai District

On-site
INR 2,000,000 - 3,000,000
Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 400,000 - 700,000
AI model optimization & acceleration Engineer
AI model optimization & acceleration Engineer

L&T Technology Services • Bengaluru

On-site
INR 1,800,000 - 3,200,000
GPU Compute & MLIR Compiler Engineer
GPU Compute & MLIR Compiler Engineer

BuildxPartners • Bengaluru

On-site
INR 2,000,000 - 3,000,000
AI Engineer Model Optimization & Acceleration
AI Engineer Model Optimization & Acceleration

Sunrise Biztech Systems • Bangalore Rural

On-site
INR 1,200,000 - 2,400,000
Software Engineer
Software Engineer

MulticoreWare, Inc. • Chennai District

On-site
INR 900,000 - 1,400,000