Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems

Chennai District

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EPAM Systems is looking for an experienced Solution Architect in Chennai, India, specializing in GPU/TPU Kernel Optimization. The successful candidate will design and optimize high-performance kernels for machine learning operations, redefining performance boundaries while shaping developer infrastructure. The role requires 12-18 years of software development experience and expertise in optimizing code using languages like Pallas, CUDA, and Triton. Excellent communication skills and knowledge of ML frameworks such as JAX and PyTorch are essential.

Qualifications

  • 12-18 years of experience in software development.
  • Expertise in optimizing TPU/GPU code using low-level kernel languages.
  • Strong knowledge of ML frameworks, model optimization, and low-precision formats.

Responsibilities

  • Design and optimize high-performance kernels targeting TPU and GPU architectures.
  • Architect infrastructure for benchmarking suites and performance analysis tools.
  • Engage with ML researchers and framework developers to enhance adoption.

Skills

GPU/TPU Kernel Optimization
Pallas language
Triton
CUDA
JAX
PyTorch
Performance Optimization
Cross-Functional Communication

Tools

MLIR
OpenXLA

Job description

We are seeking an experienced Solution Architect specializing in GPU/TPU Kernel Optimization to design and optimize high-performance kernels for cutting-edge Machine Learning operations. In this role, you will redefine performance boundaries across massive training runs and high-speed inference workloads while shaping the developer infrastructure that powers next-generation AI systems.

Responsibilities
  • Design and optimize high-performance kernels (using languages like Pallas, Mosaic and Triton) targeting Tensor Processing Unit (TPU) and Graphics Processing Unit (GPU) architectures for critical Machine Learning (ML) operations, redefining what’s possible from massive training runs to high-speed inference
  • Architect infrastructure such as benchmarking suites, autotuning frameworks, performance analysis tools, regression testing and documentation
  • Transform how the developer community interacts with increasingly critical custom kernels in key Open-Source Software (OSS) libraries
  • Track the latest advancements in hardware architectures, compiler technologies and AI models to identify new opportunities for performance optimization through custom kernels
  • Engage with ML researchers, framework developers (Just After eXecution JAX, PyTorch) and compiler engineers (Accelerated Linear Algebra XLA) to enhance adoption
  • Identify new requirements and address bottlenecks by providing appropriate solutions
Requirements
  • 12-18 years of experience in software development
  • Expertise in optimizing TPU/GPU code using low-level kernel languages like Pallas, Compute Unified Device Architecture (CUDA) or Triton
  • Knowledge of ML Frameworks (JAX/PyTorch), common operations like attention and Mixture of Experts (MoEs), including model optimization and low-precision formats
  • Understanding of modern accelerators (e.g., data movement, pipelining, heterogeneous compute and scale-out)
  • Understanding of compiler principles (optimization, code generation) and toolchains such as MLIR, OpenXLA
  • Showcase of building developer infrastructure, including Open-Source Software (OSS) libraries, flexible high-performance APIs and easy-to-consume documentation to empower the community
  • Excellent investigative and problem-solving capabilities with communication skills across cross-functional teams
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Hyderabad

On-site
INR 2,500,000 - 3,500,000
GPU Compute & MLIR Compiler Engineer
GPU Compute & MLIR Compiler Engineer

BuildxPartners • Bengaluru Urban

On-site
INR 2,500,000 - 5,000,000
Lead Engineer (HPC, GPU, CUDA)
Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix • Thane

On-site
INR 1,800,000 - 3,200,000
Hardware Engineer (Remote | $80 –$100/hr)
Hardware Engineer (Remote | $80 –$100/hr)

Synthires • India

On-site
INR 10,526,000 - 13,158,000
GPU Compute & MLIR Compiler Engineer
GPU Compute & MLIR Compiler Engineer

BuildxPartners • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Graphics/AI Performance Engineer - Staff / Senior
Graphics/AI Performance Engineer - Staff / Senior

BuildxPartners • Bengaluru

On-site
INR 1,000,000 - 3,000,000
Staff ML Compiler Engineer, TPU Performance Optimizations
Staff ML Compiler Engineer, TPU Performance Optimizations

Google • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Architect - GPU Performance
Architect - GPU Performance

NVIDIA • Bengaluru

On-site
INR 1,000,000 - 1,500,000
CUDA Kernel Optimization Specialist
CUDA Kernel Optimization Specialist

Obsidian • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Architect - Performance Verification and Analysis
Architect - Performance Verification and Analysis

NVIDIA Gruppe • Bengaluru

On-site
INR 1,200,000 - 1,800,000