Compiler Engineer

Evollabs Tech

Anupgarh

On-site

INR 1,500,000 - 2,100,000

Full time

32 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Evollabs Tech is seeking an experienced compiler engineer to design and implement MLIR dialects and lowering pipelines for our AI accelerator platform. You will lower graphs from ONNX/PyTorch into optimized kernels, focusing on fusion, tiling and vectorization.

You will profile and optimize compiler output to meet strict latency, throughput and power targets, while collaborating with hardware and runtime teams to build a portable, multi‑generation compiler stack.

Qualifications

  • 5+ years of experience in compiler engineering or high-performance systems.
  • Strong programming skills in C++ (modern standards) and Python scripting.
  • Hands-on experience with MLIR and/or LLVM and lowering/optimizing dialects.
  • Deep understanding of compiler internals: IRs, scheduling, vectorization, transformations.
  • Familiarity with AI/ML frameworks (PyTorch, ONNX, TensorFlow).
  • Understanding of tensor operators, matmuls, quantization and dynamic shapes.
  • Knowledge of computer architecture and memory hierarchy.
  • Bachelor’s or higher degree in CS/CE or equivalent.

Responsibilities

  • Define and implement MLIR dialects and lowering pipelines for AI accelerator platform.
  • Lower ML graphs into optimized kernels focusing on fusion, tiling and vectorization.
  • Profile, benchmark and optimize compiler output for latency, throughput, power targets.
  • Collaborate with hardware architects and runtime teams to co-design features.
  • Build a portable compiler stack supporting multiple hardware generations.

Skills

C++ (modern)
Python
MLIR/LLVM
Compiler internals
PyTorch
ONNX
TensorFlow
Vectorization
Hardware awareness

Education

Bachelor’s degree in Computer Science or Computer Engineering

Tools

MLIR
LLVM

Job description

About Us

We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.

About Us

We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.

As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.

You’ll Collaborate With

Silicon, firmware, runtime, datacenter-software and architecture teams to build a mlir-compiler and software stack that brings next-gen AI/NPU hardware into modern datacenter environments and production AI workloads

What You’ll Own
  • Define and implement MLIR dialects and lowering pipelines targeting our AI-accelerator/NPU platform.
  • Lower ML graphs (e.g., from ONNX/PyTorch) into optimized kernels focusing on fusion, tiling and vectorization.
  • Profile, benchmark and optimize compiler output to meet latency, throughput and power targets.
  • Work with hardware architects and runtime/framework teams to co-design compiler features.
  • Build a future‑ready, portable compiler stack that supports multiple hardware generations.
Minimum Qualifications
  • 5+ years of experience in compiler engineering or high‑performance systems.
  • Strong programming skills in C++ (modern standards), and scripting or tooling in Python.
  • Hands‑on experience with MLIR and/or LLVM (designing/extending dialects, lowering, optimization).
  • Deep understanding of compiler internals: IRs, code scheduling, vectorization, loop transformations.
  • Strong familiarity with AI/ML frameworks (e.g.,PyTorch, ONNX, TensorFlow).
  • Excellent understanding of ML operations such as tensor operators, matmuls, quantization, dynamic shapes.
  • Robust understanding of computer architecture and memory hierarchy.
  • Bachelor’s (or higher) degree in Computer Science, Computer Engineering or equivalent
Preferred Qualifications
  • Prior experience targeting NPU/AI accelerator hardware.
  • Contributions to open‑source compiler projects (MLIR, LLVM,Torch-MLIR, ONNX-MLIR).
  • Familiarity with runtime systems, scheduling, resource sharing, memory movement engines and, interconnects.
  • Previous leadership role driving compiler architecture, setting standards, mentoring teams and owning roadmap.
  • Familiarity with software development tooling: Git, CI/CD, debuggers, profilers.
  • Advanced degree (M.S./Ph.D.) preferred.
What Success Looks Like (First 6–9 Months)
  • A production compiler IR/dialect is designed and integrated, successfully lowering key ML workloads to the target NPU.
  • Key operator kernels (matmul, convolution, etc) show measurable performance gains (reduced latency, increased throughput) on the hardware or simulator.
  • The compiler pipeline handles multi‑die/chiplet topology correctly and efficiently schedules across devices.
  • Runtime interfaces and resource scheduling between host and accelerator are functional and validated with real workloads.
  • Telemetry/debug hooks and performance counters from compiled code are available and used for performance analysis by other teams.
  • The architecture and roadmap for next‑gen hardware are defined, and junior engineers are actively mentored within the compiler team.
  • Join us in our mission to democratize AI compute — where your firmware expertise becomes the

bedrock of tomorrow's AI breakthroughs.

About Us

We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.

As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.

You’ll Collaborate With

Silicon, firmware, runtime, datacenter-software and architecture teams to build a mlir-compiler and software stack that brings next-gen AI/NPU hardware into modern datacenter environments and production AI workloads

What You’ll Own
  • Define and implement MLIR dialects and lowering pipelines targeting our AI-accelerator/NPU platform.
  • Lower ML graphs (e.g., from ONNX/PyTorch) into optimized kernels focusing on fusion, tiling and vectorization.
  • Profile, benchmark and optimize compiler output to meet latency, throughput and power targets.
  • Work with hardware architects and runtime/framework teams to co-design compiler features.
  • Build a future‑ready, portable compiler stack that supports multiple hardware generations.
Minimum Qualifications
  • 5+ years of experience in compiler engineering or high‑performance systems.
  • Strong programming skills in C++ (modern standards), and scripting or tooling in Python.
  • Hands‑on experience with MLIR and/or LLVM (designing/extending dialects, lowering, optimization).
  • Deep understanding of compiler internals: IRs, code scheduling, vectorization, loop transformations.
  • Strong familiarity with AI/ML frameworks (e.g.,PyTorch, ONNX, TensorFlow).
  • Excellent understanding of ML operations such as tensor operators, matmuls, quantization, dynamic shapes.
  • Robust understanding of computer architecture and memory hierarchy.
  • Bachelor’s (or higher) degree in Computer Science, Computer Engineering or equivalent
Preferred Qualifications
  • Prior experience targeting NPU/AI accelerator hardware.
  • Contributions to open‑source compiler projects (MLIR, LLVM,Torch-MLIR, ONNX-MLIR).
  • Familiarity with runtime systems, scheduling, resource sharing, memory movement engines and, interconnects.
  • Previous leadership role driving compiler architecture, setting standards, mentoring teams and owning roadmap.
  • Familiarity with software development tooling: Git, CI/CD, debuggers, profilers.
  • Advanced degree (M.S./Ph.D.) preferred.
What Success Looks Like (First 6–9 Months)
  • A production compiler IR/dialect is designed and integrated, successfully lowering key ML workloads to the target NPU.
  • Key operator kernels (matmul, convolution, etc) show measurable performance gains (reduced latency, increased throughput) on the hardware or simulator.
  • The compiler pipeline handles multi‑die/chiplet topology correctly and efficiently schedules across devices.
  • Runtime interfaces and resource scheduling between host and accelerator are functional and validated with real workloads.
  • Telemetry/debug hooks and performance counters from compiled code are available and used for performance analysis by other teams.
  • The architecture and roadmap for next‑gen hardware are defined, and junior engineers are actively mentored within the compiler team.
  • Join us in our mission to democratize AI compute — where your firmware expertise becomes the

bedrock of tomorrow's AI breakthroughs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI20P Compiler Engineer
AI20P Compiler Engineer

Qpisemi • Bengaluru

On-site
INR 2,000,000 - 3,500,000
Staff AI Compiler Engineer
Staff AI Compiler Engineer

EnCharge AI • India

On-site
INR 1,200,000 - 2,400,000
Principal Engineer - NPU Compiler & Architecture
Principal Engineer - NPU Compiler & Architecture

NXP Semiconductors • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Senior AI Compiler Engineer
Senior AI Compiler Engineer

EnCharge AI • India

On-site
INR 2,500,000 - 6,000,000
Principal AI Compiler Engineer
Principal AI Compiler Engineer

Mulya Technologies • India

On-site
INR 5,000,000 - 9,000,000
Principal AI Compiler Engineer
Principal AI Compiler Engineer

EnCharge AI • India

On-site
INR 4,000,000 - 6,000,000
Staff ML Compiler Engineer, TPU Performance Optimizations
Staff ML Compiler Engineer, TPU Performance Optimizations

Google • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Compiler Engineer
AI Compiler Engineer

EnCharge AI • India

On-site
INR 1,800,000 - 2,800,000
Principal AI Compiler Engineer
Principal AI Compiler Engineer

Sptix • India

On-site
INR 3,000,000 - 7,000,000
GPU Compute & MLIR Compiler Engineer
GPU Compute & MLIR Compiler Engineer

BuildxPartners • Bengaluru

On-site
INR 2,000,000 - 3,000,000