Principal AI Compiler – Runtime Engineer

Jobtailor

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor, based in Palo Alto, seeks a senior engineer to build and optimize compiler and runtime infrastructure for modern AI workloads. The role focuses on efficient execution across GPUs/TPUs/NPUs and hardware accelerators, applying AI techniques to software optimization and co-design.

The candidate will collaborate with inference, systems, and hardware teams, profiling and benchmarking to remove bottlenecks and drive state-of-the-art AI infrastructure. International travel may be required.

Qualifications

  • BSc, MSc, or PhD in Computer Science, Engineering, Mathematics, or a related discipline
  • Strong programming skills in C/C++ and/or Python in Linux environments using common development tools
  • Solid understanding of machine learning fundamentals and modern AI workloads
  • Experience applying AI technologies to software engineering, performance optimization, and hardware-aware design challenges
  • Experience with AI compiler frameworks and infrastructure such as Mojo/Max, MLIR, LLVM, XLA, OpenXLA, Triton, or Gluon
  • Knowledge of compiler optimizations including operator fusion, graph transformations, scheduling, code generation, and lowering pipelines
  • Experience with AI runtimes and execution frameworks such as ONNX Runtime, TensorRT, TVM Runtime, IREE Runtime, or XLA
  • Experience optimizing machine learning workloads on GPUs, TPUs, NPUs, DSPs, or custom accelerators
  • Experience with hardware-aware software development and AI accelerator enablement
  • Experience developing high-performance kernels and operators such as GEMMs, convolutions, attention, normalization, or quantization
  • Experience with distributed AI training or inference systems
  • Experience with model execution frameworks such as Max, PyTorch, TensorFlow, JAX, or ONNX
  • Ability to travel periodically; international travel may be required
  • Nice-to-have experience with Modular (Mojo/Max), OpenXLA, StableHLO, Torch-MLIR, Triton, TVM, or IREE
  • Nice-to-have contributions to open-source projects such as LLVM, MLIR, PyTorch, OpenXLA, Triton, Gluon, or xDSL

Responsibilities

  • Build and optimize compiler and runtime infrastructure for modern AI workloads
  • Enable efficient execution of machine learning models across GPUs, NPUs, TPUs, and custom AI accelerators
  • Apply modern AI technologies and methodologies to improve software development, system optimization, and software-hardware co-design processes
  • Collaborate with inference, systems, and hardware teams to improve software-hardware co-design
  • Investigate and resolve performance bottlenecks through profiling, benchmarking, and system-level analysis
  • Contribute to state-of-the-art AI infrastructure and optimize software for emerging AI hardware
  • Help define how modern machine learning workloads are represented, compiled, and executed at scale
  • Travel periodically for collaboration with global teams and stakeholders

Skills

C/C++ Programming
Python Programming
AI Compiler Frameworks
Machine Learning Optimization
Performance Profiling

Education

BSc/MSc/PhD in CS/Engineering/Math

Tools

Mojo/Max
MLIR
LLVM
XLA
ONNX Runtime
TensorRT
Triton
Gluon
PyTorch
TensorFlow

Job description

  • Build and optimize compiler and runtime infrastructure for modern AI workloads
  • Enable efficient execution of machine learning models across GPUs, NPUs, TPUs, and custom AI accelerators
  • Apply modern AI technologies and methodologies to improve software development, system optimization, and software-hardware co-design processes
  • Collaborate with inference, systems, and hardware teams to improve software-hardware co-design
  • Investigate and resolve performance bottlenecks through profiling, benchmarking, and system-level analysis
  • Contribute to state-of-the-art AI infrastructure and optimize software for emerging AI hardware
  • Help define how modern machine learning workloads are represented, compiled, and executed at scale
  • Travel periodically for collaboration with global teams and stakeholders
Requirements
  • BSc, MSc, or PhD in Computer Science, Engineering, Mathematics, or a related discipline
  • Strong programming skills in C/C++ and/or Python in Linux environments using common development tools
  • Solid understanding of machine learning fundamentals and modern AI workloads
  • Experience applying AI technologies to software engineering, performance optimization, and hardware-aware design challenges
  • Experience with AI compiler frameworks and infrastructure such as Mojo/Max, MLIR, LLVM, XLA, OpenXLA, Triton, or Gluon
  • Knowledge of compiler optimizations including operator fusion, graph transformations, scheduling, code generation, and lowering pipelines
  • Experience with AI runtimes and execution frameworks such as ONNX Runtime, TensorRT, TVM Runtime, IREE Runtime, or XLA
  • Experience optimizing machine learning workloads on GPUs, TPUs, NPUs, DSPs, or custom accelerators
  • Experience with hardware-aware software development and AI accelerator enablement
  • Experience developing high-performance kernels and operators such as GEMMs, convolutions, attention, normalization, or quantization
  • Experience with distributed AI training or inference systems
  • Experience with model execution frameworks such as Max, PyTorch, TensorFlow, JAX, or ONNX
  • Ability to travel periodically; international travel may be required
  • Nice-to-have experience with Modular (Mojo/Max), OpenXLA, StableHLO, Torch-MLIR, Triton, TVM, or IREE
  • Nice-to-have contributions to open-source projects such as LLVM, MLIR, PyTorch, OpenXLA, Triton, Gluon, or xDSL
Core Competencies

Demonstrates expertise in building and optimizing compiler and runtime infrastructure for AI workloads, with strong programming skills in C/C++ and Python. Proficient in applying AI technologies to enhance software development and performance optimization across various hardware platforms.

Highest-signal resume keywords
  • C/C++ Programming
  • Python Programming
  • AI Compiler Frameworks
  • Machine Learning Optimization
  • Performance Profiling
Hard Skills
  • Compiler Optimization
  • Machine Learning Fundamentals
  • High-Performance Kernels
  • Distributed AI Training
  • Software-Hardware Co-Design
Industry Keywords
  • AI Workloads
  • Custom AI Accelerators
  • Performance Bottlenecks
  • Graph Transformations
  • AI Runtimes
Tools & Technologies
  • LLVM
  • MLIR
  • ONNX Runtime
  • TensorRT
  • PyTorch
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Compiler & Runtime Engineer
Principal AI Compiler & Runtime Engineer

HTEC Group • Palo Alto (CA)

Hybrid
USD 200,000 - 350,000
Medical, dental and vision coverage
401(k) with Company match
Unlimited vacation
+3
AI Compiler Engineer
AI Compiler Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 100,000 - 130,000
Machine Learning Engineer, AI Inference Solutions – Early Career
Machine Learning Engineer, AI Inference Solutions – Early Career

Jobtailor • Sunnyvale (CA)

On-site
USD 110,000 - 160,000
Software Engineer, AI Infrastructure – LVM Inference & Evaluation
Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 240,000
Staff AI Software Engineer
Staff AI Software Engineer

Jobtailor • California (MO)

On-site
USD 170,000 - 210,000
AI Kernel Writer
AI Kernel Writer

Majestic Labs ai • Los Altos (CA)

On-site
USD 70,000 - 90,000
Senior Compiler Engineer - AI
Senior Compiler Engineer - AI

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity
Generous benefits package
Forward Deployed Software Engineer – Advanced
Forward Deployed Software Engineer – Advanced

Jobtailor • Illinois

On-site
USD 150,000 - 190,000
Principal Software Architect
Principal Software Architect

Jobtailor • California (MO)

On-site
USD 150,000 - 230,000
Compiler Engineer
Compiler Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 250,000