Principal AI Compiler Engineer

HTEC

California (MO)

Hybrid

USD 200,000 - 350,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision coverage
401(k) with Company match
Unlimited vacation
Paid holidays
Parental benefits
Employee assistance program

Job summary

HTEC in the United States is seeking a Principal AI Compiler & Runtime Engineer for a hybrid US-based role. You will help define the software foundation for next-gen AI compute platforms, working on compiler tech, AI runtimes, graph optimization, and hardware-aware execution.

The position emphasizes applying AI to solve complex systems, compiler/runtime challenges, and hardware-software co-design, in collaboration with inference, systems, and hardware teams.

Qualifications

  • Strong programming skills in C/C++ and/or Python in Linux environments using common development tools.
  • Solid understanding of machine learning fundamentals and modern AI workloads.
  • Experience applying AI technologies to software engineering, performance optimization, and hardware-aware design challenges.

Responsibilities

  • Build and optimize compiler and runtime infrastructure for modern AI workloads.
  • Enable efficient execution of machine learning models across GPUs, NPUs, TPUs, and custom AI accelerators.
  • Apply modern AI technologies and methodologies to improve software development, system optimization, and software-hardware co-design processes.
  • Collaborate with inference, systems, and hardware teams to improve software-hardware co-design.
  • Investigate and resolve performance bottlenecks through profiling, benchmarking, and system-level analysis.

Skills

C/C++
Python
ML fundamentals
AI optimization

Education

BSc/MSc/PhD in CS/Engineering/Math

Job description

Principal AI Compiler & Runtime Engineer

US (Hybrid)

About the Role

Be part of the team creating the software foundation for next-generation AI compute platforms. In this role, you'll work on compiler technologies, AI runtimes, graph optimization, and hardware-aware execution in close collaboration with inference engineers, ML scientists, and hardware specialists. You'll leverage modern AI technologies and methodologies to accelerate software development, improve software-hardware co-design, and optimize the deployment and execution of AI workloads on next-generation compute platforms.

This position offers the opportunity to contribute to state-of-the-art AI infrastructure, optimize software for emerging AI hardware, and help define how modern machine learning workloads are represented, compiled, and executed at scale.

We are particularly interested in engineers who have applied AI technologies to solve complex systems, compiler, runtime, or hardware challenges, rather than solely using AI as a software productivity tool.

How You'll Contribute
  • Build and optimize compiler and runtime infrastructure for modern AI workloads
  • Enable efficient execution of machine learning models across GPUs, NPUs, TPUs, and custom AI accelerators
  • Apply modern AI technologies and methodologies to improve software development, system optimization, and software-hardware co-design processes
  • Collaborate with inference, systems, and hardware teams to improve software-hardware co-design
  • Investigate and resolve performance bottlenecks through profiling, benchmarking, and system-level analysis
Required Skills
  • BSc, MSc, or PhD in Computer Science, Engineering, Mathematics, or a related discipline
  • Strong programming skills in C/C++ and/or Python in Linux environments using common development tools
  • Solid understanding of machine learning fundamentals and modern AI workloads
  • Experience applying AI technologies to software engineering, performance optimization, and hardware-aware design challenges
Examples of Relevant Backgrounds Include
  • AI compiler frameworks and infrastructure (e.g., Mojo/Max, MLIR, LLVM, XLA, OpenXLA, Triton, Gluon)
  • Compiler optimizations such as operator fusion, graph transformations, scheduling, code generation, and lowering pipelines
  • AI runtimes and execution frameworks (e.g., ONNX Runtime, TensorRT, TVM Runtime, IREE Runtime, XLA)
  • Deep understanding of AI systems, with experience applying AI technologies to solve engineering, performance, systems, or hardware-software optimization challenges
  • Performance optimization of machine learning workloads on GPUs, TPUs, NPUs, DSPs, or custom accelerators
  • Hardware-aware software development and AI accelerator enablement
  • Development of high-performance kernels and operators (e.g., GEMMs, convolutions, attention, normalization, quantization)
  • Distributed AI training or inference systems
  • Model execution frameworks such as Max, PyTorch, TensorFlow, JAX, or ONNX
Nice to Have
  • Experience with Modular (Mojo/Max), OpenXLA, StableHLO, Torch-MLIR, Triton, TVM, or IREE
  • Experience developing software for AI accelerators or machine learning hardware platforms
  • Contributions to open-source projects such as LLVM, MLIR, PyTorch, OpenXLA, Triton, Gluon, or xDSL
  • Experience building software for AI accelerators or AI compute platforms
Location & Travel
  • Hybrid role based in US
  • Ability to travel periodically for collaboration with global teams and stakeholders
  • International travel may be required

Salary Range: $200,000–$350,000 gross USD

Our benefits package for this position includes medical, dental, and vision coverage; 401(k) with Company match; unlimited vacation; paid holidays; paid leave for personal events; parental benefits; and an employee assistance program.

All compensation and benefits are subject to the terms of the applicable plan documents and Company policies, which may be amended or discontinued at the Company's discretion. Eligibility for and the value of any bonus or equity award is not guaranteed.

  • HTEC Group, Inc. is an equal opportunity employer.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Compiler & Runtime Engineer
Principal AI Compiler & Runtime Engineer

HTEC Group • Palo Alto (CA)

Hybrid
USD 200,000 - 350,000
Medical, dental and vision coverage
401(k) with Company match
Unlimited vacation
+3
Principal AI Compiler & Runtime Engineer
Principal AI Compiler & Runtime Engineer

Coley S Home Remodeling • Palo Alto (CA)

Hybrid
USD 200,000 - 350,000
Medical, dental, vision
401(k) with company match
Unlimited vacation
+2
Tech Lead, AI Compiler
Tech Lead, AI Compiler

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 180,000 - 260,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 385,000
Senior AI Compiler & Runtime Architect (Hybrid)
Senior AI Compiler & Runtime Architect (Hybrid)

HTEC • California (MO)

Hybrid
USD 200,000 - 350,000
Medical, dental, and vision coverage
401(k) with Company match
Unlimited vacation
+3
Senior Compiler Engineer - AI
Senior Compiler Engineer - AI

NVIDIA Gruppe • Austin (TX)

On-site
USD 184,000 - 288,000
Equity
Benefits
Staff Compiler Engineer
Staff Compiler Engineer

Neurophos • Sunnyvale (CA), Austin (TX)

On-site
USD 140,000 - 200,000
Health plan premiums
Unlimited PTO
401(k) matching
+3
Compiler Code Gen Engineer
Compiler Code Gen Engineer

Lemurian Labs Inc. • Santa Clara (CA)

On-site
USD 170,000 - 230,000
Equity
Medical/dental/vision
Retirement savings plan
+1
Staff Compiler Engineer
Staff Compiler Engineer

Neurophos Inc. • Town of Texas (WI), Northern (KY)

Hybrid
USD 140,000 - 210,000
Health insurance
Unlimited PTO
401(k) matching & stock options
+1
Senior Software Engineer – AI Compiler & Runtime Infrastructure
Senior Software Engineer – AI Compiler & Runtime Infrastructure

Xcelerium • Irvine (CA)

Hybrid
USD 190,000 - 260,000
Comprehensive medical, dental, and vision coverage
401(k) plan
Paid time off (PTO)
+1