Accelerator Compiler and Tool Chain Lead

Velaura

Santa Clara (CA)

On-site

USD 200,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity participation
Performance-based incentives
Medical, dental, and vision coverage
Paid time off
Flexible work arrangements
Professional development opportunities

Job summary

Velaura is seeking an Accelerator Compiler Lead to own the compiler and model-lowering stack for our AI accelerator. You will drive the path from customer AI models to optimized executable artifacts for our NPU, including graph import, operator lowering, IR transformations, and quantization integration.

You will lead a team of compiler and ML systems engineers, partner with hardware and firmware teams, and shape compiler-visible performance features to deliver scalable, production-grade software

Qualifications

  • Deep experience with compiler development, ML graph compilers, or code generation for accelerators.
  • Strong understanding of ML model formats, graph IRs, operator lowering, tensor layouts, quantization, and runtime/compiler interfaces.
  • Strong C++ and Python programming skills and experience building production-quality compiler or systems software.
  • Experience with compiler frameworks such as MLIR, LLVM, TVM, XLA, IREE, Glow, TensorRT-like or OpenVINO-like systems.
  • Strong understanding of correctness risks in compiler optimizations, graph rewrites, mixed precision, operator fusion, and hardware-specific lowering.
  • Ability to work closely with hardware architects, firmware engineers, runtime engineers, model-integration teams, and SQA.
  • Experience leading technical teams or major architecture areas.

Responsibilities

  • Lead architecture and development of the AI accelerator compiler stack.
  • Own model ingestion and graph lowering from frameworks and exchange formats such as PyTorch export flows, ONNX, TensorFlow Lite, or similar.
  • Define operator coverage strategy, lowering rules, graph transformations, fusion, partitioning, and fallback behavior.
  • Develop compiler optimization passes for tensor layout, tiling, memory movement, mixed precision, operator fusion, and hardware-specific scheduling.
  • Work closely with accelerator runtime and driver teams to define executable artifact formats, metadata, memory planning requirements, profiling hooks, and runtime constraints.
  • Partner with hardware architecture and NPU firmware teams on ISA, command streams, tensor layouts, data movement, hardware constraints, and compiler-visible performance features.
  • Own quantization compiler integration, including calibration metadata, precision selection, scale handling, layout constraints, and accuracy/performance tradeoffs.
  • Build compiler diagnostics that help customers understand unsupported operators, shape constraints, graph rewrites, quantization issues, and performance bottlenecks.
  • Establish compiler verification and regression strategy for graph transformations,IR lowering, numerical behavior, model accuracy, and performance.
  • Hire, mentor, and lead a team of compiler and ML systems engineers.

Skills

Accelerator compiler development
ML graph compiler expertise
C++ and Python
Production-grade software
Team leadership

Tools

MLIR
LLVM
TVM
XLA
IREE
OpenVINO-like

Job description

Role Overview

We are looking for an Accelerator Compiler Lead to own the compiler andmodel-lowering stack for Velaura’s AI accelerator.This role will lead the path fromcustomer AI models to optimized executable artifacts for our NPU, including graphimport, operator lowering, compiler IR, graph transformations, quantization integration,code generation, graph partitioning, and compiler diagnostics. The ideal candidate hasbuilt or shipped compiler infrastructure for ML accelerators, GPUs, DSPs, or otherheterogeneous compute targets.

Responsibilities
  • Lead architecture and development of the AI accelerator compiler stack.
  • Own model ingestion and graph lowering from frameworks and exchange formats such as PyTorch export flows, ONNX, TensorFlow Lite, or similar.
  • Define operator coverage strategy, lowering rules, graph transformations, fusion, partitioning, and fallback behavior.
  • Develop compiler optimization passes for tensor layout, tiling, memory movement, mixed precision, operator fusion, and hardware-specific scheduling.
  • Work closely with accelerator runtime and driver teams to define executable artifact formats, metadata, memory planning requirements, profiling hooks, and runtime constraints.
  • Partner with hardware architecture and NPU firmware teams on ISA, command streams, tensor layouts, data movement, hardware constraints, and compiler-visible performance features.
  • Own quantization compiler integration, including calibration metadata, precision selection, scale handling, layout constraints, and accuracy/performance tradeoffs.
  • Build compiler diagnostics that help customers understand unsupported operators, shape constraints, graph rewrites, quantization issues, and performance bottlenecks.
  • Establish compiler verification and regression strategy for graph transformations,IR lowering, numerical behavior, model accuracy, and performance.
  • Hire, mentor, and lead a team of compiler and ML systems engineers.
Required Qualifications
  • Deep experience with compiler development, ML graph compilers, or code generation for accelerators, GPUs, DSPs, or heterogeneous compute systems.
  • Strong understanding of ML model formats, graph IRs, operator lowering, tensorlayouts, quantization, and runtime/compiler interfaces.
  • Strong C++ and Python programming skills and experience buildingproduction-quality compiler or systems software.
  • Experience with compiler frameworks or technologies such as MLIR, LLVM,TVM, XLA, IREE, Glow, TensorRT-like systems, OpenVINO-like systems, orequivalent.
  • Strong understanding of correctness risks in compiler optimizations, graph rewrites, mixed precision, operator fusion, and hardware-specific lowering.
  • Ability to work closely with hardware architects, firmware engineers, runtime engineers, model-integration teams, and SQA.
  • Experience leading technical teams or major architecture areas.
Preferred Qualifications
  • Experience with NPU, GPU, DSP, or AI accelerator compiler stacks.
  • Experience with quantization-aware compilation, mixed precision, sparsity, pruning, graph partitioning, or hardware-specific scheduling.
  • Experience supporting ONNX, PyTorch export, TensorFlow Lite, JAX/XLA,TorchDynamo/TorchInductor, or other model import flows.
  • Familiarity with robotics, computer vision, CNNs, transformers, detection,segmentation, depth, SLAM-adjacent perception, or edge AI workloads.
  • Experience building customer-facing compiler diagnostics and model-portingtools.
  • Experience with model-zoo release processes, accuracy validation, and reproducible benchmark artifacts.
  • Open-source compiler contributions or experience working with externalframework communities.

$200,000 - $500,000 a year

Compensation & Benefits

At Velaura, we believe exceptional talent deserves exceptional rewards. Compensationfor this role includes a competitive base salary, performance-based incentives, andequity participation, allowing team members to share in the company’s long-termsuccess.

Your base pay will depend on your skills, qualifications, experience, and location.

In addition to compensation, Velaura offers a comprehensive benefits package that may include medical, dental, and vision coverage; paid time off; flexible work arrangements; professional development opportunities; and other benefits designed to support the well-being and growth of our team.

Velaura is committed to pay equity and transparency and regularly benchmarks compensation to ensure we remain competitive in the market.

Why Velaura?

Velaura is building next-generation compute technology for cloud, edge, and PhysicalAI. Our solutions will enable robots, autonomous systems, drones, and other intelligentmachines to operate efficiently in the physical world. This is an opportunity to help build foundational technology at a time when the industry is undergoing fundamental change. You will work alongside experienced leaders, architects, engineers, and operators who have delivered industry-defining products across mobile, cloud, semiconductor, and AI platforms. If you enjoy solving difficult problems, working across disciplines, and helping shape thefuture of Physical AI, we would love to hear from you.

Equal Employment Opportunity and Accommodations

Velaura is an Equal Opportunity Employer that is committed to inclusion and diversity.Qualified applicants will receive consideration for employment without regard to race,color, religion, national origin, gender, sexual orientation, gender identity, disability orprotected veteran status. We also take affirmative action to offer employment opportunities to minorities, women, individuals with disabilities, and protected veterans.

Velaura is committed to working with qualified individuals with physical or mentaldisabilities. Applicants who would like to contact us regarding the accessibility of ourwebsite or who need special assistance or a reasonable accommodation for any part of the application or hiring process may contact us at: careers@velaura.ai. This contact information is for accommodation requests only. Evaluation of requests for reasonableaccommodation will be determined on a case-by-case basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI SoC Runtime Software Architect
Principal AI SoC Runtime Software Architect

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Design Verification Engineer- AI Accelerator Lead
Design Verification Engineer- AI Accelerator Lead

Velaura AI, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
Design Verification Engineer- AI Accelerator Lead
Design Verification Engineer- AI Accelerator Lead

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 350,000
Senior RTL Engineer
Senior RTL Engineer

Velaura • California (MO)

On-site
USD 200,000 - 500,000
Senior RTL Engineer
Senior RTL Engineer

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Equity participation
Performance-based incentives
Comprehensive benefits
RTL Lead – Physical AI Compute
RTL Lead – Physical AI Compute

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Medical, dental, and vision coverage
Paid time off
Flexible work arrangements
+2
Power Management Architect, Physical AI SoC
Power Management Architect, Physical AI SoC

Velaura • Santa Clara (CA)

On-site
USD 125,000 - 500,000
Competitive base salary
Equity participation
Benefits package
AI Systems Architect (Models & Hardware Co-Design)
AI Systems Architect (Models & Hardware Co-Design)

Velaura • Boston (MA)

On-site
USD 200,000 - 500,000
Medical, dental, and vision coverage
Paid time off
Flexible work arrangements
+2
Platform Software Lead – Physical AI
Platform Software Lead – Physical AI

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Equity participation
Medical/dental/vision
Paid time off
+1
CAD & Engineering Infrastructure Lead
CAD & Engineering Infrastructure Lead

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Equity participation
Medical, dental, and vision coverage
Paid time off
+2