An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Black Sesame is seeking an ML Accelerator Compiler Developer in San Jose to design and optimize compiler toolchains for ML accelerators, focusing on high performance and portability. You will implement ML-specific optimizations and collaborate with hardware teams to tailor compiler passes for target architectures.
The role involves working with ML frameworks like PyTorch and ONNX, driving operator fusion, quantization-aware compilation, and scheduling, while profiling and debugging across
Home > Jobs > ML Accelerator Compiler Developer
San Jose, US
2025-03-13
Strong proficiency in compiler development, including experience with LLVM, MLIR, TVM, or similar frameworks.
Expertise in Machine Learning model execution, optimization, and deployment.
Strong programming skills in C++, Python, and assembly-level optimizations.
Knowledge of parallel computing, vectorization, and memory hierarchy optimizations.
Familiarity with deep learning frameworks (TensorFlow, PyTorch, ONNX).
Strong analytical skills for performance profiling and debugging.
Experience in graph optimizations, quantization, and code generation.
Knowledge of heterogeneous computing, DSPs, and low-level hardware programming.
Familiarity with AI model deployment and inference optimization techniques.
Background in high-performance computing (HPC).
As an ML Accelerator Compiler Developer, you will be responsible for designing and optimizing compilers for Machine Learning (ML) accelerators. You will work on enhancing performance, efficiency, and portability of ML models by developing compiler toolchains, optimizations, and code generation techniques tailored for specialized hardware architectures.
Develop and optimize compiler toolchains for ML accelerators, including front-end parsing, intermediate representation (IR) transformations, and backend code generation.
Implement and enhance ML-specific optimizations such as operator fusion, memory layout transformations, quantization-aware compilation, and scheduling.
Collaborate with hardware architects to co-design compiler optimizations aligned with accelerator capabilities.
Work on ML frameworks (PyTorch, ONNX) to integrate compiler passes for efficient execution on target hardware.
Improve performance through domain-specific optimizations, autotuning, and parallelization techniques.
Debug and analyze performance bottlenecks across software and hardware stacks.
Develop automated testing, benchmarking, and profiling tools for validating compiler optimizations.
Job Location: 2290 N 1st. St. STE100 San Jose, CA 95131