AI Accelerator Hardware Enablement Engineer

Modular

United States

Remote

USD 180,000 - 270,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

RSU grants
Healthcare benefits
Team onsites

Job summary

Modular is seeking a Hardware Enablement Engineer to bring up and optimize support for new accelerator architectures across its software stack, from Mojo kernels to model serving. You will work closely with vendor tools and cross-functional teams to enable high performance on emerging hardware.

You will write and optimize Mojo kernels, investigate cross-layer performance issues, and build portability infrastructure while collaborating with hardware partners to ensure robust integration and

Qualifications

  • 5+ years in high-performance computing, compiler engineering, accelerator software, systems engineering, or a closely related domain.
  • Strong C++ skills and experience contributing to complex, multi-component software systems.
  • Hands-on experience with at least one heterogeneous programming model such as CUDA, SYCL, OpenCL, or a comparable accelerator programming environment.
  • Understanding of how AI operators are implemented at a low level, demonstrated through experience writing or modifying GPU kernels, custom operators, accelerator code, or working with ML frameworks such as PyTorch at the C++/systems layer.
  • Working knowledge of hardware concepts such as memory hierarchy, parallel execution, synchronization, data movement, and the relationship between software performance and the underlying architecture.
  • Experience debugging problems that cross software abstraction boundaries and an ability to systematically reason about correctness and performance.

Responsibilities

  • Bring up and validate support for new hardware architectures across the Modular software stack, working across kernels, compiler infrastructure, runtime, graph execution, and model serving.
  • Write, port, and optimize Mojo kernels for novel accelerator architectures, establishing correctness first and iterating toward strong performance.
  • Investigate performance and correctness issues across layers of the stack, including kernel execution, memory movement, compiler lowering, graph decisions, synchronization, and vendor runtime behavior.
  • Develop a detailed understanding of new hardware platforms, including their ISA, execution model, memory hierarchy, synchronization mechanisms, compiler constraints, and vendor toolchains.
  • Help map important AI operators and workloads onto new architectures, understanding how operations such as matrix multiplication, convolution, reductions, and other model primitives execute efficiently on the target hardware.
  • Contribute to portability infrastructure, compiler integration, tooling, testing, and debugging workflows that make it easier to support additional hardware platforms over time.
  • Collaborate directly with hardware vendor engineers to understand platform capabilities, investigate issues, build integration tests, and improve end-to-end hardware/software integration.
  • Benchmark and profile workloads as hardware support matures, identifying bottlenecks and helping move new platforms from initial correctness toward production-quality performance.
  • Share knowledge about emerging architectures with the broader team through technical documentation, demos, design discussions, and engineering write-ups.
  • Participate in company events such as onsites and hackathons while contributing to Modular's collaborative and open engineering culture.

Skills

C++
CUDA
OpenCL
SYCL
GPU kernels
AI hardware
Performance debugging

Education

BS in CS/EE

Tools

MLIR/LLVM

Job description

Modular is seeking a Hardware Enablement Engineer to bring up and optimize support for new accelerator architectures across its software stack, from Mojo kernels to model serving. You will work closely with vendor tools and cross-functional teams to enable high performance on emerging hardware.

You will write and optimize Mojo kernels, investigate cross-layer performance issues, and build portability infrastructure while collaborating with hardware partners to ensure robust integration and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mojo Libraries Engineer — Hardware Enablement & Kernels
Mojo Libraries Engineer — Hardware Enablement & Kernels

Modular • United States

Hybrid
USD 148,000 - 270,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Mojo Hardware Enablement Engineer
Mojo Hardware Enablement Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 148,000 - 270,000
Healthcare coverage
Retirement plans
RSU grants
+5
Engineering Manager - Hardware Bringup (Remote)
Engineering Manager - Hardware Bringup (Remote)

Modular, a Qualcomm company • United States

On-site
USD 216,000 - 372,000
RSU grants
Onsite team events
Paid time off
+3
Remote Mojo Compiler Engineer MLIR/LLVM
Remote Mojo Compiler Engineer MLIR/LLVM

Modular • United States

Remote
USD 180,000 - 270,000
Competitive compensation
RSU grants
Team building events
+1
AI Accelerator Hardware Engineer - Compute Modules
AI Accelerator Hardware Engineer - Compute Modules

Meta • Menlo Park (CA)

On-site
USD 140,000 - 180,000
Mojo Libraries Engineer
Mojo Libraries Engineer

Modular • United States

On-site
USD 148,000 - 270,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Remote Mojo Compiler Backend Lead (IC)
Remote Mojo Compiler Backend Lead (IC)

Modular, a Qualcomm company • United States

Hybrid
USD 216,000 - 372,000
RSU grants
Team onsite events
Excellent benefits
Driver Engineer, Multi-Accelerator AI Runtime
Driver Engineer, Multi-Accelerator AI Runtime

Modular, a Qualcomm company • United States

Hybrid
USD 148,000 - 270,000
RSU grants
Team onsite events in Los Altos, CA
Travel 2–4 times per year
AI Hardware Co-Design Engineer – Accelerate Next-Gen Silicon
AI Hardware Co-Design Engineer – Accelerate Next-Gen Silicon

OpenAI • Seattle (WA)

Hybrid
USD 180,000 - 280,000
Relocation assistance
Hybrid work model
ML Accelerator Performance Tools Intern
ML Accelerator Performance Tools Intern

Etched • San Jose (CA)

On-site
USD 20,000 - 31,000