Software Engineer, Hardware Enablement

Modular

Boston (MA)

Hybrid

USD 198,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401k matching
Premier insurance plans
Flexible paid time off
Team onsites and meetups

Job summary

Modular is seeking a Hardware Enablement Engineer to bring up and optimize support for new accelerator platforms across Mojo kernels, compiler infrastructure, runtime, and MAX model serving. You will write and port Mojo kernels for novel architectures, validate correctness, and collaborate with hardware vendors to ensure portable, high-performance AI workloads.

The role requires 5+ years in HPC/accelerator software, strong C++ skills, and experience with CUDA/SYCL/OpenCL.

Qualifications

  • 5+ years of experience in high-performance computing, compiler engineering, accelerator software, or related field.
  • Strong C++ skills and experience with large multi-component systems.
  • Hands-on experience with heterogeneous programming models (CUDA, SYCL, OpenCL).
  • Understanding AI operators at a low level through GPU kernels, custom operators, or ML frameworks at the C++/systems layer.
  • Knowledge of memory hierarchy, parallel execution, synchronization, and vendor toolchains.
  • Proven debugging skills across software abstraction boundaries with a focus on correctness and performance.
  • Ability to quickly learn new hardware platforms, architecture manuals, and vendor docs.
  • Strong communication and collaboration skills across compiler, runtime, kernel, and hardware teams.
  • Pragmatic mindset prioritizing correctness during bring-up and iterative performance improvement.

Responsibilities

  • Bring up and validate support for new hardware architectures across the Modular stack (kernels, compiler, runtime, graph, model serving).
  • Write, port, and optimize Mojo kernels for new accelerators, ensuring correctness and performance.
  • Investigate performance and correctness issues across stack layers (kernel, memory, lowering, graph, synchronization).
  • Understand new hardware platforms' ISA, memory hierarchy, synchronization, and vendor toolchains.
  • Map AI operators to new architectures ensuring efficient execution of matrix ops, convolutions, and reductions.
  • Contribute to portability, compiler integration, tooling, testing, and debugging workflows for multi-hardware support.
  • Collaborate with hardware vendor engineers to run integration tests and resolve platform issues.
  • Benchmark and profile workloads as support matures to push from correctness to production performance.
  • Share knowledge through documentation, demos, and engineering write-ups.
  • Participate in onsites and hackathons, supporting Modular's open engineering culture.

Skills

C++
CUDA
SYCL
OpenCL
GPU kernels
MLIR/LLVM
Hardware enablement
Debugging
Cross-team collaboration

Tools

Triton
CuTe
CUTLASS

Job description

About Modular:

Modular is building the next generation of AI infrastructure, bringing together programming languages, compilers, runtimes, frameworks, and developer tools to make AI development faster, more portable, and more accessible.

At the center of this work are Mojo, our systems programming language designed for the AI era, and MAX, our unified AI platform. Together, they enable developers to build and deploy high-performance AI workloads across CPUs, GPUs, and next-generation AI accelerators.

As part of our mission to build AI's unified compute layer, we're expanding the Modular software stack across a growing range of hardware platforms. Our goal is to give developers a consistent programming and deployment experience while allowing hardware vendors to bring differentiated architectures into a unified software ecosystem.

About the Role:

ML developers today face significant friction when taking trained models into deployment. The ecosystem remains fragmented, with incomplete and patchwork solutions that often require extensive performance tuning and model-specific optimization.

At Modular, we're building a unified AI platform designed to radically improve how developers build, optimize, and deploy AI models across heterogeneous hardware.

We're looking for a motivated Hardware Enablement Engineer to help bring that vision to new accelerator platforms.

You’ll work across the Modular software stack — from Mojo kernels and compiler infrastructure to the graph compiler, runtime, and MAX model serving — to bring up and optimize support for new hardware architectures.

This is a highly cross-functional systems role. You’ll learn how new accelerators execute computation, move data, synchronize work, and interact with vendor toolchains, then help translate those capabilities into a performant and reliable experience within Modular's stack.

You’ll collaborate closely with engineers across Modular as well as external hardware partners, developing deep expertise in novel architectures while contributing directly to our hardware portability story.

LOCATION: We welcome candidates who are based in and have work authorization in the United Kingdom, Norway, or the United States Eastern Time Zone. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Edinburgh, UK or Boston, MA office, with a minimum of two days per week on-site. Onboarding for new hires is conducted in person.

What you will do:
  • Bring up and validate support for new hardware architectures across the Modular software stack, working across kernels, compiler infrastructure, runtime, graph execution, and model serving.
  • Write, port, and optimize Mojo kernels for novel accelerator architectures, establishing correctness first and iterating toward strong performance.
  • Investigate performance and correctness issues across layers of the stack, including kernel execution, memory movement, compiler lowering, graph decisions, synchronization, and vendor runtime behavior.
  • Develop a detailed understanding of new hardware platforms, including their ISA, execution model, memory hierarchy, synchronization mechanisms, compiler constraints, and vendor toolchains.
  • Help map important AI operators and workloads onto new architectures, understanding how operations such as matrix multiplication, convolution, reductions, and other model primitives execute efficiently on the target hardware.
  • Contribute to portability infrastructure, compiler integration, tooling, testing, and debugging workflows that make it easier to support additional hardware platforms over time.
  • Collaborate directly with hardware vendor engineers to understand platform capabilities, investigate issues, build integration tests, and improve end-to-end hardware/software integration.
  • Benchmark and profile workloads as hardware support matures, identifying bottlenecks and helping move new platforms from initial correctness toward production-quality performance.
  • Share knowledge about emerging architectures with the broader team through technical documentation, demos, design discussions, and engineering write-ups.
  • Participate in company events such as onsites and hackathons while contributing to Modular's collaborative and open engineering culture.
What you Bring to the table:
  • 5+ years of experience in high-performance computing, compiler engineering, accelerator software, systems engineering, or a closely related domain in industry or research.
  • Strong C++ skills and experience contributing to complex, multi-component software systems.
  • Hands‑on experience with at least one heterogeneous programming model such as CUDA, SYCL, OpenCL, or a comparable accelerator programming environment.
  • Understanding of how AI operators are implemented at a low level, demonstrated through experience writing or modifying GPU kernels, custom operators, accelerator code, or working with ML frameworks such as PyTorch at the C++/systems layer.
  • Working knowledge of hardware concepts such as memory hierarchy, parallel execution, synchronization, data movement, and the relationship between software performance and the underlying architecture.
  • Experience debugging problems that cross software abstraction boundaries and an ability to systematically reason about correctness and performance.
  • Ability and enthusiasm to learn unfamiliar hardware platforms quickly, including becoming comfortable reading architecture manuals, ISA documentation, and vendor technical materials.
  • Strong communication and collaboration skills, including the ability to work effectively across compiler, runtime, kernel, and hardware teams.
  • A pragmatic engineering mindset that prioritizes correctness during initial bringup while understanding how to iterate methodically toward performance.
Helpful, but not required:
  • Experience with non‑GPU accelerator architectures, such as DSPs, NPUs, AI ASICs, or other specialized compute hardware.
  • Familiarity with MLIR or LLVM compiler infrastructure.
  • Experience with GPU DSLs or libraries such as Triton, CUTLASS, or CuTe.
  • Previous experience bringing up software support for a new accelerator or hardware platform.
  • Experience working directly with hardware vendors or silicon engineering teams.
  • Experience profiling and optimizing kernels or accelerator workloads using hardware-specific performance tools.
  • Exposure to graph compilers and understanding how model operations are lowered, scheduled, fused, or mapped onto target hardware.
  • Experience with model serving or inference optimization workflows.
  • Familiarity with runtime concepts such as device management, memory allocation, queues, synchronization, or asynchronous execution.
  • Experience working on portability layers that target multiple hardware architectures.
What Modular brings to the table:
  • Amazing Team. We are a progressive and agile team with some of the industry's best engineering and product leaders. You'll work alongside engineers across compilers, kernels, runtimes, AI infrastructure, and hardware acceleration.
  • Unique Technical Scope. You'll have the opportunity to work directly with emerging accelerator architectures and see your work span the full software stack — from low‑level kernels and compiler decisions through production AI workloads.
  • Real Hardware Impact. Rather than optimizing for a single architecture, you’ll help determine how Modular's platform expands to an increasingly diverse ecosystem of CPUs, GPUs, and AI accelerators.
  • World‑class Benefits. In order to attract the best, we need to offer the best. Premier insurance plans, up to 5% 401k matching, flexible paid time off, and more are available to you! Please note that specific benefit packages may vary based on your location.
  • Competitive Compensation. We offer very strong compensation packages and want people to be focused on their best work. We believe in tailoring compensation plans to meet the needs of our workforce.
  • Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as other cities. Traveling 2–4 times a year is expected for all roles.

Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and share a purpose to fundamentally change how software is built for AI and accelerated computing.

The estimated base salary range for this role to be performed in the United States is $198,000–$242,000 USD.

The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job‑related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. Total compensation may also include annual target bonus, equity, and benefits, as applicable.

For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply, as we may have openings at lower or higher levels than the one advertised.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Hardware Enablement
Software Engineer, Hardware Enablement

Modular • United States

Hybrid
USD 198,000 - 242,000
Premier insurance plans
401k matching (up to 5%)
Flexible paid time off
+1
Engineering Manager, Hardware Bringup United States / Canada / Europe · Remote
Engineering Manager, Hardware Bringup United States / Canada / Europe · Remote

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 261,000 - 330,000
Stock options
Excellent benefits
Team onsite events
Driver Engineer
Driver Engineer

Modular • United States

Hybrid
USD 167,000 - 242,000
Open-source project
On-site in Los Altos, CA
Travel 2-4 times/year
+2
Driver Engineer
Driver Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 167,000 - 242,000
Premier insurance
401k matching
Flexible PTO
+1
AI Inference Tools Engineer
AI Inference Tools Engineer

Modular • United States

Hybrid
USD 167,000 - 242,000
Stock options
401k matching (up to 5%)
Flexible PTO
+1
Senior AI Kernel Engineer
Senior AI Kernel Engineer

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Developer Advocate, MAX Inference & Serving
Developer Advocate, MAX Inference & Serving

Modular • United States

Hybrid
USD 150,000 - 225,000
Premier insurance
401k matching
Flexible PTO
+1
Senior Community Engineer
Senior Community Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premium health insurance
401k matching
Flexible PTO
+1
Senior Community Engineer
Senior Community Engineer

Modular • United States

Hybrid
USD 150,000 - 225,000
Stock options
401k matching
Premium insurance
+1
Developer Advocate, Mojo
Developer Advocate, Mojo

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premier insurance plans
Stock options
401k matching (up to 5%)
+1