Driver Engineer

Modular

United States

Hybrid

USD 167,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Open-source project
On-site in Los Altos, CA
Travel 2-4 times/year
Competitive compensation
Great benefits

Job summary

Modular is seeking a Driver Engineer to join its team improving the low-level infrastructure between the compiler and runtime stack across NVIDIA, AMD, Apple Silicon, and other accelerators. You will help design device abstractions, manage asynchronous execution, memory allocation, and cross-platform driver interfaces, enabling high performance and portability for MAX and Mojo.

You will collaborate with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams, contribute to open-source

Qualifications

  • 3+ years of experience developing high-performance, low-latency production systems software in modern C++ (C++17/20).
  • Strong systems programming fundamentals around ownership, lifetimes, concurrency, multithreading, synchronization, memory management, and parallel execution.
  • Hands-on experience with at least one accelerator driver-level API (CUDA Driver API, HIP, Metal, Vulkan compute).
  • Practical understanding of accelerators: devices, contexts, streams/queues, events, synchronization, module loading, kernel launch, and host/device interaction.
  • Working knowledge of accelerator execution and memory models, including stream ordering, asynchronous execution, host-device transfers, and pinned memory.
  • Experience designing and maintaining APIs used by other engineers with good ergonomics and compatibility.
  • Strong debugging skills in asynchronous and concurrent systems; diagnosing race conditions, resource lifetimes, memory issues, and cross-boundary failures.
  • Experience with systems debugging tools such as GDB, LLDB, sanitizers.

Responsibilities

  • Design, implement, and extend core driver abstractions — including Device, Context, Queue, Memory, and Function — across diverse hardware backends.
  • Build clean abstractions around vendor driver APIs while exposing hardware-specific functionality where needed for performance or advanced capabilities.
  • Develop infrastructure for asynchronous execution, including queues, streams, events, synchronization, resource lifetimes, and error propagation.
  • Build and improve memory-management infrastructure spanning device allocation, host/device transfers, pinned memory, asynchronous allocation, and other accelerator memory-management capabilities.
  • Contribute to multi-accelerator and multi-node communication and collective primitives underpinning large-model inference.
  • Work with high-performance interconnect and networking technologies (NVLink, RDMA, InfiniBand, RoCE, EFA, sockets, UCX) where relevant.
  • Improve diagnostics, observability, and error reporting across the asynchronous execution stack.
  • Debug difficult correctness and reliability issues involving concurrency, synchronization, memory lifetimes, device state, resource leaks, and asynchronous execution.
  • Partner with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams to refine driver and runtime surfaces.
  • Ensure driver abstractions remain performant and maintainable as Modular expands to more accelerator architectures.
  • Participate in architecture discussions, code reviews, and collaborative software development to maintain high engineering bar.
  • Contribute to our fully open-source project and public codebase.

Skills

C++17/20
Systems programming
Multithreading
Asynchronous execution
Memory management
Device/Context/Queue concepts
Debugging asynchronous systems

Tools

CUDA Driver API
HIP
Metal
Vulkan compute
GDB
LLDB
Sanitizers

Job description

About Modular:

Modular is building the next generation of AI infrastructure, bringing together programming languages, compilers, runtimes, frameworks, and developer tools to make AI development faster, more portable, and more accessible.

At the center of this work are Mojo, our systems programming language designed for the AI era, and MAX, our unified AI platform. Together, they enable developers to build and deploy high-performance AI workloads across CPUs, GPUs, and next-generation AI accelerators.

We're building these technologies in the open and rethinking the software foundations required to make heterogeneous computing reliable, performant, and accessible to developers across an increasingly diverse hardware ecosystem.

About the Role:

As a Driver Engineer, you'll work on the critical systems layer that sits between Modular's compiler and runtime stack and the underlying silicon — the devices, contexts, queues, synchronization primitives, memory allocators, networking, and collectives that every kernel and graph execution depends on.

Your work directly affects how reliably and efficiently MAX and Mojo execute across NVIDIA, AMD, Apple Silicon, and emerging accelerators, and how cleanly kernel and graph authors can target those platforms from Mojo and Python.

You'll tackle systems problems involving asynchronous execution, device and resource management, memory movement, synchronization, communication, error handling, and hardware abstraction. You'll help design interfaces that expose the capabilities of different accelerators without forcing the rest of the software stack to reason about every vendor-specific implementation detail.

As Modular expands across more accelerators and increasingly distributed AI workloads, you'll also help evolve the infrastructure needed for multi-accelerator and multi-node execution, including high-performance communication and collective operations.

This is the team that helps make "it just works on a new accelerator" actually true.

LOCATION: Candidates based in the US or Canada are welcome to apply. You can work from our office in Los Altos, CA or remotely from home.

What you will do:
  • Design, implement, and extend core driver abstractions — including Device, Context, Queue, Memory, and Function — across diverse hardware backends.
  • Build clean abstractions around vendor driver APIs while exposing hardware-specific functionality where it is necessary for performance, correctness, or advanced capabilities.
  • Develop infrastructure for asynchronous execution, including queues, streams, events, synchronization, resource lifetimes, and error propagation.
  • Build and improve memory-management infrastructure spanning device allocation, host/device transfers, pinned memory, asynchronous allocation, and other accelerator memory-management capabilities.
  • Contribute to multi-accelerator and multi-node communication and collective primitives that underpin large-model inference.
  • Work with high-performance interconnect and networking technologies such as NVLink, RDMA, InfiniBand, RoCE, EFA, sockets, and communication libraries such as UCX, where relevant.
  • Improve diagnostics, observability, and error reporting across the asynchronous execution stack, turning low-level driver failures into actionable information for kernel and graph authors.
  • Debug difficult correctness and reliability issues involving concurrency, synchronization, memory lifetimes, device state, resource leaks, and asynchronous execution.
  • Partner closely with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams to design and refine the driver and runtime surfaces they depend on.
  • Help ensure driver abstractions remain performant and maintainable as Modular expands to additional accelerator architectures and increasingly complex execution environments.
  • Participate actively in architecture discussions, design reviews, code reviews, and collaborative software development to maintain a high engineering bar.
  • Contribute directly to our fully open-source project, with your work becoming part of Modular's public codebase.
What you Bring to the table:
  • 3+ years of experience developing high-performance, low-latency production systems software in modern C++ (C++17/20).
  • Strong systems programming fundamentals, particularly around ownership, object and resource lifetimes, concurrency, multithreading, synchronization, memory management, and parallel execution.
  • Hands-on experience with at least one accelerator driver-level API, such as the CUDA Driver API, HIP, Metal, Vulkan compute, or a comparable device programming interface.
  • Practical understanding of accelerator concepts such as devices, contexts, streams or queues, events, synchronization, module loading, kernel launch, and host/device interaction.
  • Working knowledge of accelerator execution and memory models, including stream ordering, asynchronous execution, hostdevice transfers, and pinned memory.
  • Experience designing and maintaining APIs or libraries used by other engineers, with strong instincts around naming, layering, abstraction boundaries, ergonomics, and compatibility.
  • Strong debugging skills in asynchronous and concurrent systems, including the ability to investigate race conditions, resource lifetime issues, memory problems, synchronization bugs, and failures that cross software abstraction boundaries.
  • Experience using systems debugging tools such as GDB, LLDB, sanitizers, or comparable low-level diagnostic tooling.
  • Ability to work effectively in a complex, multi-component software stack where problems may span runtime, driver, operating system, hardware, and higher-level framework boundaries.
  • A proactive and collaborative engineering mindset, intellectual curiosity, and an ability to work closely with teams consuming the infrastructure you build.
Helpful, but not required:
  • Experience developing asynchronous runtimes, custom memory allocators, or low-level resource-management systems.
  • Familiarity with accelerator capabilities such as asynchronous memory allocation, IPC, peer-to-peer access, or unified/managed memory.
  • Experience with multi-GPU or multi-accelerator topologies, including NVLink, xGMI, NUMA, or comparable technologies.
  • Experience with RDMA-based networking and distributed communication, including InfiniBand, RoCE, EFA, UCX, NIXL, MPI, NVSHMEM, ROC_SHMEM, or OpenSHMEM.
  • Exposure to multiple accelerator ecosystems, including AMD ROCm, Apple Metal, or other non-NVIDIA platforms.
  • Experience implementing or integrating collective communication primitives for distributed or multi-accelerator workloads.
  • Experience with zero-copy tensor interoperability, including DLPack, the CUDA Array Interface, or similar mechanisms.
  • Experience building diagnostics, observability, tracing, or reliability infrastructure for asynchronous systems.
  • Familiarity with Mojo.
  • Recent open-source contributions to systems, compiler, runtime, or AI infrastructure projects such as LLVM, PyTorch, JAX/XLA, TVM, IREE, vLLM, or TensorRT-LLM.
  • An advanced degree in Computer Science or a related technical field.
What Modular brings to the table:
  • Amazing Team. We are a progressive and agile team with some of the industry's best engineering and product leaders. You'll work alongside engineers across compilers, runtimes, kernels, AI infrastructure, distributed systems, and hardware acceleration.
  • Foundational Systems Impact. You'll build the low-level infrastructure that every kernel and graph execution relies on, directly influencing the reliability, performance, and portability of MAX and Mojo.
  • Multi-Accelerator Scope. Rather than building for a single vendor or architecture, you'll help create systems abstractions that allow Modular's software stack to operate across an increasingly diverse hardware ecosystem.
  • Open-Source Impact. You'll have the opportunity to build publicly, contribute directly to Modular's open-source codebase, and create infrastructure used by developers working across modern AI hardware.
  • World-class Benefits. In order to attract the best, we need to offer the best. Premier insurance plans, up to 5% 401k matching, flexible paid time off, and more are available to you! Please note that specific benefit packages may vary based on your location.
  • Competitive Compensation. We offer very strong compensation packages and want people to be focused on their best work. We believe in tailoring compensation plans to meet the needs of our workforce.
  • Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as other cities. Traveling 2-4 times a year is expected for all roles.

Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and share a purpose to fundamentally change how software is built for AI and accelerated computing.

The estimated base salary range for this role to be performed in the US, regardless of the state, is $167,000.00-$242,000.00 USD.

The estimated base salary range for this role to be performed in Canada, regardless of the province, is $164,000.00-$237,000.00 CAD.

The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of total compensation.

For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply, as we may have openings at lower or higher levels than the one advertised.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Driver Engineer
Driver Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 167,000 - 242,000
Premier insurance
401k matching
Flexible PTO
+1
Software Engineer, Hardware Enablement
Software Engineer, Hardware Enablement

Modular • Boston (MA)

Hybrid
USD 198,000 - 242,000
401k matching
Premier insurance plans
Flexible paid time off
+1
Engineering Manager, Hardware Bringup United States / Canada / Europe · Remote
Engineering Manager, Hardware Bringup United States / Canada / Europe · Remote

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 261,000 - 330,000
Stock options
Excellent benefits
Team onsite events
Software Engineer, Hardware Enablement
Software Engineer, Hardware Enablement

Modular • United States

Hybrid
USD 198,000 - 242,000
Premier insurance plans
401k matching (up to 5%)
Flexible paid time off
+1
Developer Advocate, MAX Inference & Serving
Developer Advocate, MAX Inference & Serving

Modular • United States

Hybrid
USD 150,000 - 225,000
Premier insurance
401k matching
Flexible PTO
+1
Senior Community Engineer
Senior Community Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premium health insurance
401k matching
Flexible PTO
+1
Senior Community Engineer
Senior Community Engineer

Modular • United States

Hybrid
USD 150,000 - 225,000
Stock options
401k matching
Premium insurance
+1
Developer Advocate, Mojo
Developer Advocate, Mojo

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premier insurance plans
Stock options
401k matching (up to 5%)
+1
Developer Advocate, Mojo
Developer Advocate, Mojo

Modular • United States

Hybrid
USD 150,000 - 225,000
Stock options
Premier insurance plans
Flexible PTO
Senior AI Kernel Engineer
Senior AI Kernel Engineer

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1