About Modular:
Modular is building the next generation of AI infrastructure, bringing together programming languages, compilers, runtimes, frameworks, and developer tools to make AI development faster, more portable, and more accessible.
At the center of this work are Mojo, our systems programming language designed for the AI era, and MAX, our unified AI platform. Together, they enable developers to build and deploy high-performance AI workloads across CPUs, GPUs, and next-generation AI accelerators.
We're building these technologies in the open and rethinking the software foundations required to make heterogeneous computing reliable, performant, and accessible to developers across an increasingly diverse hardware ecosystem.
About the Role:
As a Driver Engineer, you'll work on the critical systems layer that sits between Modular's compiler and runtime stack and the underlying silicon — the devices, contexts, queues, synchronization primitives, memory allocators, networking, and collectives that every kernel and graph execution depends on.
Your work directly affects how reliably and efficiently MAX and Mojo execute across NVIDIA, AMD, Apple Silicon, and emerging accelerators, and how cleanly kernel and graph authors can target those platforms from Mojo and Python.
You'll tackle systems problems involving asynchronous execution, device and resource management, memory movement, synchronization, communication, error handling, and hardware abstraction. You'll help design interfaces that expose the capabilities of different accelerators without forcing the rest of the software stack to reason about every vendor-specific implementation detail.
As Modular expands across more accelerators and increasingly distributed AI workloads, you'll also help evolve the infrastructure needed for multi-accelerator and multi-node execution, including high-performance communication and collective operations.
This is the team that helps make "it just works on a new accelerator" actually true.
LOCATION: Candidates based in the US or Canada are welcome to apply. You can work from our office in Los Altos, CA or remotely from home.
What you will do:
- Design, implement, and extend core driver abstractions — including Device, Context, Queue, Memory, and Function — across diverse hardware backends.
- Build clean abstractions around vendor driver APIs while exposing hardware-specific functionality where it is necessary for performance, correctness, or advanced capabilities.
- Develop infrastructure for asynchronous execution, including queues, streams, events, synchronization, resource lifetimes, and error propagation.
- Build and improve memory-management infrastructure spanning device allocation, host/device transfers, pinned memory, asynchronous allocation, and other accelerator memory-management capabilities.
- Contribute to multi-accelerator and multi-node communication and collective primitives that underpin large-model inference.
- Work with high-performance interconnect and networking technologies such as NVLink, RDMA, InfiniBand, RoCE, EFA, sockets, and communication libraries such as UCX, where relevant.
- Improve diagnostics, observability, and error reporting across the asynchronous execution stack, turning low-level driver failures into actionable information for kernel and graph authors.
- Debug difficult correctness and reliability issues involving concurrency, synchronization, memory lifetimes, device state, resource leaks, and asynchronous execution.
- Partner closely with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams to design and refine the driver and runtime surfaces they depend on.
- Help ensure driver abstractions remain performant and maintainable as Modular expands to additional accelerator architectures and increasingly complex execution environments.
- Participate actively in architecture discussions, design reviews, code reviews, and collaborative software development to maintain a high engineering bar.
- Contribute directly to our fully open-source project, with your work becoming part of Modular's public codebase.
What you Bring to the table:
- 3+ years of experience developing high-performance, low-latency production systems software in modern C++ (C++17/20).
- Strong systems programming fundamentals, particularly around ownership, object and resource lifetimes, concurrency, multithreading, synchronization, memory management, and parallel execution.
- Hands-on experience with at least one accelerator driver-level API, such as the CUDA Driver API, HIP, Metal, Vulkan compute, or a comparable device programming interface.
- Practical understanding of accelerator concepts such as devices, contexts, streams or queues, events, synchronization, module loading, kernel launch, and host/device interaction.
- Working knowledge of accelerator execution and memory models, including stream ordering, asynchronous execution, hostdevice transfers, and pinned memory.
- Experience designing and maintaining APIs or libraries used by other engineers, with strong instincts around naming, layering, abstraction boundaries, ergonomics, and compatibility.
- Strong debugging skills in asynchronous and concurrent systems, including the ability to investigate race conditions, resource lifetime issues, memory problems, synchronization bugs, and failures that cross software abstraction boundaries.
- Experience using systems debugging tools such as GDB, LLDB, sanitizers, or comparable low-level diagnostic tooling.
- Ability to work effectively in a complex, multi-component software stack where problems may span runtime, driver, operating system, hardware, and higher-level framework boundaries.
- A proactive and collaborative engineering mindset, intellectual curiosity, and an ability to work closely with teams consuming the infrastructure you build.
Helpful, but not required:
- Experience developing asynchronous runtimes, custom memory allocators, or low-level resource-management systems.
- Familiarity with accelerator capabilities such as asynchronous memory allocation, IPC, peer-to-peer access, or unified/managed memory.
- Experience with multi-GPU or multi-accelerator topologies, including NVLink, xGMI, NUMA, or comparable technologies.
- Experience with RDMA-based networking and distributed communication, including InfiniBand, RoCE, EFA, UCX, NIXL, MPI, NVSHMEM, ROC_SHMEM, or OpenSHMEM.
- Exposure to multiple accelerator ecosystems, including AMD ROCm, Apple Metal, or other non-NVIDIA platforms.
- Experience implementing or integrating collective communication primitives for distributed or multi-accelerator workloads.
- Experience with zero-copy tensor interoperability, including DLPack, the CUDA Array Interface, or similar mechanisms.
- Experience building diagnostics, observability, tracing, or reliability infrastructure for asynchronous systems.
- Familiarity with Mojo.
- Recent open-source contributions to systems, compiler, runtime, or AI infrastructure projects such as LLVM, PyTorch, JAX/XLA, TVM, IREE, vLLM, or TensorRT-LLM.
- An advanced degree in Computer Science or a related technical field.
What Modular brings to the table:
- Amazing Team. We are a progressive and agile team with some of the industry's best engineering and product leaders. You'll work alongside engineers across compilers, runtimes, kernels, AI infrastructure, distributed systems, and hardware acceleration.
- Foundational Systems Impact. You'll build the low-level infrastructure that every kernel and graph execution relies on, directly influencing the reliability, performance, and portability of MAX and Mojo.
- Multi-Accelerator Scope. Rather than building for a single vendor or architecture, you'll help create systems abstractions that allow Modular's software stack to operate across an increasingly diverse hardware ecosystem.
- Open-Source Impact. You'll have the opportunity to build publicly, contribute directly to Modular's open-source codebase, and create infrastructure used by developers working across modern AI hardware.
- World-class Benefits. In order to attract the best, we need to offer the best. Premier insurance plans, up to 5% 401k matching, flexible paid time off, and more are available to you! Please note that specific benefit packages may vary based on your location.
- Competitive Compensation. We offer very strong compensation packages and want people to be focused on their best work. We believe in tailoring compensation plans to meet the needs of our workforce.
- Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as other cities. Traveling 2-4 times a year is expected for all roles.
Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and share a purpose to fundamentally change how software is built for AI and accelerated computing.
The estimated base salary range for this role to be performed in the US, regardless of the state, is $167,000.00-$242,000.00 USD.
The estimated base salary range for this role to be performed in Canada, regardless of the province, is $164,000.00-$237,000.00 CAD.
The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of total compensation.
For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply, as we may have openings at lower or higher levels than the one advertised.