Driver Engineer, AI Accelerator Stack (Remote)

Modular

United States

Hybrid

USD 167,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Open-source project
On-site in Los Altos, CA
Travel 2-4 times/year
Competitive compensation
Great benefits

Job summary

Modular is seeking a Driver Engineer to join its team improving the low-level infrastructure between the compiler and runtime stack across NVIDIA, AMD, Apple Silicon, and other accelerators. You will help design device abstractions, manage asynchronous execution, memory allocation, and cross-platform driver interfaces, enabling high performance and portability for MAX and Mojo.

You will collaborate with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams, contribute to open-source

Qualifications

  • 3+ years of experience developing high-performance, low-latency production systems software in modern C++ (C++17/20).
  • Strong systems programming fundamentals around ownership, lifetimes, concurrency, multithreading, synchronization, memory management, and parallel execution.
  • Hands-on experience with at least one accelerator driver-level API (CUDA Driver API, HIP, Metal, Vulkan compute).
  • Practical understanding of accelerators: devices, contexts, streams/queues, events, synchronization, module loading, kernel launch, and host/device interaction.
  • Working knowledge of accelerator execution and memory models, including stream ordering, asynchronous execution, host-device transfers, and pinned memory.
  • Experience designing and maintaining APIs used by other engineers with good ergonomics and compatibility.
  • Strong debugging skills in asynchronous and concurrent systems; diagnosing race conditions, resource lifetimes, memory issues, and cross-boundary failures.
  • Experience with systems debugging tools such as GDB, LLDB, sanitizers.

Responsibilities

  • Design, implement, and extend core driver abstractions — including Device, Context, Queue, Memory, and Function — across diverse hardware backends.
  • Build clean abstractions around vendor driver APIs while exposing hardware-specific functionality where needed for performance or advanced capabilities.
  • Develop infrastructure for asynchronous execution, including queues, streams, events, synchronization, resource lifetimes, and error propagation.
  • Build and improve memory-management infrastructure spanning device allocation, host/device transfers, pinned memory, asynchronous allocation, and other accelerator memory-management capabilities.
  • Contribute to multi-accelerator and multi-node communication and collective primitives underpinning large-model inference.
  • Work with high-performance interconnect and networking technologies (NVLink, RDMA, InfiniBand, RoCE, EFA, sockets, UCX) where relevant.
  • Improve diagnostics, observability, and error reporting across the asynchronous execution stack.
  • Debug difficult correctness and reliability issues involving concurrency, synchronization, memory lifetimes, device state, resource leaks, and asynchronous execution.
  • Partner with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams to refine driver and runtime surfaces.
  • Ensure driver abstractions remain performant and maintainable as Modular expands to more accelerator architectures.
  • Participate in architecture discussions, code reviews, and collaborative software development to maintain high engineering bar.
  • Contribute to our fully open-source project and public codebase.

Skills

C++17/20
Systems programming
Multithreading
Asynchronous execution
Memory management
Device/Context/Queue concepts
Debugging asynchronous systems

Tools

CUDA Driver API
HIP
Metal
Vulkan compute
GDB
LLDB
Sanitizers

Job description

Modular is seeking a Driver Engineer to join its team improving the low-level infrastructure between the compiler and runtime stack across NVIDIA, AMD, Apple Silicon, and other accelerators. You will help design device abstractions, manage asynchronous execution, memory allocation, and cross-platform driver interfaces, enabling high performance and portability for MAX and Mojo.

You will collaborate with Kernels, Graph Compiler/Runtime, Serving, and Mojo library teams, contribute to open-source

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Driver Engineer - AI Accelerator Systems (Remote)
Driver Engineer - AI Accelerator Systems (Remote)

Modular, a Qualcomm company • United States

Hybrid
USD 167,000 - 242,000
Premier insurance
401k matching
Flexible PTO
+1
Driver Engineer
Driver Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 167,000 - 242,000
Premier insurance
401k matching
Flexible PTO
+1
Hardware Enablement Engineer for AI Accelerators
Hardware Enablement Engineer for AI Accelerators

Modular • Boston (MA)

Hybrid
USD 198,000 - 242,000
401k matching
Premier insurance plans
Flexible paid time off
+1
Engineering Manager - Hardware Bringup (Remote)
Engineering Manager - Hardware Bringup (Remote)

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 261,000 - 330,000
Stock options
Excellent benefits
Team onsite events
Driver Engineer
Driver Engineer

Modular • United States

Hybrid
USD 167,000 - 242,000
Open-source project
On-site in Los Altos, CA
Travel 2-4 times/year
+2
Remote Mojo Compiler Engineer — AI Language & GPUs
Remote Mojo Compiler Engineer — AI Language & GPUs

Modular • United States

Remote
USD 135,000 - 242,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Kernel Driver Architect for AI Accelerators
Kernel Driver Architect for AI Accelerators

The Consensus • San Jose (CA)

On-site
USD 180,000 - 280,000
Medical, dental, and vision benefits
Relocation to San Jose
Housing subsidy
+3
Senior Linux Kernel Driver Engineer – PCIe AI Accelerators
Senior Linux Kernel Driver Engineer – PCIe AI Accelerators

MatX Inc. • Mountain View (CA)

Hybrid
USD 250,000 - 600,000
4 weeks PTO
Holidays & remote time
Medical insurance
+3
Senior AI Graph Compiler Architect (Remote)
Senior AI Graph Compiler Architect (Remote)

Modular • United States

On-site
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Low-Level AI Accelerator Runtime Engineer
Low-Level AI Accelerator Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000