GPU Kernel Engineer

Coda Robotics

San Francisco (CA)

On-site

USD 100,000 - 120,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Coda Robotics is looking for an experienced engineer to join their founding team, focusing on low-level compute kernels to enhance robotic foundation models. The ideal candidate will have substantial experience in systems programming (C/C++, assembly), expertise in GPU optimizations, and familiarity with ML framework internals. Responsibilities include leading engineering teams, integrating kernel optimizations, and pioneering high-velocity development culture. Compensation includes a salary range of $100,000 - $120,000 plus equity and benefits.

Qualifications

  • Proven experience in low-level systems programming (C/C++, assembly) targeting CPU and GPU architectures.
  • Expertise in developing and optimizing compute kernels (CUDA, ROCm, OpenCL, SIMD intrinsics).
  • Deep understanding of performance profiling tools (nvprof, perf, Intel VTune).

Responsibilities

  • Lead a team of kernel and system engineers focused on performance-critical code.
  • Design, implement, and optimize custom compute kernels for CPU and GPU.
  • Integrate kernel optimizations into distributed ML frameworks (e.g., PyTorch, TensorFlow).

Skills

Low-level systems programming
CUDA
C/C++
Memory optimization
GPU architectures
Performance profiling tools

Job description

Coda Robotics is scaling the compute infrastructure that powers next‑generation robotic foundation models. As training and inference workloads grow, we need kernel‑level innovations to reduce latency, memory usage, and energy consumption. You will join Coda's founding team to architect and optimize low‑level compute kernels, drivers, and runtime components—making model training and inference significantly cheaper and faster.

Responsibilities
  • Lead a team of kernel and system engineers focused on performance-critical code
  • Design, implement, and optimize custom compute kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators
  • Find bottlenecks in memory hierarchy, thread scheduling, and data movement
  • Integrate kernel optimizations into distributed ML frameworks (e.g., PyTorch, TensorFlow) and orchestrate deployment in cloud and edge environments
  • Explore OS and driver‑level enhancements—such as zero‑copy I/O, custom scheduling, and power management—to further boost throughput
  • Define and own the technical roadmap for kernel and runtime subsystems, balancing performance, maintainability, and portability
  • Drive rapid iteration, testing, and benchmarking cycles to validate improvements and de‑risk large‑scale rollouts
  • Champion a high‑velocity culture that values bold technical ambition, clear accountability, and measurable impact
Requirements
  • Proven experience in low‑level systems programming (C/C++, assembly) targeting CPU and GPU architectures
  • Expertise in developing and optimizing compute kernels (CUDA, ROCm, OpenCL, SIMD intrinsics)
  • Deep understanding of performance profiling tools (nvprof, perf, Intel VTune) and techniques for memory and compute optimization
  • Strong familiarity with ML framework internals (PyTorch, TensorFlow) and integration of custom operations
  • Experience with compiler design or code generation (LLVM, MLIR) is a plus
Compensation

Base salary range: $100,000 - $120,000 per year, plus strong equity and benefits

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead GPU Kernel Engineer for High-Performance ML
Lead GPU Kernel Engineer for High-Performance ML

Coda Robotics • San Francisco (CA)

On-site
USD 100,000 - 120,000
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
+2
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 210,000 - 420,000
Health, dental, and vision benefits
Equity
Relocation support
GPU Kernel Engineer
GPU Kernel Engineer

Sciforium • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Foundation Capital • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture