AI Kernel Writer

Majestic Labs ai

Los Altos (CA)

On-site

USD 70,000 - 90,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

A cutting-edge AI technology firm in California is seeking an entry-level programmer to design and implement high-performance compute kernels for AI primitives. Responsibilities include optimizing for throughput and memory hierarchy, collaborating with teams, and writing reusable code in C++ and CUDA. Ideal candidates have a Bachelor's or Master's in Computer Science and a strong background in parallel programming and optimization techniques.

Qualifications

  • Entry-level position in AI compute kernel design and implementation.
  • Strong background in optimization and parallel programming required.
  • Hands-on experience profiling and optimizing AI workloads is a plus.

Responsibilities

  • Design and implement high-performance compute kernels for AI primitives.
  • Optimize for throughput, latency, and memory hierarchy.
  • Collaborate with compiler and runtime teams to integrate kernels.

Skills

Parallel programming (CUDA, Triton, SYCL, OpenCL)
C++11 or higher
Performance analysis and parallel debugging
Memory layout and vectorization
Optimization of irregular algorithms

Education

Bachelor's or Master's in Computer Science or related field

Tools

Perfetto
VTune
Tracy
Valgrind
GNU Debugger

Job description

Responsibilities
  • Design and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.
  • Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).
  • Collaborate with compiler and runtime teams to integrate kernels into Triton, PyTorch, or SYCL pipelines.
  • Profile and tune kernels using tools like Perfetto, VTune, Tracy, or custom simulators.
  • Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, etc.) and stochastic rounding.
  • Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.
  • Write clear, well-structured, and reusable code (C++/CUDA/Triton/LLVM MLIR).
Requirements
  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field from a recognized university.
  • Strong background in parallel programming (CUDA, Triton, SYCL, OpenCL, Metal, POSIX Threads, or OpenMP).
  • Experience with optimization of irregular algorithms, such as graph computations or sparse numerical linear algebra, combining high-level data structure design with low-level SIMD and synchronization optimizations.
  • Deep understanding of memory layout, vectorization, thread/block scheduling, and cache behavior.
  • Proficiency in C++11 or higher, with strong knowledge of standard algorithms, data structures, and generic programming paradigms.
  • Experience with code generation for high-performance computations and knowledge of frameworks like BLAS/BLIS/Torch.
  • Skilled in performance analysis and parallel debugging using tools such as Valgrind, GNU Debugger, or CI testing frameworks.
  • Hands‑on experience profiling and optimizing compute or AI workloads (e.g., GEMM, softmax, attention).
  • Solid grasp of numerical stability, precision formats, and mixed precision arithmetic.
  • Collaborative work style with the ability to operate effectively in multicultural, cross-disciplinary environments.
Seniority level
  • Entry level
Employment type
  • Full-time
Job function
  • Marketing, Public Relations, and Writing/Editing
Industries
  • Software Development

Referrals increase your chances of interviewing at Majestic Labs ai by 2x

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Compiler – Runtime Engineer
Principal AI Compiler – Runtime Engineer

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Product Manager - AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US
Product Manager - AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US

Majestic Labs ai • Los Altos (CA)

On-site
USD 150,000 - 210,000
Principal Software Engineer - Kernels
Principal Software Engineer - Kernels

d-Matrix • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Kernel Engineer (Compute / Accelerator)
Kernel Engineer (Compute / Accelerator)

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

AMD • San Jose (CA)

On-site
USD 190,000 - 270,000
Product Manager – AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US
Product Manager – AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US

Majestic Labs • Los Altos (CA)

On-site
USD 150,000 - 190,000
Senior AI Kernel Engineer
Senior AI Kernel Engineer

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Staff AI Performance Engineer
Staff AI Performance Engineer

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000