AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. in San Francisco seeks an infrastructure research engineer to design, optimize, and maintain the compute foundations powering large-scale language model training.

You will develop high-performance ML kernels (CUDA, CuTe, Triton) and improve the distributed compute stack for scalable AI systems. You’ll collaborate with researchers and systems architects, prototype kernel implementations, and profile performance across hardware generations to shape numerical and

Qualifications

  • Bachelor’s degree or equivalent experience in a technical field.
  • Strong engineering skills with maintainable code and debugging in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their system architectures.
  • Ability to collaborate with cross-functional teams and take initiative across stacks.

Responsibilities

  • Design and implement custom ML kernels (CUDA, CuTe, Triton) for core LLM operations such as attention and matrix multiplication.
  • Design and optimize compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency.
  • Collaborate with research teams to align kernel optimizations with model architecture goals.
  • Develop and maintain a library of reusable kernels and performance benchmarks.
  • Contribute to infrastructure stability, reproducibility, and high compute utilization.
  • Document insights via talks, papers, or open-source contributions.

Skills

Kernel development
Performance profiling
GPU programming
Collaboration
Strong coding skills
Deep learning systems

Education

Bachelor’s degree in CS/EE/related

Tools

CUDA
CuTe
Triton
PyTorch/JAX

Job description

Thinking Machines Lab Inc. in San Francisco seeks an infrastructure research engineer to design, optimize, and maintain the compute foundations powering large-scale language model training.

You will develop high-performance ML kernels (CUDA, CuTe, Triton) and improve the distributed compute stack for scalable AI systems. You’ll collaborate with researchers and systems architects, prototype kernel implementations, and profile performance across hardware generations to shape numerical and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Kernel Engineer for Scalable AI Training
Infrastructure Kernel Engineer for Scalable AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Infrastructure Research Engineer - Distributed AI Training
Infrastructure Research Engineer - Distributed AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Engineer: Scalable AI Research Infrastructure
ML Engineer: Scalable AI Research Infrastructure

Tilde Research • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Infrastructure Research Engineer: Scalable ML Training
Infrastructure Research Engineer: Scalable ML Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Infra Research Engineer: Distributed AI Training Systems
Infra Research Engineer: Distributed AI Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1