Performance Engineer - ML Training & CUDA Kernels

Cohere

New York (NY)

Hybrid

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Health benefits
RRSP matching
Parental leave
Enrichment benefits
Vacation
Travel budget
Home office stipend

Job summary

Cohere is a leading security‑first AI company focused on scalable language model training and deployment. You will optimize training throughput and accelerator utilization within a multidisciplinary team of engineers and researchers.

The role emphasizes low‑level kernel development (CUDA, Triton) and distributed training strategies to advance model performance across the Cohere platform.

Qualifications

  • Proven strong software engineering skills.
  • Proficient in Python and ML frameworks (JAX, PyTorch, XLA/MLIR).
  • Experience writing GPU kernels with CUDA, Triton.
  • Experience with large-scale distributed training strategies.
  • Familiarity with autoregressive models like Transformers.

Responsibilities

  • Design and implement high-performance software for training scalable models.
  • Analyze architectural implications on throughput and quality.
  • Write low-level CUDA and Triton kernels for accelerator optimization.
  • Research and prototype training infrastructure and tooling.
  • Collaborate with researchers to push model performance.

Skills

Software engineering
Python
JAX
PyTorch
XLA/MLIR
CUDA
Triton
Distributed training
Transformers

Job description

Cohere is a leading security‑first AI company focused on scalable language model training and deployment. You will optimize training throughput and accelerator utilization within a multidisciplinary team of engineers and researchers.

The role emphasizes low‑level kernel development (CUDA, Triton) and distributed training strategies to advance model performance across the Cohere platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Cerebras Systems • United States

On-site
USD 110,000 - 140,000
Non-corporate work culture
Equal opportunity employer
Continuous learning and growth opportunities
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Cerebras • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture
High-Performance AI Research Engineer (CUDA/ML)
High-Performance AI Research Engineer (CUDA/ML)

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship and career development
+1
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Kernel Engineer
Kernel Engineer

Cerebras • Raleigh (NC)

On-site
USD 100,000 - 140,000
Equal opportunity work environment
Continuous learning and support
Diverse team culture
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000