Performance Engineer - ML Training & CUDA Kernels

Cohere

New York (NY)

Hybrid

USD 150,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health benefits
RRSP matching
Parental leave
Enrichment benefits
Vacation
Travel budget
Home office stipend

Job summary

Cohere is a leading security‑first AI company focused on scalable language model training and deployment. You will optimize training throughput and accelerator utilization within a multidisciplinary team of engineers and researchers.

The role emphasizes low‑level kernel development (CUDA, Triton) and distributed training strategies to advance model performance across the Cohere platform.

Qualifications

  • Proven strong software engineering skills.
  • Proficient in Python and ML frameworks (JAX, PyTorch, XLA/MLIR).
  • Experience writing GPU kernels with CUDA, Triton.
  • Experience with large-scale distributed training strategies.
  • Familiarity with autoregressive models like Transformers.

Responsibilities

  • Design and implement high-performance software for training scalable models.
  • Analyze architectural implications on throughput and quality.
  • Write low-level CUDA and Triton kernels for accelerator optimization.
  • Research and prototype training infrastructure and tooling.
  • Collaborate with researchers to push model performance.

Skills

Software engineering
Python
JAX
PyTorch
XLA/MLIR
CUDA
Triton
Distributed training
Transformers

Job description

Cohere is a leading security‑first AI company focused on scalable language model training and deployment. You will optimize training throughput and accelerator utilization within a multidisciplinary team of engineers and researchers.

The role emphasizes low‑level kernel development (CUDA, Triton) and distributed training strategies to advance model performance across the Cohere platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Systems Engineer: Training Frameworks & Tooling
Senior ML Systems Engineer: Training Frameworks & Tooling

Cohere • United States

Remote
USD 150,000 - 230,000
Senior ML Training Frameworks & Tools Engineer
Senior ML Training Frameworks & Tools Engineer

Cohere • New York (NY)

Remote
USD 180,000 - 280,000
Health & dental benefits
Parental leave top-up
6 weeks vacation
Senior Staff Engineer, Model Efficiency - Remote
Senior Staff Engineer, Model Efficiency - Remote

Cohere • United States

Remote
USD 150,000 - 210,000
High-Performance ML Runtime & Kernel Engineer
High-Performance ML Runtime & Kernel Engineer

Cerebras Systems • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Algorithm Mapping and Performance Engineer, Core ML
ML Algorithm Mapping and Performance Engineer, Core ML

Cerebras • United States

Remote
USD 140,000 - 220,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Kernel Performance Engineer - AI Accelerator
Senior ML Kernel Performance Engineer - AI Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Performance Engineer — Hybrid, High-Impact Role
ML Performance Engineer — Hybrid, High-Impact Role

Trading Interview • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 220,000
Hybrid working opportunities
Generous time off
Free meals daily
+2