Member of Technical Staff, Training Performance Engineer

Cohere

Greater London

On-site

GBP 70,000 - 90,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Deepstreamtech is seeking a Performance Engineer in Greater London to enhance the performance of their advanced language models. The role demands strong software engineering and machine learning expertise, focusing on optimizing training metrics and developing efficient systems. You will work on CUDA and other tools to drive innovation in natural language processing and collaborate with leading researchers in the field.

Qualifications

  • Extremely strong software engineering skills.
  • Proficiency in Python and related ML frameworks such as JAX, Pytorch.
  • Experience writing kernels for GPUs using CUDA, triton.

Responsibilities

  • Optimize performance of advanced language models and systems.
  • Improve key model training metrics like training throughput.
  • Design high-performant software for training.

Skills

Software engineering skills
Proficiency in Python
Experience with ML frameworks (JAX, Pytorch)
Experience with GPU programming
Experience with distributed training
Familiarity with Transformers
Research paper publication

Job description

Requirements
  • Extremely strong software engineering skills
  • Proficiency in Python and related ML frameworks such as JAX, Pytorch and XLA/MLIR
  • Experience writing kernels for GPUs using CUDA, triton, etc
  • Experience using large-scale distributed training strategies
  • Familiarity with autoregressive sequence models, such as Transformers
  • (Desirable) Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)
  • If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!
What the job involves
  • As a Performance Engineer in the Pre-Training team you will be responsible for optimizing the performance of our advanced language models and systems
  • Their primary focus is on improving key model training metrics, such as training throughput, ensuring high accelerator utilization
  • The team combines expertise in software engineering, machine learning, and low-level kernel design and development to design robust systems and enhance model performance
  • You will work on identifying and removing performance bottlenecks, develop cutting‑edge training and profiling tools to help Cohere's mission of providing efficient and reliable language understanding and generation capabilities and drive innovation in the field of natural language processing
  • Design and write high-performant and scalable software for training
  • Understand architectural modifications and design choices and their effects on training throughput and quality
  • Write low-level CUDA, triton kernels to squeeze every last bit of performance from our accelerators
  • Research, implement, and experiment with ideas on our supercompute and data infrastructure
  • Learn from and work with the best researchers in the field
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Member of Technical Staff, Data Analysis and Evaluation
Member of Technical Staff, Data Analysis and Evaluation

Cohere • Greater London

On-site
GBP 50,000 - 70,000
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Principal ML Performance Engineer (GPU Optimization)
Principal ML Performance Engineer (GPU Optimization)

Sponsor Finder • Boston

On-site
GBP 120,000 - 180,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Member of Technical Staff, Pre-Training Data
Member of Technical Staff, Pre-Training Data

Cohere • City Of London

On-site
GBP 60,000 - 90,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+1