Lead Training Infra Engineer for Scalable AI Models

Cohere

New York (NY)

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health and dental benefits
RRSP/401K plan

Job summary

Cohere is a leading security-first enterprise AI company focused on building foundation models and end-to-end products for real-world business problems. We welcome candidates for a role that blends research and production, with no remote location restrictions and multiple offices worldwide.

The role emphasizes contributing to training pipelines, shipping frontier models to production, and bridging research with scalable production code.

Qualifications

  • Extremely strong software engineering skills.
  • Proficiency in Python and ML frameworks such as JAX, PyTorch and XLA/MLIR.
  • Experience with distributed training infrastructures (Kubernetes, Slurm) and associated frameworks (Ray)
  • Experience using large-scale distributed training strategies.
  • Hands on experience on training large model at scale and having contributed to the tooling and/or setup of the training infrastructure
  • Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP).

Responsibilities

  • Design and write high-performant and scalable software for training.
  • Improve our training setup from an infrastructure and codebase performance standpoint.
  • Craft and implement tools to speed up our training cycles and improve the overall efficacy of our training infrastructure
  • Research, implement, and experiment with ideas on our supercompute and data infrastructure.
  • Learn from and work with the best researchers in the field.

Skills

Software engineering
Python ML
Distributed training
Kubernetes & Slurm
Large-scale training
Research tooling

Tools

JAX
PyTorch
XLA/MLIR
Ray
Kubernetes
Slurm

Job description

Cohere is a leading security-first enterprise AI company focused on building foundation models and end-to-end products for real-world business problems. We welcome candidates for a role that blends research and production, with no remote location restrictions and multiple offices worldwide.

The role emphasizes contributing to training pipelines, shipping frontier models to production, and bridging research with scalable production code.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer - AI Systems & Post-Training RL
Staff Engineer - AI Systems & Post-Training RL

Cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend and health benefits
Dental benefits
RRSP/401K/Pension
+6
Staff Engineer, Training Infrastructure & ML Pipelines
Staff Engineer, Training Infrastructure & ML Pipelines

Cohere • United States

Remote
USD 140,000 - 200,000
Lunch stipend
Health and dental benefits
401K / Pension
+5
Staff Engineer, Production AI & Scalable LLMs
Staff Engineer, Production AI & Scalable LLMs

cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend
Health benefits
RRSP/401K
+4
Staff Training Infra Engineer for Scalable ML Pipelines
Staff Training Infra Engineer for Scalable ML Pipelines

WeHireYou • Paris (IN)

Hybrid
USD 102,000 - 147,000
Lunch stipend
Health & dental benefits
Parental leave
+4
Senior AI Infrastructure Architect for Scalable Training
Senior AI Infrastructure Architect for Scalable Training

LinkedIn • Sunnyvale (CA)

Hybrid
USD 198,000 - 326,000
Staff ML Engineer: Post-Training & Production
Staff ML Engineer: Post-Training & Production

Cohere • New York (NY)

Remote
USD 150,000 - 210,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Senior AI Infra Engineer for Scalable Enterprise Platforms
Senior AI Infra Engineer for Scalable Enterprise Platforms

Seekr • Austin (TX)

Hybrid
USD 180,000 - 260,000
Equity Ownership
Unlimited PTO and holidays
Hybrid work environment (Reston, VA &\
+2
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000