Training Infra Engineer for Scalable AI Models

Cohere Inc.

New York, Northern (NY, KY)

Hybrid

USD 160,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health and dental benefits
RRSP/401K match
Parental leave top-up
Enrichment benefits
Education & learning stipend
Annual offsite
Six weeks paid vacation
Home office stipend
Coworking allowance

Job summary

Cohere is seeking a highly skilled engineer to join our role-focused team, contributing to model training pipelines and production-grade code alongside researchers. You will help push frontier models from research into scalable, production-ready systems with ample compute and data resources.

We welcome remote applicants and offer a global, collaborative environment with strong emphasis on training infrastructure, tooling, and performance improvements across distributed systems.

Qualifications

  • Extremely strong software engineering skills.
  • Proficiency in Python and ML frameworks (JAX, PyTorch, XLA/MLIR).
  • Experience with distributed training infrastructures (Kubernetes, Slurm) and Ray.
  • Experience with large-scale distributed training strategies.
  • Hands-on experience training large models and contributing to training tooling.

Responsibilities

  • Design and write high-performant and scalable software for training.
  • Improve our training setup from an infrastructure and codebase performance standpoint.
  • Craft and implement tools to speed up our training cycles and improve the overall efficacy of our training infrastructure
  • Research, implement, and experiment with ideas on our supercompute and data infrastructure.
  • Learn from and work with the best researchers in the field.

Skills

Strong software engineering
Python & ML frameworks
Distributed training infra
Large-scale training
Research tooling experience
Top-tier paper (bonus)

Tools

Kubernetes
Slurm
Ray

Job description

Cohere is seeking a highly skilled engineer to join our role-focused team, contributing to model training pipelines and production-grade code alongside researchers. You will help push frontier models from research into scalable, production-ready systems with ample compute and data resources.

We welcome remote applicants and offer a global, collaborative environment with strong emphasis on training infrastructure, tooling, and performance improvements across distributed systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Training Infra Engineer for Scalable AI Models
Lead Training Infra Engineer for Scalable AI Models

Cohere • New York (NY)

Remote
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP/401K plan
Staff Engineer, Training Infrastructure & ML Pipelines
Staff Engineer, Training Infrastructure & ML Pipelines

Cohere • United States

Remote
USD 180,000 - 260,000
Lunch stipend
Health & dental benefits
401K / Pension
+2
Staff Training Infra Engineer for Scalable ML Pipelines
Staff Training Infra Engineer for Scalable ML Pipelines

WeHireYou • Paris (IN)

Hybrid
USD 102,000 - 147,000
Lunch stipend
Health & dental benefits
Parental leave
+4
Staff Engineer, Production AI & Scalable LLMs
Staff Engineer, Production AI & Scalable LLMs

cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend
Health benefits
RRSP/401K
+4
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
Staff Engineer - AI Systems & Post-Training RL
Staff Engineer - AI Systems & Post-Training RL

Cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend and health benefits
Dental benefits
RRSP/401K/Pension
+6
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000