Lead ML Infra Architect for High-Performance NLP

Cohere

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
Parental leave top-up
Enrichment benefits & learning stipend
6 weeks paid vacation
Travel budget to other offices / off‑s
Home office stipend

Job summary

Cohere is seeking a Lead Member of Technical Staff to drive the Model Serving platform, delivering low-latency NLP model deployments via scalable, reliable APIs. You will mentor engineers, define architecture direction, and coordinate across teams to meet customers’ needs.

The role focuses on leading design and deployment of high-performance, distributed ML infrastructure with strong emphasis on Kubernetes, cloud platforms, and accelerator technologies.

Qualifications

  • 8+ years of engineering leadership running production infrastructure at scale.
  • Architecture and design of large, highly available distributed systems on Kubernetes with GPU workloads.
  • Deep expertise with Kubernetes development and production standards.
  • Experience across multi-cloud (GCP, Azure, AWS, OCI) and on-prem/hybrid environments.
  • Leading design, deployment, and troubleshooting of Linux-based computing environments at scale.
  • Ownership of compute/storage/network resources and cost management at an organizational level.
  • Strong collaboration and mentoring skills across cross-functional teams.
  • GPGPU/accelerator knowledge to optimize latency and throughput at scale.
  • Proficiency in high-performance servers using Golang and C++.

Skills

Engineering leadership
Kubernetes
Distributed systems
GPU workloads
Cloud platforms
Linux
Cost management
Mentoring
Adaptability
GPUs (accelerators)
C++
Golang
High-performance servers

Job description

Cohere is seeking a Lead Member of Technical Staff to drive the Model Serving platform, delivering low-latency NLP model deployments via scalable, reliable APIs. You will mentor engineers, define architecture direction, and coordinate across teams to meet customers’ needs.

The role focuses on leading design and deployment of high-performance, distributed ML infrastructure with strong emphasis on Kubernetes, cloud platforms, and accelerator technologies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Infrastructure Engineer (Kubernetes + GPUs)
Lead ML Infrastructure Engineer (Kubernetes + GPUs)

Cohere • California (MO)

Hybrid
USD 180,000 - 260,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Software Engineer — ML Platform & Inference
Staff Software Engineer — ML Platform & Inference

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 280,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Senior Lead, Scalable ML Inference Infrastructure
Senior Lead, Scalable ML Inference Infrastructure

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Lead ML Infrastructure & Evaluation
Lead ML Infrastructure & Evaluation

Cursor • New York (NY)

On-site
USD 180,000 - 260,000
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Infra Engineer: Build Scalable ML Services & MLOps
ML Infra Engineer: Build Scalable ML Services & MLOps

Stripe • United States

Hybrid
CAD 172,000 - 258,000
Equity
Retirement plans
Health benefits
+1
Engineering Manager, ML Infrastructure & Scale
Engineering Manager, ML Infrastructure & Scale

Cursor • California (MO)

On-site
USD 180,000 - 260,000
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO