Lead ML Infrastructure Engineer (Kubernetes + GPUs)

Cohere

California (MO)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health and dental benefits
RRSP matching / 401K
Parental leave top-up
Learning stipend
Paid vacation (6 weeks)
Travel to other offices
Home office stipend ($500)

Job summary

Cohere is seeking a Lead Member of Technical Staff to drive the architecture and deployment of scalable NLP model-serving platforms. You will own the end-to-end design of low-latency, high-throughput API endpoints and guide cross-functional teams to deliver production-grade infrastructure.

You will mentor engineers, set technical direction across multiple teams, and optimize compute/storage networks for cost efficiency while supporting GPU-accelerated workloads in hybrid multi-cloud environments.

Qualifications

  • 8+ years of engineering experience leading production infrastructure at scale.
  • Experience architecting large, highly available distributed systems.
  • Kubernetes experience in dev and production environments.
  • Multi-cloud/ hybrid serving environments (GCP, Azure, AWS, OCI).
  • Proven ability to lead design, deployment and operations of Linux-based computing environments at scale.
  • Resource and cost management at organizational level.
  • Strong collaboration and communication, mentoring engineers.
  • Experience with GPUs/accelerators to improve latency and throughput.

Responsibilities

  • Lead architecture and design of large, distributed systems with Kubernetes and GPU workloads.
  • Own deployment, support, and troubleshooting of the AI platform for production NLP models via APIs.
  • Mentor engineers and drive cross-functional initiatives to build mission-critical systems.
  • Guide infrastructure decisions across multi-cloud and on-prem/hybrid environments.
  • Serve as a primary technical contact for customers for customized deployments.

Skills

Engineering leadership
Distributed systems
Kubernetes
Cloud platforms
Linux systems
Cost optimization
GPU acceleration
Go / C++

Tools

Kubernetes
Golang
C++

Job description

Cohere is seeking a Lead Member of Technical Staff to drive the architecture and deployment of scalable NLP model-serving platforms. You will own the end-to-end design of low-latency, high-throughput API endpoints and guide cross-functional teams to deliver production-grade infrastructure.

You will mentor engineers, set technical direction across multiple teams, and optimize compute/storage networks for cost efficiency while supporting GPU-accelerated workloads in hybrid multi-cloud environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Infra Architect for High-Performance NLP
Lead ML Infra Architect for High-Performance NLP

Cohere • New York (NY)

On-site
USD 180,000 - 240,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior ML Infra Platform Engineer — Kubernetes & GPUs
Senior ML Infra Platform Engineer — Kubernetes & GPUs

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Senior GPU HPC Infrastructure Engineer for ML Pipelines
Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health • United States

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Health and dental benefits
Parental leave
+3
Engineering Manager, GPU Infrastructure & Platforms
Engineering Manager, GPU Infrastructure & Platforms

cohere • United States

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K matching
+5
ML Infra Tech Lead: High-Performance GPU & K8s
ML Infra Tech Lead: High-Performance GPU & K8s

Reducto • Santa Fe (NM)

On-site
USD 190,000 - 230,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)

Lambda • San Francisco (CA)

Hybrid
USD 314,000 - 465,000
Health, dental, and vision coverage
401k with 2% company match
Wellness stipend
+1
Senior Kubernetes Engineer: GPU AI Infra Platform
Senior Kubernetes Engineer: GPU AI Infra Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 150,000 - 210,000
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000