Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health

United States

Hybrid

USD 180,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
Parental leave
Education & learning stipend
6 weeks vacation
Remote office stipend

Job summary

Cohere is seeking a Staff Software Engineer to design and operate ML/HPC infrastructure that powers front-line AI workloads. You will deploy and manage GPU/TPU clusters and scale distributed training across multiple clouds, ensuring reliability and performance for researchers.

You will collaborate closely with AI researchers, build self-service tooling, and promote observability, IaC, and automation to improve efficiency and resilience across the platform.

Qualifications

  • Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and HPC environments.

Responsibilities

  • Build and scale ML-optimized HPC infrastructure across multiple clouds, ensuring high throughput and low latency for AI workloads.

Skills

ML/HPC infra
Kubernetes at scale
Python
Go
Linux internals
RDMA/NCCL
Research collaboration
Self-directed problem solving

Tools

Kubernetes
RDMA
NCCL
PyTorch
JAX
TensorFlow
Go
Python
Linux

Job description

Cohere is seeking a Staff Software Engineer to design and operate ML/HPC infrastructure that powers front-line AI workloads. You will deploy and manage GPU/TPU clusters and scale distributed training across multiple clouds, ensuring reliability and performance for researchers.

You will collaborate closely with AI researchers, build self-service tooling, and promote observability, IaC, and automation to improve efficiency and resilience across the platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, GPU Infrastructure & Platforms
Engineering Manager, GPU Infrastructure & Platforms

cohere • United States

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K matching
+5
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Lead ML Infrastructure Engineer (Kubernetes + GPUs)
Lead ML Infrastructure Engineer (Kubernetes + GPUs)

Cohere • California (MO)

Hybrid
USD 180,000 - 260,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Lead AI Infra Engineer: GPU Compute & ML Pipelines, Equity
Lead AI Infra Engineer: GPU Compute & ML Pipelines, Equity

Fuel Talent • Seattle (WA)

On-site
USD 180,000 - 210,000
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Engineering Manager, GPU Infrastructure - Remote
Engineering Manager, GPU Infrastructure - Remote

Cohere • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health/dental benefits
RRSP matching
+6