Senior GPU HPC Infra Engineer for AI Training

Showcify

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health & dental benefits
RRSP/401K
Parental leave
Enrichment benefits
6 weeks paid vacation
Office travel budget & home office

Job summary

Cohere is seeking a Staff Software Engineer to design and operate ML/HPC infrastructure at scale. You will deploy GPU/TPU superclusters, optimize cost and performance, and ensure reliable AI workloads across clouds.

You will collaborate with researchers to translate needs into robust infra, build self-service tools, and drive observability and IaC practices across the org. This requires strong systems programming and cross-team mentorship.

Qualifications

  • Deep expertise in ML/HPC infrastructure and GPU/TPU clusters.
  • Experience deploying distributed training with JAX, PyTorch, or TensorFlow.
  • Ability to design scalable Kubernetes-based AI infra at scale.
  • Proficiency in Python and Go for ML tooling and systems.
  • Strong knowledge of Linux internals and RDMA networking.
  • Track record working with AI researchers to solve infra challenges.
  • Self-directed, proactive problem solving in fast-moving teams.

Responsibilities

  • Build and scale ML-optimized HPC infrastructure and GPU/TPU clusters across clouds.
  • Collaborate with cloud providers to optimize cost, reliability, and performance.
  • Troubleshoot bottlenecks and minimize disruptions to AI/ML workflows.
  • Create self-service tools and interfaces for researchers to monitor and debug jobs.
  • Drive infrastructure innovation for distributed training and ML workloads.
  • Promote observability, automation, and IaC across teams.
  • Share knowledge via code reviews and documentation.

Skills

ML/HPC infra
Kubernetes at scale
Python
Go
Linux internals
RDMA networking
Research collaboration
Problem solving

Tools

JAX
PyTorch
TensorFlow

Job description

Cohere is seeking a Staff Software Engineer to design and operate ML/HPC infrastructure at scale. You will deploy GPU/TPU superclusters, optimize cost and performance, and ensure reliable AI workloads across clouds.

You will collaborate with researchers to translate needs into robust infra, build self-service tools, and drive observability and IaC practices across the org. This requires strong systems programming and cross-team mentorship.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU HPC Infrastructure Engineer for ML Pipelines
Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health • United States

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Health and dental benefits
Parental leave
+3
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Engineering Manager, GPU Infrastructure & Platforms
Engineering Manager, GPU Infrastructure & Platforms

cohere • United States

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K matching
+5
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives