Staff Software Engineer, AI Inference Platform

Visa Hunt

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Health & dental benefits
Parental leave top-up
Enrichment benefits
Vacation time
Travel to offices

Job summary

Cohere is seeking a Members of Technical Staff to join the Model Serving team. You will deploy optimized NLP models to production with low latency and high throughput, interfacing with multiple teams and customers to tailor deployments for specific needs.

The role emphasizes building scalable, reliable AI infrastructure on Kubernetes across multi-cloud environments, with GPU workloads and performance tuning in Linux-based systems. Strong collaboration and problem-solving are essential.

Qualifications

  • 5+ years creating production infrastructure at scale.
  • Designing large, highly available distributed systems with Kubernetes and GPUs.
  • Kubernetes dev and prod coding and support experience.
  • Experience with GCP, Azure, AWS, OCI, multi-cloud.
  • Linux-based compute environments: design, deploy, troubleshoot.
  • Compute/storage/network resource and cost management.
  • Strong collaboration and troubleshooting for critical systems.
  • Grit and adaptability for evolving technical challenges.
  • Familiarity with GPUs/TPUs and latency/throughput effects.
  • Strong distributed systems knowledge.
  • Golang or C++ high-performance servers experience.

Responsibilities

  • Deploy and operate scalable AI platform components.
  • Collaborate across teams to tailor deployments for customers.
  • Optimize latency and throughput for NLP models in production.
  • Ensure high availability and reliability in multi-cloud environments.

Skills

Distributed systems
Kubernetes
Golang/C++
Multi-cloud
Linux troubleshooting
Collaboration
Latency optimization
GPU accelerators
GPUs/TPUs
Distributed systems knowledge
High-performance servers

Tools

Kubernetes
GCP
AWS
Azure
OCI
Linux
GPUs

Job description

Cohere is seeking a Members of Technical Staff to join the Model Serving team. You will deploy optimized NLP models to production with low latency and high throughput, interfacing with multiple teams and customers to tailor deployments for specific needs.

The role emphasizes building scalable, reliable AI infrastructure on Kubernetes across multi-cloud environments, with GPU workloads and performance tuning in Linux-based systems. Strong collaboration and problem-solving are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer — Scalable Inference Infra
Staff Software Engineer — Scalable Inference Infra

Cohere • New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
SRE for AI Platform & ML Inference Infra
SRE for AI Platform & ML Inference Infra

Cohere • New York (NY)

Hybrid
USD 140,000 - 200,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Cohere • New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Staff Software Engineer, AI Inference & Kubernetes
Staff Software Engineer, AI Inference & Kubernetes

Coreweave • United States

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Voluntary supplemental life insurance
+14
Staff AI Systems Engineer - Modeling & Research
Staff AI Systems Engineer - Modeling & Research

Cohere • United States

Remote
USD 180,000 - 240,000
Weekly lunch stipend
Full health and dental benefits
RRSP matching, 401K
Staff AI Engineer: Scale Training to Production (Remote)
Staff AI Engineer: Scale Training to Production (Remote)

Jobzhr • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Full health and dental benefits
RRSP matching/401K
+5
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Jaide Health • San Francisco (CA)

On-site
USD 130,000 - 170,000
Open and inclusive culture
Weekly lunch stipend and snacks
Full health and dental benefits
+3
Lead ML Infra Architect for High-Performance NLP
Lead ML Infra Architect for High-Performance NLP

Cohere • New York (NY)

On-site
USD 180,000 - 240,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5