SRE for AI Platform & ML Inference Infra

Cohere

New York (NY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
RRSP matching
Parental leave
Education stipend
Vacation

Job summary

Cohere is seeking a Site Reliability Engineer to join the Model Serving team. You will build and operate high-performance, scalable AI platform components delivering large language models through reliable API endpoints.

You will work with cross-functional teams to deploy NLP models to production in low-latency, high-throughput environments, participate in on-call rotations, and help shape the infrastructure roadmap while mentoring teammates.

Qualifications

  • 5+ years of engineering experience running production infrastructure at a large scale.
  • Experience designing large, highly available distributed systems with Kubernetes and GPU workloads.
  • Experience in designing, deploying, supporting, and troubleshooting Linux-based computing environments.

Responsibilities

  • Build self-service systems that automate managing, deploying and operating services.
  • Automate environment observability and resilience. Enable developers to troubleshoot and resolve problems.
  • Hit defined SLOs, including participation in an on-call rotation.

Skills

Kubernetes
Distributed systems
Golang
C++
Linux
Cloud platforms

Tools

Kubernetes operators
Cloud providers (GCP, AWS, Azure, OCI)

Job description

Cohere is seeking a Site Reliability Engineer to join the Model Serving team. You will build and operate high-performance, scalable AI platform components delivering large language models through reliable API endpoints.

You will work with cross-functional teams to deploy NLP models to production in low-latency, high-throughput environments, participate in on-call rotations, and help shape the infrastructure roadmap while mentoring teammates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI Inference Platform & ML Systems
SRE: AI Inference Platform & ML Systems

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Senior SRE: AI Inference Platform - Remote
Senior SRE: AI Inference Platform - Remote

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
An open and inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
AI SRE: Scale, Resilience & GPU Inference
AI SRE: Scale, Resilience & GPU Inference

Seekr • Austin (TX)

Hybrid
USD 140,000 - 200,000
Equity Ownership – RSUs
Unlimited PTO
14 paid company holidays
+4
Staff Software Engineer — ML Platform & Inference
Staff Software Engineer — ML Platform & Inference

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 280,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
SRE: Automation & Infra for AI Inference
SRE: Automation & Infra for AI Inference

Cerebras • United States

On-site
USD 125,000 - 170,000
Site Reliability Engineer, Inference Infrastructure
Site Reliability Engineer, Inference Infrastructure

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior AI-Driven SRE for Cloud Reliability
Senior AI-Driven SRE for Cloud Reliability

Cerebras • Mountain View (CA)

Hybrid
USD 100,000 - 150,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment