Site Reliability Engineer, Inference Infrastructure

Cohere

New York (NY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
RRSP matching
Parental leave
Education stipend
Vacation

Job summary

Cohere is seeking a Site Reliability Engineer to join the Model Serving team. You will build and operate high-performance, scalable AI platform components delivering large language models through reliable API endpoints.

You will work with cross-functional teams to deploy NLP models to production in low-latency, high-throughput environments, participate in on-call rotations, and help shape the infrastructure roadmap while mentoring teammates.

Qualifications

  • 5+ years of engineering experience running production infrastructure at a large scale.
  • Experience designing large, highly available distributed systems with Kubernetes and GPU workloads.
  • Experience in designing, deploying, supporting, and troubleshooting Linux-based computing environments.

Responsibilities

  • Build self-service systems that automate managing, deploying and operating services.
  • Automate environment observability and resilience. Enable developers to troubleshoot and resolve problems.
  • Hit defined SLOs, including participation in an on-call rotation.

Skills

Kubernetes
Distributed systems
Golang
C++
Linux
Cloud platforms

Tools

Kubernetes operators
Cloud providers (GCP, AWS, Azure, OCI)

Job description

Why this role?

Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.

As a Site Reliability Engineer you will:

  • Build self-service systems that automate managing, deploying and operating services.
  • This includes our custom Kubernetes operators that support language model deployments.
  • Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.
  • Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.
  • Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.
  • Develop our team through knowledge sharing and an active review process.

You may be a good fit if you have:

  • 5+ years of engineering experience running production infrastructure at a large scale
  • Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
  • Experience with Kubernetes dev and production coding and support
  • Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
  • Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments
  • Experience in compute/storage/network resource and cost management
  • Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork
  • The grit and adaptability to solve complex technical challenges that evolve day to day
  • Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.
  • Strong understanding or working experience with distributed systems.
  • Experience in Golang, C++ or other languages designed for high-performance scalable servers).
Full-Time Employees at Cohere enjoy these Perks:
  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
  • Full health and dental benefits, including a separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months, for either parent.
  • Annual enrichment benefits:

    Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.

    Education & learning stipend for conferences, courses, and coaching.

  • 6 weeks of paid vacation (30 working days!)
  • Budget for traveling to other offices if you are remote, plus an annual company offsite.
How and Where We Work:
  • Cohere is remote-friendly. We have offices in Toronto, San Francisco, New York City, London, Paris, Montreal, and more coming soon.
  • For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.
  • For those not near an office: a co-working benefit so you can work alongside others in your city.
  • Everyone receives a $500 home office stipend to set up your workspace properly.

We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Inference Infrastructure
Site Reliability Engineer, Inference Infrastructure

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+4
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Jaide Health • San Francisco (CA)

On-site
USD 130,000 - 170,000
Open and inclusive culture
Weekly lunch stipend and snacks
Full health and dental benefits
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Cohere • New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Lead Member of Technical Staff, Inference Infrastructure
Lead Member of Technical Staff, Inference Infrastructure

Visa Hunt • San Francisco (CA)

Hybrid
USD 210,000 - 320,000
Lunch stipend
Health benefits
RRSP matching
+4
Lead Member of Technical Staff, Inference Infrastructure
Lead Member of Technical Staff, Inference Infrastructure

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Visa Hunt • New York (NY)

Hybrid
USD 180,000 - 230,000
Lunch stipend
Health benefits
RRSP/401K matching
+3
Member of Technical Staff, Training Infra Engineer
Member of Technical Staff, Training Infra Engineer

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Health and dental benefits
RRSP matching
6 weeks paid vacation
+1
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 160,000 - 230,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+4
Data Engineer, Data Foundations
Data Engineer, Data Foundations

Cohere • United States

Hybrid
USD 120,000 - 180,000
Weekly lunch stipend ($75/£75)
Full health and dental benefits
RRSP matching / 401K
+5