Senior Site Reliability Engineer – Scalable AI Infra

Tavily Inc.

New York (NY)

Hybrid

USD 156,000 - 262,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
Remote work reimbursement up to $85/month
Company-paid short-term and long-term disability insurance

Job summary

Tavily Inc. in New York City is seeking a Senior Site Reliability Engineer to manage Kubernetes clusters and own the full infrastructure. You will improve CI/CD pipelines and ensure systems are reliable and scalable. This role offers the chance to work on real scaling challenges and directly impact AI technologies.

With 5-8 years in DevOps or SRE roles and strong skills in Kubernetes, you will have a significant impact in a dynamic and innovative company.

Qualifications

  • 5-8 years in a DevOps or SRE role, working in production environments.
  • Proven experience designing and operating large-scale, distributed systems.
  • Strong Kubernetes experience in a managed cloud environment.

Responsibilities

  • Managing Kubernetes clusters across multiple environments and regions.
  • Owning infrastructure as code for all resources.
  • Maintaining and improving CI/CD pipelines and GitOps-based deployments.

Skills

Kubernetes
Infrastructure as Code (Terraform)
CI/CD
Large-scale Distributed Systems
Observability Stacks

Job description

Tavily Inc. in New York City is seeking a Senior Site Reliability Engineer to manage Kubernetes clusters and own the full infrastructure. You will improve CI/CD pipelines and ensure systems are reliable and scalable. This role offers the chance to work on real scaling challenges and directly impact AI technologies.

With 5-8 years in DevOps or SRE roles and strong skills in Kubernetes, you will have a significant impact in a dynamic and innovative company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
AI Infrastructure Engineer – Scalable Backend & Real-Time Data
AI Infrastructure Engineer – Scalable Backend & Real-Time Data

Tavily • New York (NY)

On-site
USD 100,000 - 150,000
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Senior Site Reliability Engineer (Agentic Search)
Senior Site Reliability Engineer (Agentic Search)

Tavily Inc. • New York (NY)

Hybrid
USD 156,000 - 262,000
100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
+2
Site Reliability Engineer, Cloud Infra for AI Platform
Site Reliability Engineer, Cloud Infra for AI Platform

Anyscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Lead Site Reliability Engineer — Scalable Infra for ML Pipelines
Lead Site Reliability Engineer — Scalable Infra for ML Pipelines

Treeswift Inc • New York (NY)

Hybrid
USD 160,000 - 215,000