Senior SRE for AI Platform Reliability & MLOps

Tiger Analytics

Washington (District of Columbia)

Hybrid

USD 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tiger Analytics is seeking a Site Reliability Engineer (SRE) in Washington, DC to ensure the resilience and performance of data-driven AI platforms. The role involves a blend of software engineering and systems architecture, focusing on MLOps to maintain high availability for AI models. Responsibilities include managing SLAs, Kubernetes scaling, and automating operations. Ideal candidates will have expertise in Kubernetes, Docker, and Python. Benefits include significant career development opportunities in a fast-growing environment.

Qualifications

  • Expert-level knowledge of Kubernetes and Docker required.
  • Strong proficiency in Python and Bash for automation.
  • Familiarity with MLOps tools like Kubeflow or Vertex AI is needed.

Responsibilities

  • Ensure production ecosystems are resilient, scalable, and performant.
  • Define and maintain SLOs and SLIs for AI/ML services.
  • Optimize resources for Large Language Models.

Skills

Kubernetes
Docker
Python
Bash
MLOps tools (Kubeflow, Vertex AI)
Data Systems management
Networking (VPCs, Load Balancers)

Tools

Terraform
Prometheus
Grafana
Google Cloud Operations Suite

Job description

Tiger Analytics is seeking a Site Reliability Engineer (SRE) in Washington, DC to ensure the resilience and performance of data-driven AI platforms. The role involves a blend of software engineering and systems architecture, focusing on MLOps to maintain high availability for AI models. Responsibilities include managing SLAs, Kubernetes scaling, and automating operations. Ideal candidates will have expertise in Kubernetes, Docker, and Python. Benefits include significant career development opportunities in a fast-growing environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — MLOps for Scalable AI Infrastructure
Senior SRE — MLOps for Scalable AI Infrastructure

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Senior SRE, MLOps & AI Infra Architect
Senior SRE, MLOps & AI Infra Architect

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Senior SRE: AI-Powered Reliability on AWS & Kubernetes
Senior SRE: AI-Powered Reliability on AWS & Kubernetes

BetterUp • Austin (TX)

Hybrid
USD 147,000 - 185,000
SRE: AI Inference Platform & ML Systems
SRE: AI Inference Platform & ML Systems

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000