Sr. Site Reliability Engineer

Tiger Analytics

Washington (District of Columbia)

Hybrid

USD 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tiger Analytics is seeking a Site Reliability Engineer (SRE) in Washington, DC to ensure the resilience and performance of data-driven AI platforms. The role involves a blend of software engineering and systems architecture, focusing on MLOps to maintain high availability for AI models. Responsibilities include managing SLAs, Kubernetes scaling, and automating operations. Ideal candidates will have expertise in Kubernetes, Docker, and Python. Benefits include significant career development opportunities in a fast-growing environment.

Qualifications

  • Expert-level knowledge of Kubernetes and Docker required.
  • Strong proficiency in Python and Bash for automation.
  • Familiarity with MLOps tools like Kubeflow or Vertex AI is needed.

Responsibilities

  • Ensure production ecosystems are resilient, scalable, and performant.
  • Define and maintain SLOs and SLIs for AI/ML services.
  • Optimize resources for Large Language Models.

Skills

Kubernetes
Docker
Python
Bash
MLOps tools (Kubeflow, Vertex AI)
Data Systems management
Networking (VPCs, Load Balancers)

Tools

Terraform
Prometheus
Grafana
Google Cloud Operations Suite

Job description

Role Overview

We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant. This role is a hybrid of software engineering and systems architecture, with a specialized focus on MLOps—bridging the gap between model development and production-grade reliability.

Key Responsibilities
  • Reliability & Performance Engineering
  • SLA/SLO Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical AI/ML services
  • Error Budgeting: Manage error budgets to balance the velocity of feature releases from the ML team with the stability of the production environment
  • Scalability: Architect and manage auto-scaling strategies for Kubernetes (GKE) to handle fluctuating workloads during model training and high-volume inference
  • Model Serving Reliability: Ensure the high availability of Vertex AI endpoints and custom inference services
  • GPU/TPU Optimization: Monitor and optimize compute resource utilization (accelerators) to ensure cost-efficient performance for Large Language Models (LLMs)
  • Pipeline Resilience: Support and stabilize ML pipelines (Vertex AI Pipelines/Kubeflow) to ensure seamless data flow from ingestion to model retraining
  • Infrastructure as Code (IaC): Use Terraform or Pulumi to provision and manage consistent, version-controlled cloud environments
  • CI/CD & GitOps: Design and optimize robust deployment pipelines for both application code and ML models using GitHub Actions, Cloud Build, or ArgoCD
  • Task Automation: Develop custom Python or Go scripts to automate repetitive operational tasks, self-healing mechanisms, and resource cleanup
  • Observability: Build and manage comprehensive dashboards using Prometheus, Grafana, or Google Cloud Operations Suite (Stackdriver)
  • Incident Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages
  • Blameless Post-Mortems: Conduct deep-dive root cause analysis (RCA) to ensure systemic issues are identified and permanently remediated through code
Requirements
  • Orchestration: Expert-level knowledge of Kubernetes (K8s) and Docker.
  • MLOps Stack: Familiarity with tools such as Kubeflow, Vertex AI, MLflow, or DVC.
  • Scripting: Strong proficiency in Python (for automation) and Bash; knowledge of Go is a plus.
  • Data Systems: Experience managing the reliability of data-heavy services (BigQuery, Pub/Sub, or Vector Databases like Pinecone/Milvus).
  • Networking: Solid understanding of VPCs, Load Balancers, DNS, and secure service mesh (Istio/Anthos).
Benefits

Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, fast-growing, challenging and entrepreneurial environment, with a high degree of individual responsibility.

Equal Employment Opportunity

Tiger Analytics provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, national origin, ancestry, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Paid parental leave
+1
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer (Noida, BLR, India)
Senior Site Reliability Engineer (Noida, BLR, India)

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Machine Learning Engineer (with Vertex AI Experience)
Machine Learning Engineer (with Vertex AI Experience)

Tiger Analytics • United States

On-site
USD 120,000 - 150,000
Career development opportunities
Entrepreneurial environment