Sr. Site Reliability Engineer

Tiger Analytics, LLC

Washington (District of Columbia)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career development opportunities

Job summary

Tiger Analytics, LLC is seeking a Site Reliability Engineer (SRE) in Washington, D.C. to ensure the reliability and performance of complex AI platforms. This hybrid role combines software engineering and systems architecture with a strong focus on MLOps. Key responsibilities include managing SLAs, optimizing AI infrastructure, automating tasks, and incident response. The position offers significant career growth opportunities in a fast-paced environment.

Qualifications

  • Expert-level knowledge of Kubernetes and Docker.
  • Strong proficiency in Python for automation and scripting.
  • Familiarity with MLOps tools like Kubeflow and Vertex AI.

Responsibilities

  • Define and maintain Service Level Objectives and Indicators.
  • Manage auto-scaling strategies for Kubernetes infrastructure.
  • Ensure high availability of Vertex AI endpoints and services.

Skills

Kubernetes (GKE)
Terraform
Python
MLOps
CI/CD

Tools

Prometheus
Grafana
Vertex AI
Kubeflow
Google Cloud Operations Suite

Job description

Role Overview

We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant. This role is a hybrid of software engineering and systems architecture, with a specialized focus on MLOps—bridging the gap between model development and production-grade reliability.

Key Responsibilities
1. Reliability & Performance Engineering
  • SLA/SLO Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical AI/ML services.
  • Error Budgeting: Manage error budgets to balance the velocity of feature releases from the ML team with the stability of the production environment.
  • Scalability: Architect and manage auto-scaling strategies for Kubernetes (GKE) to handle fluctuating workloads during model training and high-volume inference.
2. MLOps & AI Infrastructure
  • Model Serving Reliability: Ensure the high availability of Vertex AI endpoints and custom inference services.
  • GPU/TPU Optimization: Monitor and optimize compute resource utilization (accelerators) to ensure cost-efficient performance for Large Language Models (LLMs).
  • Pipeline Resilience: Support and stabilize ML pipelines (Vertex AI Pipelines/Kubeflow) to ensure seamless data flow from ingestion to model retraining.
3. Automation & Orchestration (Eliminating "Toil")
  • Infrastructure as Code (IaC): Use Terraform or Pulumi to provision and manage consistent, version-controlled cloud environments.
  • CI/CD & GitOps: Design and optimize robust deployment pipelines for both application code and ML models using GitHub Actions, Cloud Build, or ArgoCD.
  • Task Automation: Develop custom Python or Go scripts to automate repetitive operational tasks, self-healing mechanisms, and resource cleanup.
4. Monitoring, Alerting & Incident Response
  • Observability: Build and manage comprehensive dashboards using Prometheus, Grafana, or Google Cloud Operations Suite (Stackdriver).
  • Incident Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages.
  • Blameless Post-Mortems: Conduct deep-dive root cause analysis (RCA) to ensure systemic issues are identified and permanently remediated through code.

Orchestration: Expert-level knowledge of Kubernetes (K8s) and Docker.

MLOps Stack: Familiarity with tools such as Kubeflow, Vertex AI, MLflow, or DVC.

Scripting: Strong proficiency in Python (for automation) and Bash; knowledge of Go is a plus.

Data Systems: Experience managing the reliability of data-heavy services (BigQuery, Pub/Sub, or Vector Databases like Pinecone/Milvus).

Networking: Solid understanding of VPCs, Load Balancers, DNS, and secure service mesh (Istio/Anthos).

Benefits

Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, fast-growing, challenging and entrepreneurial environment, with a high degree of individual responsibility.

Tiger Analytics provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, national origin, ancestry, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Paid parental leave
+1
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior SRE, MLOps & AI Infra Architect
Senior SRE, MLOps & AI Infra Architect

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer (Noida, BLR, India)
Senior Site Reliability Engineer (Noida, BLR, India)

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000