Senior SRE — MLOps for Scalable AI Infrastructure

Tiger Analytics, LLC

Washington (District of Columbia)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career development opportunities

Job summary

Tiger Analytics, LLC is seeking a Site Reliability Engineer (SRE) in Washington, D.C. to ensure the reliability and performance of complex AI platforms. This hybrid role combines software engineering and systems architecture with a strong focus on MLOps. Key responsibilities include managing SLAs, optimizing AI infrastructure, automating tasks, and incident response. The position offers significant career growth opportunities in a fast-paced environment.

Qualifications

  • Expert-level knowledge of Kubernetes and Docker.
  • Strong proficiency in Python for automation and scripting.
  • Familiarity with MLOps tools like Kubeflow and Vertex AI.

Responsibilities

  • Define and maintain Service Level Objectives and Indicators.
  • Manage auto-scaling strategies for Kubernetes infrastructure.
  • Ensure high availability of Vertex AI endpoints and services.

Skills

Kubernetes (GKE)
Terraform
Python
MLOps
CI/CD

Tools

Prometheus
Grafana
Vertex AI
Kubeflow
Google Cloud Operations Suite

Job description

Tiger Analytics, LLC is seeking a Site Reliability Engineer (SRE) in Washington, D.C. to ensure the reliability and performance of complex AI platforms. This hybrid role combines software engineering and systems architecture with a strong focus on MLOps. Key responsibilities include managing SLAs, optimizing AI infrastructure, automating tasks, and incident response. The position offers significant career growth opportunities in a fast-paced environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for AI Platform Reliability & MLOps
Senior SRE for AI Platform Reliability & MLOps

Tiger Analytics • Washington

Hybrid
USD 110,000 - 160,000
Senior SRE, MLOps & AI Infra Architect
Senior SRE, MLOps & AI Infra Architect

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior SRE for AI-Driven Healthcare Platform
Senior SRE for AI-Driven Healthcare Platform

Oracle • United States

Remote
USD 79,000 - 159,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid parental leave
+2
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Principal SRE - AI-Driven Platform & AIOps
Principal SRE - AI-Driven Platform & AIOps

Oracle • United States

Remote
USD 86,000 - 200,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Flexible Vacation
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth