Senior SRE, MLOps & AI Infra Architect

Tiger Analytics

Washington

Hybrid

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tiger Analytics is seeking a Site Reliability Engineer (SRE) to maintain and ensure the reliability of our production AI platforms. The role requires expertise in Kubernetes, MLOps, and automation tools like Terraform. You will define and monitor service level objectives, architect scalable solutions, and play a crucial part in performance engineering for AI services. This hybrid position provides significant career development opportunities in a fast-growing environment.

Qualifications

  • Experience in reliability and performance engineering, particularly with AI/ML services.
  • Familiarity with automation tools for infrastructure management.
  • Strong background in networking and cloud architectures.

Responsibilities

  • Define, monitor, and maintain SLOs for AI/ML services.
  • Architect auto-scaling strategies for Kubernetes.
  • Ensure high availability of Vertex AI endpoints.

Skills

Kubernetes (GKE)
Python
Terraform
MLOps
CI/CD
Grafana
Prometheus

Tools

Vertex AI
Kubeflow
Google Cloud Operations Suite (Stackdriver)

Job description

Tiger Analytics is seeking a Site Reliability Engineer (SRE) to maintain and ensure the reliability of our production AI platforms. The role requires expertise in Kubernetes, MLOps, and automation tools like Terraform. You will define and monitor service level objectives, architect scalable solutions, and play a crucial part in performance engineering for AI services. This hybrid position provides significant career development opportunities in a fast-growing environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for AI Platform Reliability & MLOps
Senior SRE for AI Platform Reliability & MLOps

Tiger Analytics • Washington

Hybrid
USD 110,000 - 160,000
Senior SRE — MLOps for Scalable AI Infrastructure
Senior SRE — MLOps for Scalable AI Infrastructure

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000