Senior Applied AI SRE: Reliability & Observability Lead

PowerToFly

Tennessee

On-site

USD 120,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Deloitte seeks an Applied AI Site Reliability Engineer III to drive reliability, performance, and cost discipline for cloud-native platforms. You will own SLOs, build robust observability, and automate deployments while operating AI workloads safely at scale.

Join a cross-functional team focused on resilient systems, incident response, and continuous improvement. This role requires hands-on engineering with strong scripting and collaboration across security, risk, and product groups.

Qualifications

  • 5+ years of software engineering and SRE experience in production
  • Experience operating large-scale, cloud-native systems
  • Ability to define and own SLIs, SLOs, and SLAs, and manage incident command
  • Experience with AI/ML workloads in production is a plus
  • Strong automation and scripting capabilities
  • Excellent collaboration with cross-functional teams

Responsibilities

  • Drive reliability, performance, and cost outcomes via SLOs and error budgets
  • Lead design of observability, performance and resilience testing
  • Own admissions to production with automated reliability checks
  • Develop runbooks and automation to meet reliability KPIs
  • Collaborate with security, risk, and product teams to ensure compliance
  • Operate AI/ML workloads in production focusing on drift, cost, and latency

Skills

SRE Experience
Cloud Platforms
Observability
Incident Mgmt
Automation
Cross-functional Collaboration

Education

Bachelor's degree in CS/Software

Tools

Python
Go
Bash
Java
C#/.NET
Kubernetes
Terraform
ArgoCD
CI/CD
OpenTelemetry
Prometheus
Grafana
Datadog
Dynatrace
CloudWatch
Azure Monitor
GCP Ops

Job description

Deloitte seeks an Applied AI Site Reliability Engineer III to drive reliability, performance, and cost discipline for cloud-native platforms. You will own SLOs, build robust observability, and automate deployments while operating AI workloads safely at scale.

Join a cross-functional team focused on resilient systems, incident response, and continuous improvement. This role requires hands-on engineering with strong scripting and collaboration across security, risk, and product groups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Applied AI SRE: Reliability & Observability
Senior Applied AI SRE: Reliability & Observability

Relha LLC • Northern (KY)

Hybrid
USD 140,000 - 190,000
Travel 10%
Reasonable accommodations
Senior Applied AI SRE - Production Reliability & Observability
Senior Applied AI SRE - Production Reliability & Observability

Relha LLC • Northern (KY)

Hybrid
USD 130,000 - 180,000
Lead Applied AI SRE - Reliability, Observability, AI Ops
Lead Applied AI SRE - Reliability, Observability, AI Ops

Relha LLC • Tampa (FL), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Senior SRE, AI Production Reliability & Observability
Senior SRE, AI Production Reliability & Observability

EPAM Systems • Town of Poland (NY)

On-site
USD 140,000 - 210,000
Senior SRE Leader: AI-Driven Reliability & Resilience
Senior SRE Leader: AI-Driven Reliability & Resilience

Jobtailor • Arlington (TX)

On-site
USD 180,000 - 240,000
Senior SRE: Lead Reliability for AI Platform (Hybrid)
Senior SRE: Lead Reliability for AI Platform (Hybrid)

Docebo • United States

Hybrid
USD 140,000 - 210,000
Health benefits
Paid vacation days
Docebo Days
+3
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model