AI Reliability Engineer (AI SRE)

DeWinter Group

Campbell (CA)

Remote

USD 68,880 - 241,080

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading firm in AI solutions is looking for an experienced AI Reliability Engineer (AI SRE) for a 12-month remote contract role. In this high-impact position, you will focus on ensuring the reliability and performance of critical AI systems by defining SLOs, implementing automated resilience measures, and leading incident response. Ideal candidates will have over 4 years of SRE experience, expertise in system monitoring, and proficiency in Python and Kubernetes.

Qualifications

  • 4+ years of experience in Site Reliability Engineering (SRE).
  • Deep expertise in system monitoring, incident management, and cloud resilience.
  • Demonstrated ability to work autonomously and manage your own time effectively to meet project goals.

Responsibilities

  • Defining and maintaining Service Level Objectives (SLOs) for AI inference latency and availability.
  • Building automated 'circuit breakers' and fallback logic.
  • Leading incident response and root-cause analysis (RCA) for AI system failures.
  • Developing stress-testing and chaos engineering scenarios for AI agent swarms.
  • Optimizing the 'cold start' and scaling time for serverless AI functions.

Skills

Site Reliability Engineering
System monitoring
Incident management
Cloud resilience
Python
Go
Kubernetes
Observability stacks

Tools

Datadog
New Relic

Job description

Title: AI Reliability Engineer (AI SRE)

Job Type: Contract

Contract Length: 12 Months

Pay Range: $50/hr – $175/hr

Start Date: ASAP

Location: Remote

About the Opportunity:

Our client, a leader in AI testing and Generative AI solutions, is looking for a skilled AI Reliability Engineer (AI SRE) to join their team for a 12-month engagement. This project involves ensuringthe reliability, availability, and performance of mission‑critical AI systems by defining SLOs, implementing automated resilience measures, and leading incident response. This is a high‑impact role that requires a self‑motivated professional who can hit the ground running and deliver results quickly.

Key Responsibilities & Deliverables:
  • Defining and maintaining Service Level Objectives (SLOs) for AI inference latency and availability.
  • Building automated "circuit breakers" and fallback logic (e.g., switching to a smaller model if the primary fails).
  • Leading incident response and root-cause analysis (RCA) for complex AI system failures.
  • Developing stress‑testing and chaos engineering scenarios specifically for AI agent swarms.
  • Optimizing the "cold start" and scaling time for serverless AI functions.
Required Skills & Experience:
  • 4+ years of experience in Site Reliability Engineering (SRE).
  • Deep expertise in system monitoring, incident management, and cloud resilience. This isn't a learning role—you need to be a subject matter expert.
  • Demonstrated ability to work autonomously and manage your own time effectively to meet project goals.
  • Experience with Python/Go, Kubernetes, and observability stacks (Datadog, New Relic).
  • Strong communication skills to provide clear and concise status updates to the project team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Reliability Engineer (SRE) for Gen AI Systems
Remote AI Reliability Engineer (SRE) for Gen AI Systems

DeWinter Group • Campbell (CA)

Remote
AI Site Reliability Engineer (AI SRE)
AI Site Reliability Engineer (AI SRE)

BilgeAdam Technologies GmbH • Town of Poland (NY)

On-site
USD 120,000 - 180,000
Production Systems Engineer – SRE, Monitoring & Root Cause Analysis (AI) | Remote
Production Systems Engineer – SRE, Monitoring & Root Cause Analysis (AI) | Remote

Crossing Hurdles • United States

Remote
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

asobbi • California (MO)

On-site
USD 170,000 - 220,000
Fully remote (US timezone)
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000