Senior SRE: AI-Driven Cloud Reliability & Automation

Hidden Jobs

United States

Remote

USD 191,000 - 226,000

Full time

41 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity incentive
Flexible PTO
Health insurance
401(k) with company match
Telehealth

Job summary

Hidden Jobs is seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of cloud infrastructure powering customer-facing products and AI/ML workloads. You will lead incident response, develop automated, monitored processes, and drive cost-efficient, secure operations.

On a Platform Engineering team, you will define SLOs, build observability, and translate requirements into IaC.

Qualifications

  • 4+ years of hands-on experience operating production cloud infrastructure at scale in SRE/DevOps/platform engineering.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS preferred.
  • Strong track record defining SLOs, building monitoring/alerting, and leading incident response.
  • Solid Python or Go fundamentals applied to infrastructure automation; Kubernetes API a plus.
  • Proven ability to drive cloud cost-efficiency and performance optimization across compute/storage/networking.
  • Experience supporting AI/ML or data-intensive workloads; security/compliance aware (HIPAA/SOC 2) is a plus.
  • Familiarity with AI-assisted engineering workflows is a plus.

Responsibilities

  • Own end-to-end reliability, performance, and resilience of AWS and Kubernetes environments and define SLOs.
  • Participate in on-call rotation, lead incident response, and perform blameless post-incident reviews.
  • Build and maintain monitoring, alerting, and observability systems before user impact.
  • Translate high-performance scaling needs into infrastructure-as-code with Terraform.
  • Apply AI tools to reduce operational toil and automate repetitive tasks.
  • Establish deployment/observability standards to enable faster, more reliable feature delivery.
  • Ensure infrastructure meets security and healthcare compliance obligations.

Skills

SRE/DevOps experience
SLOs & monitoring
Python/Go scripting
Cloud cost optimization
Security/compliance
AI/ML workloads

Tools

Kubernetes
Terraform
AWS
Datadog
GitLab
Postgres
NATS
Istio

Job description

Hidden Jobs is seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of cloud infrastructure powering customer-facing products and AI/ML workloads. You will lead incident response, develop automated, monitored processes, and drive cost-efficient, secure operations.

On a Platform Engineering team, you will define SLOs, build observability, and translate requirements into IaC.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

NDEAVOUR CONSULTING • United States

Hybrid
USD 120,000 - 150,000
Remote Office
Parking Space
Fun Office Space
+7
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

1611 Ally Bank • Charlotte (NC)

On-site
USD 110,000 - 180,000
Annual incentive plan
Relocation assistance
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Sight Machine • United States

Hybrid
USD 170,000 - 250,000
Hybrid work flexibility
Catered Lunches, Snacks and Beverages
Commuter Savings Program
+2
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1