Senior SRE: AI-Driven Cloud Automation & Kubernetes

Attaindata

Chicago (IL)

Hybrid

USD 140,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Attain is seeking a Senior/Staff Site Reliability Engineer for Consumer Apps in Chicago. You’ll design, build, and scale reliable, secure systems with a focus on automation and AI-driven tooling.

You’ll collaborate with backend, frontend, data science, and product teams to improve performance and developer experience while shaping the architecture for future growth. You’ll work on Kubernetes, Istio, cloud platforms, and CI/CD pipelines, building observable, resilient services that support

Qualifications

  • 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP).
  • Demonstrated fluency directing AI coding agents to build, operate, and debug real infrastructure; and robust and experienced judgment on verification of their work.
  • A track record of replacing manual operations with durable automation.
  • Experience with Docker, Kubernetes, Istio or a similar service mesh.
  • Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner.
  • Experience with stream technologies such as Kafka and Amazon Kinesis.
  • Experience with pub/sub technologies such as AWS SNS and Google Pub/Sub.
  • Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run.
  • Experience with infrastructure-as-code tools such as Terraform.
  • Experience with observability tools such as Datadog, Prometheus, and Grafana.
  • Strong computer science and software engineering fundamentals.
  • Experience with SOC2 and PCI Compliance processes and requirements.

Responsibilities

  • Use AI agents as a force multiplier for yourself and others.
  • Create, improve, and maintain internal agentic tools and harnesses.
  • Add automation to both existing and new systems until manual processes disappear.
  • Develop Helm charts for deploying services and jobs in our Kubernetes cluster.
  • Define metrics, network policies, and routing rules for our Istio service mesh.
  • Monitor and maintain our GCP BigQuery, Spanner, and CloudSQL databases.
  • Pipe metrics to our Google-managed Prometheus instance and build out Grafana dashboards and alerts to increase visibility on our systems.
  • Experiment with GCP offerings, 3rd party vendors, AI tooling, and open-source projects to automate and secure day-to-day operations.
  • Pair with engineering leads to instrument and monitor critical functionality.
  • Participate in architecture design and capacity planning discussions to ensure scalable, maintainable, reliable, and secure systems.
  • Build, maintain, and improve our CI/CD pipeline.

Skills

Automation
AI agents
Observability
CI/CD
Security/compliance
Architecture

Tools

Docker
Kubernetes
Istio
MySQL
Google BigQuery
Google Spanner
Kafka
Amazon Kinesis
AWS Lambda
Google Cloud Functions
Google Cloud Run
Terraform
Datadog
Prometheus
Grafana

Job description

Attain is seeking a Senior/Staff Site Reliability Engineer for Consumer Apps in Chicago. You’ll design, build, and scale reliable, secure systems with a focus on automation and AI-driven tooling.

You’ll collaborate with backend, frontend, data science, and product teams to improve performance and developer experience while shaping the architecture for future growth. You’ll work on Kubernetes, Istio, cloud platforms, and CI/CD pipelines, building observable, resilient services that support

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff SRE: AI‑Driven Cloud Automation Leader
Senior/Staff SRE: AI‑Driven Cloud Automation Leader

Attain • Chicago (IL)

Hybrid
USD 150,000 - 210,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Sight Machine • United States

Hybrid
USD 170,000 - 250,000
Hybrid work flexibility
Catered Lunches, Snacks and Beverages
Commuter Savings Program
+2
Sr/Staff Site Reliability Engineer, Consumer Apps Chicago, IL
Sr/Staff Site Reliability Engineer, Consumer Apps Chicago, IL

Attaindata • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Sr/Staff Site Reliability Engineer, Consumer Apps
Sr/Staff Site Reliability Engineer, Consumer Apps

Attain • Chicago (IL)

Hybrid
USD 150,000 - 210,000
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave