Sr/Staff Site Reliability Engineer, Consumer Apps

Attain

United States

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Attain is seeking a Senior/Staff Site Reliability Engineer to design, build, and scale the infrastructure powering its fintech platform. You will own automation, deploy AI-assisted tooling, and drive reliability across cloud-native systems in a fast-growing environment.

You will work with AWS/GCP, Docker/Kubernetes, Istio, and modern observability stacks to reduce toil, improve performance, and enable secure, scalable services for millions of users in the U.S. hybrid/remote-friendly setting.

Qualifications

  • 6+ years building large-scale cloud-native infrastructure on AWS and/or GCP.

Responsibilities

  • Use AI agents to automate infrastructure tasks end-to-end.
  • Define metrics, policies, and routing for Istio service mesh.
  • Develop and maintain CI/CD pipelines and automation tooling.
  • Monitor and optimize cloud resources across GCP/AWS and databases.
  • Collaborate with multiple teams on architecture and capacity planning.

Skills

AI agents
Cloud infrastructure
Docker & Kubernetes
Terraform
Observability tools
SQL databases
Service mesh Istio
Pub/Sub
Serverless

Tools

AWS
GCP
MySQL
Google BigQuery
Google Spanner
Kafka
Kinesis
AWS Lambda
Google Cloud Functions
Google Cloud Run
Terraform
Datadog
Prometheus
Grafana
Docker
Kubernetes
Istio

Job description

About Attain

Built for consumers and companies, alike

Klover’s engineering team powers one of the fastest-growing fintech platforms in the U.S., supporting over one million active users each month. Our systems process and move more than $1.5 billion annually, enabling real-time access to financial tools, rewards, and services that help people improve their day-to-day lives.

As part of this team, you’ll help design, build, and scale the systems that underpin Klover’s core products and platform. You’ll work on high-impact, production-grade systems that prioritize reliability, security, and performance, and that integrate with a broad ecosystem of internal and external services. The work you do will directly shape how users interact with Klover’s products, access their money, and experience transparent, low-fee financial services.

Klover engineers collaborate closely with colleagues across backend, frontend, data science, and product teams to deliver scalable, high-quality solutions for a rapidly growing user base. You’ll have the opportunity to work with modern technologies and architectures while helping define and evolve the next generation of inclusive, data-powered financial products—building systems and interfaces that emphasize reliability, privacy, and performance at scale.

Attain Office Hybrid Schedule (where applicable):

  • _Chicago, IL & New York, NY:_4 days in-office; 1 day remote
  • We are also potentially open to discussing a remote arrangement for the right candidate
About the Role

As a Senior/Staff Site Reliability Engineer, you will play a critical role in building out and maintaining the infrastructure that powers all of our systems, as well as all of the supporting tools to ensure that those systems are running smoothly. Automation is the core of this role. We expect you to hunt down manual toil wherever it hides and engineer it out of existence — and to do that with the best tooling available, including AI agents you direct to write, test, and ship infrastructure code.

We treat fluency with modern AI as a first-class SRE skill, along with cloud platform literacy, a strong dedication to observability, and a laser-focus on improving developer experience. You will work closely with nearly every engineering team at Attain, helping to ensure that our systems are operating at peak efficiency, and preparing us to handle the scale of our future growth.

What a typical week might look like
  • Use AI agents as a force multiplier for yourself and others
  • Create, improve, and maintain internal agentic tools and harnesses
  • Add automation to both existing and new systems until manual processes, and the toil that comes with them, simply go away
  • Write Terraform modules for deploying infrastructure resources via our GitLab pipelines
  • Develop Helm charts for deploying services and jobs in our Kubernetes cluster
  • Define metrics, network policies, and routing rules for our Istio service mesh
  • Monitor and maintain our GCP BigQuery, Spanner, and CloudSQL databases
  • Pipe metrics to our Google-managed Prometheus instance and build out Grafana dashboards and alerts to increase visibility on our systems
  • Experiment with GCP offerings, 3rd party vendors, AI tooling, and open-source projects to further automate and secure day-to-day operations
  • Pair with engineering leads to instrument and monitor critical functionality
  • Participate in architecture design and capacity planning discussions to ensure that our systems are scalable, maintainable, reliable, and secure
  • Build, maintain, and improve our CI/CD pipeline
You’ll be a great fit for the role if
  • You reach for automation before you reach for a runbook, and a manual process is something you want to delete, not document
  • You treat AI agents as power tools and have real opinions about how to drive them — especially when to stop trusting them
  • You are comfortable wearing many hats
  • You have a willingness to learn and teach in a fast-paced, collaborative environment
  • You have a strong desire to automate things
  • You readily provide constructive feedback, and also proactively seek feedback to improve yourself
  • You like to get your hands dirty and tinker with/stress test new technologies
Preferred Qualifications
  • 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP)
  • Demonstrated fluency directing AI coding agents (e.g. Claude Code, Cursor, or similar) to build, operate, and debug real infrastructure; and robust and experienced judgment on verification of their work
  • A track record of replacing manual operations with durable automation
  • Experience working with the containerization technologies Docker, Kubernetes, and Istio or a similar service mesh technology
  • Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner
  • Experience with stream technologies such as Kafka and Amazon Kinesis
  • Experience with pub sub technologies such as AWS SNS and Google Pub/Sub
  • Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run
  • Experience with infrastructure-as-code tools such as Terraform
  • Experience with observability tools such as Datadog, Prometheus, and Grafana
  • Strong computer science and software engineering fundamentals
  • Experience with SOC2 and PCI Compliance processes and requirements

We are excited to hear from you.

At Attain, we are passionate about finding people to continuously help us grow our organization. We encourage you to apply, even if your experience doesn’t match every detail on the job description. If we don’t see something that immediately fits, we will keep your resume on file for future opportunities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr/Staff Site Reliability Engineer, Consumer Apps Chicago, IL
Sr/Staff Site Reliability Engineer, Consumer Apps Chicago, IL

Attaindata • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Sr/Staff Site Reliability Engineer, Consumer Apps
Sr/Staff Site Reliability Engineer, Consumer Apps

Attain • Chicago (IL)

Hybrid
USD 150,000 - 210,000
Staff Software Engineer, Technical Lead – Merryfield
Staff Software Engineer, Technical Lead – Merryfield

Attain • Chicago (IL)

Hybrid
USD 100,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Attain • Chicago (IL)

On-site
USD 150,000 - 210,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Attentive • Wilmington (DE)

On-site
USD 180,000 - 240,000
Health & wellness
Equity
Senior/Staff Backend Engineer, Platform, Consumer Apps - Klover
Senior/Staff Backend Engineer, Platform, Consumer Apps - Klover

attain • Chicago (IL)

On-site
USD 120,000 - 180,000
Health benefits
Retirement plan
Paid time off
Staff Engineer (Dev Ops)
Staff Engineer (Dev Ops)

Automation Anywhere • San Jose (CA)

Hybrid
USD 190,000 - 210,000
Flexible work schedule
Remote work options
Unlimited PTO
+4
Junior Talent Partner Chicago, IL
Junior Talent Partner Chicago, IL

Attain • Chicago (IL), Northern (KY)

Hybrid
USD 55,000 - 75,000
Site Reliability Engineer [AQ-18653]
Site Reliability Engineer [AQ-18653]

Aquent • Austin (TX)

On-site
USD 140,000 - 190,000
Senior IT Ops Automation Engineer (Hybrid, AI-Driven)
Senior IT Ops Automation Engineer (Hybrid, AI-Driven)

CloudZero • Boston (MA)

On-site