Site Reliability Engineering Manager

DriveWealth

Chicago (IL)

Hybrid

USD 160,000 - 230,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health & wellness packages
Unlimited vacation
Remote/hybrid options
Professional development support

Job summary

DriveWealth is seeking a Manager of Site Reliability Engineering to lead a team of SRE Automation Engineers and stay hands-on as a principal engineer. You’ll reduce toil, build scalable automation, and guide reliability initiatives across our Kubernetes-based, multi-region platform for a regulated brokerage environment.

You’ll drive the automation agenda with Rundeck and Airflow, mentor engineers, and own incident response, RCA culture, and SRE governance while collaborating with engineering

Qualifications

  • Leads SRE Automation Engineers, balancing people leadership with hands-on engineering depth.
  • Drives automation strategy and builds an automation practice in a regulated brokerage environment.
  • Develops internal tooling and IaC standards using Terraform and GitOps with ArgoCD.

Responsibilities

  • Lead and mentor a team of SRE Automation Engineers, guiding career development and performance.
  • Design and implement internal tooling and automation, including Rundeck and Airflow orchestration.
  • Define SLIs/SLOs, error budgets, and postmortems to drive reliability and toil reduction.
  • Oversee IaC standards and GitOps workflows for Kubernetes and cloud infrastructure.
  • Collaborate with engineering leadership to align reliability priorities with business goals.

Skills

SRE leadership
Automation engineering
Team mentorship
Google SRE principles
Kubernetes
GitOps
Python
Golang
Terraform
Airflow
Rundeck
ArgoCD
CI/CD
Linux
Networking
Observability
Cloud (AWS)
Incident response

Tools

Rundeck
Airflow
ArgoCD
Terraform
Kubernetes
Grafana
Prometheus
AWS CLI
boto3
Kafka
MQ
SQS
Ansible

Job description


  • As the Manager of Site Reliability Engineering, you’ll lead a team of SRE Automation Engineers while remaining a hands‑on technical authority for our Brokerage-as-a-Service platform. This isn’t a purely people‑management seat, you’re expected to bring the same principal‑level SRE depth to automation design and engineering as an individual contributor, while also building the team, setting technical direction, and developing your engineers’ careers

  • This role is centered on reducing manual toil through engineering, applying Google’s SRE principles: SLOs, error budgets, blameless postmortems, and systematic toil reduction, adapted to a regulated brokerage environment. You’ll carry two responsibilities at once: driving the automation agenda, building and orchestrating workflows in Rundeck and Airflow to eliminate repetitive work, and growing your team of SRE Automation Engineers into a high‑functioning automation practice

  • You’ll guide the design of internal SRE platforms, automate complex workflows, and ensure our Kubernetes‑based and colo ecosystems can handle the demands of global financial markets, while owning the people side of the team: mentorship, performance, and growth, and the day‑to‑day management of the team’s Jira board. While this role includes participation in on‑call rotations supporting our 24/7 global operations, your primary mission is to build systems that make manual intervention obsolete, and a team capable of sustaining that mission

  • Team Leadership & Development: Manage, mentor, and grow a team of SRE Automation Engineers—setting technical direction, running 1:1s, owning performance management and career development, and managing the team’s Jira board to prioritize and track sprint work

  • Engineering & Automation: Lead the design and development of internal tooling and automation—including Rundeck and Airflow‑based orchestration—to eliminate repetitive manual toil and improve developer velocity, staying hands‑on with the most complex, highest‑leverage automation work yourself

  • SRE Practice & Governance: Adapt Google’s SRE principles to our environment—defining SLIs, SLOs, and error budgets, and using them to guide engineering and operational priorities

  • Infrastructure as Code: Set architectural standards for modular, reusable IaC using Terraform and oversee GitOps workflows via ArgoCD

  • Platform Governance: Review software architecture and Kubernetes metrics to ensure high availability, capacity planning, and cost‑optimization across AWS regions, and hold the team accountable to those standards

  • Incident Engineering: Lead incident response for critical events, drive complex root‑cause analysis (RCA), and champion a blameless post‑mortem culture across the organization

  • Collaboration & Stakeholder Management: Partner with engineering leadership to align SRE priorities with business goals, and foster adoption of new tools, security standards, and reliability best practices across teams


Benefits


  • Health & wellness packages: We’ve built our benefits offering as a holistic package. We provide multiple health insurance carrier options and plans (medical, dental, vision) with access to HSA and/or FSA tax savings tools. We also provide income protection including life insurance, AD&D, short‑ and long‑term disability, as well as extended coverage resources such as fertility benefits and mental wellness resources

  • Vacation & time off: To build a successful organization, employees need time away to rest and recharge, so we provide paid time off including paid holidays and unlimited vacation time. We also strongly believe in work‑life balance and, therefore, offer fully remote and hybrid work positions based on role requirements

  • Professional developmnet: We believe one of the greatest contributions we can make is investing in our employees’ professional development. Therefore, we offer financial support toward continuing education courses, academic coursework, professional conferences, earning professional certifications, and membership fees to professional organizations



  • Code Proficiency: Strong scripting and development skills in Python or Golang, along with Bash and Ansible

  • Modern CI/CD & GitOps: Experience building secure, automated delivery pipelines and operating GitOps workflows (ArgoCD)

  • Security Mindset: Experience with secrets management, vulnerability scanning, and securing the software supply chain

  • Cloud Native Expertise (AWS): Strong grasp of AWS core services, security, and high‑availability patterns. Proficiency with boto3 and AWS CLI for automation

  • AI & Prompt Engineering: Familiarity with using LLMs, Public MCPs, or Bedrock Agent Core to enhance SRE workflows

  • People Leadership: Prior experience managing or leading SRE/DevOps engineers, ideally in a fintech or highly regulated environment. Able to flex between hands‑on principal‑level engineering and coaching and developing a team

  • Linux & Networking Mastery: Proficient in Linux administration with a deep understanding of the TCP/IP stack, OSI model, DNS, and network troubleshooting

  • Data & Middleware & Orchestration: Hands‑on experience with Rundeck and Airflow for job orchestration and automation, plus experience managing Kafka, MQ, or SQS

  • Google SRE Fundamentals: Working knowledge of Google’s SRE practices—SLIs/SLOs, error budgets, toil reduction, and blameless postmortems—and experience adapting them to a regulated environment

  • Observability: Experience with Grafana/Similar tools, Prometheus, Understanding of logs shipping, management and metric first alerting

  • Production Kubernetes: Hands‑on experience managing production‑grade clusters, including RBAC, autoscaling, Helm, and multi‑cluster patterns

  • FinTech Background: Experience working in highly regulated financial environments or with FIX/API connectivity

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

On-site
USD 150,000 - 200,000