Senior Manager, Site Reliability Engineering

Intuit Inc.

Mountain View (CA)

On-site

USD 222,000 - 301,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intuit Inc. in Mountain View seeks a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10–15 engineers responsible for the availability and operational health of Fintech Platform services on AWS.

You will set technical direction, review designs, and coach engineers through complex production issues. A core priority is AI Ops: embedding autonomous operations to 3x impact, reducing toil, and accelerating value delivery.

Qualifications

  • 8+ years in systems engineering or SRE with 3+ years leading teams.
  • Proven AWS production infra experience at scale (EC2, EKS/ECS, VPC, RDS, etc).
  • Track record of high-availability outcomes for mission-critical, customer-facing systems.
  • Deep incident management experience with postmortems and prevention.
  • Strong foundation in distributed systems, networking, and IaC (Terraform/CloudFormation).
  • Experience with observability tools (Datadog, Splunk, Prometheus).
  • Ability to balance hands-on depth with people leadership and executive communication.

Responsibilities

  • Own end-to-end operational excellence for Fintech Platform services and achieve near-99.999% availability.
  • Lead and grow a team of 10–15 SREs, hiring, mentoring, and setting goals.
  • Participate in architecture reviews, contribute to design decisions, and write/read code or IaC as needed.
  • Define and execute an AI Ops roadmap to automate detection, diagnosis, and remediation.
  • Identify toil, replace with autonomous agents, and measure impact on velocity.
  • Drive incident management maturity and rigorous root-cause analysis.
  • Build scalable AWS infrastructure with resiliency, auto-remediation, and multi-region failover.
  • Define and report on SLOs/SLIs and quality metrics to guide investments.
  • Collaborate with product, security, and compliance teams on reliability.
  • Establish on-call playbooks and escalation paths to reduce MTTR.
  • Oversee capacity planning, cost optimization, and roadmap decisions.
  • Represent Infrastructure & SRE in leadership forums and executive reviews.
  • Foster a culture of operational rigor and continuous improvement.

Skills

Leadership
Incident management
AI Ops
Reliability engineering
Communication
Strategy

Education

Bachelor’s degree in CS/Engineering

Tools

Terraform
CloudFormation
Datadog
Prometheus/Grafana
Kubernetes
EC2
EKS

Job description

About the Team

Intuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps QuickBooks, TurboTax, Credit Karma, and Mailchimp running for hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency tooling, and incident response capability that underpins Intuit's money-movement and fintech services — where availability, data integrity, and trust are non-negotiable.

The Opportunity

We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10–15 systems and reliability engineers responsible for the availability, performance, and operational health of Fintech Platform services running in AWS. This leader owns the strategy and execution behind operational excellence: driving toward a 99.999% availability bar, maturing incident management practices, and building self-healing, well-instrumented infrastructure at scale.

This is a player-coach role. You will set technical direction and organizational strategy while staying close to the systems — reviewing designs, joining incident bridges, and coaching engineers through complex production issues. You'll partner closely with software engineering, product, security, and other SRE/infrastructure leaders across Intuit to raise the bar on reliability company-wide.

A defining priority for this role is AI Ops: embedding AI-driven, autonomous operations into how the team runs infrastructure. You will lead the shift from manual, human-triggered response toward self-healing systems that detect, diagnose, and remediate issues autonomously — reducing developer toil, cutting MTTR, and freeing engineering capacity to focus on higher-value work. Done well, this delivers 3x the operational impact of the team today and directly accelerates the pace at which we deliver value to customers.

Responsibilities

Responsibilities
  • Own end-to-end operational excellence for Fintech Platform services: define and drive the strategy for achieving and sustaining 99.999% availability across customer-facing and internal systems.

  • Lead, grow, and directly manage a team of 10–15 systems/site reliability engineers — hiring, mentoring, setting goals, and developing the next generation of technical leaders.

  • Act as a hands‑on technical leader: participate in architecture and design reviews, write and review code/IaC where needed, and dive into production systems alongside the team.

  • AI Ops: Driving 3x Impact Through Autonomous Operations-

    Define and execute an AI Ops roadmap that embeds autonomous detection, diagnosis, and remediation into production systems, targeting a 3x improvement

  • Identify high-toil, repetitive operational workflows and systematically replace them with autonomous agents and automation, freeing engineers to focus on higher‑leverage engineering work.

  • Measure and report on toil reduction, automation coverage, and velocity gains, tying AI Ops investment directly to faster, safer delivery of customer value.

  • Drive incident management maturity — own the incident command process, lead or oversee response for high‑severity (P1/P2) incidents, and ensure rigorous root‑cause analysis and blameless postmortems.

  • Build and scale AWS cloud infrastructure (compute, networking, storage, container orchestration) with a focus on resiliency, auto‑remediation, chaos engineering, and multi‑AZ/multi‑region failover.

  • Define and report on SLOs/SLIs, error budgets, and availability metrics; use data to prioritize reliability investments and reduce toil through automation.

  • Partner with software engineering, product management, security, and compliance teams to embed reliability, observability, and operational readiness into the software development lifecycle.

  • Establish and continuously improve on‑call practices, runbooks, alerting, and escalation paths to reduce MTTD/MTTR.

  • Own capacity planning, cost optimization, and infrastructure roadmap decisions for the systems under your purview.

  • Represent Infrastructure & SRE in leadership forums, change advisory boards, and executive incident reviews; communicate risk and operational posture clearly to senior stakeholders.

  • Champion a culture of operational rigor, psychological safety, and continuous improvement across the team.

Qualifications

Qualifications
  • 8+ years of experience in systems engineering, site reliability engineering, or infrastructure engineering, with 3+ years directly managing engineering teams.

  • Proven, hands‑on experience operating production infrastructure in AWS at scale (EC2, EKS/ECS, VPC, RDS/DynamoDB, IAM, CloudWatch, Auto Scaling, and related services).

  • Track record of driving high‑availability outcomes (99.9%+ and above) for mission‑critical, customer‑facing systems, ideally in fintech, payments, or another regulated/high‑trust domain.

  • Deep experience with incident management — running incident command, leading postmortems, and building organizational muscle around detection, response, and prevention.

  • Strong technical foundation in distributed systems, networking, containerization/orchestration (Kubernetes), and infrastructure‑as‑code (Terraform, CloudFormation, or similar).

  • Experience with observability and reliability tooling (e.g., Datadog, Splunk, PagerDuty, Prometheus/Grafana) and building SLO‑driven operations.

  • Demonstrated ability to balance hands‑on technical depth with people leadership — comfortable reviewing a design doc and coaching a direct report in the same day.

  • Excellent communication skills, with experience presenting risk, status, and strategy to senior/executive leadership.

  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

  • Experience designing or scaling AI Ops / autonomous remediation capabilities (AIOps platforms, ML‑based anomaly detection, agentic automation) that measurably reduced toil or MTTR.

  • Experience operating within a regulated fintech, banking, or payments environment (PCI, SOC 2, money movement/ACH systems).

  • Prior experience building or scaling a chaos engineering or resilience testing practice.

  • Familiarity with cost and capacity management for large‑scale multi‑account AWS environments.

  • Experience leading through major incidents involving cross‑functional executive stakeholders.

Intuit provides a competitive compensation package with a strong pay for performance rewards approach. This position may be eligible for a cash bonus, equity rewards and benefits, in accordance with our applicable plans and programs (see more about our compensation and benefits at Intuit®: Careers | Benefits). Pay offered is based on factors such as job‑related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender.

The expected base pay range for this position is: Mountain View $222,000 - $300,500

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Site Reliability Engineering
Senior Manager, Site Reliability Engineering

Intuit • Mountain View (CA)

On-site
USD 222,000 - 301,000
Staff App Ops Engineer
Staff App Ops Engineer

Intuit • Mountain View (CA)

On-site
USD 203,000 - 274,000
Staff App Ops Engineer
Staff App Ops Engineer

Intuit Inc. • New York (NY)

On-site
USD 180,000 - 240,000
Staff App Ops Engineer
Staff App Ops Engineer

Intuit • New York (NY)

On-site
USD 140,000 - 210,000
Fintech Futures - Senior Staff Software Engineer
Fintech Futures - Senior Staff Software Engineer

Intuit Inc. • Mountain View (CA)

On-site
USD 221,000 - 299,000
Fintech Futures - Senior Staff Software Engineer
Fintech Futures - Senior Staff Software Engineer

Intuit • Mountain View (CA)

On-site
USD 221,000 - 299,000
Principal Software Engineer, Fintech Agentic AI
Principal Software Engineer, Fintech Agentic AI

Intuit • Mountain View (CA)

On-site
USD 262,000 - 354,000
Cash bonus
Equity rewards
Benefits package
Staff Software Engineer - Back End
Staff Software Engineer - Back End

Intuit Inc. • Mountain View (CA)

On-site
USD 203,000 - 274,000
Senior Staff Software Engineer
Senior Staff Software Engineer

Intuit • Mountain View (CA)

On-site
USD 221,000 - 299,000
Staff Business Systems Analyst - Finance Systems
Staff Business Systems Analyst - Finance Systems

Intuit Inc. • San Diego (CA)

On-site
USD 157,000 - 212,000