Senior Azure SRE - Platform Reliability & Automation Lead

Hobbsnews

Plano (TX)

On-site

USD 152,600 - 191,500

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Discretionary incentive eligible
Benefits eligible

Job summary

Bank of America is seeking a Senior Azure Site Reliability Engineer to design and mature reliability capabilities across the enterprise Azure platform. You will build automation, define SLOs, mentor engineers, and drive production readiness and incident management initiatives.

You will collaborate with cross-functional teams to improve observability, implement error budgets, and deliver reliable cloud services with strong governance and security alignment.

Qualifications

  • Advanced experience in Azure platform engineering, SRE, cloud infrastructure, or enterprise cloud operations.
  • Deep knowledge of Microsoft Azure architecture, including networking, identity, compute, PaaS, monitoring, security, governance, and resiliency patterns.
  • Strong experience designing and developing Terraform modules and infrastructure-as-code automation in enterprise environments.
  • Strong understanding of Azure landing zones, hub-and-spoke networking, ExpressRoute or enterprise connectivity, private endpoints, DNS, routing, firewalls, and workload isolation.
  • Advanced observability experience with Azure Monitor, Log Analytics, Dynatrace, dashboards, alerting, metrics, and platform telemetry.
  • Experience with SRE operating models, SLIs, SLOs, incident response, problem management, toil reduction, and production-readiness reviews.
  • Strong scripting or programming experience with Python, PowerShell, Bash, Java, or similar languages.
  • Experience operating highly available Azure IaaS and PaaS services in enterprise-scale environments.
  • Ability to influence architecture and engineering decisions across multiple technical teams.
  • Strong communication skills with the ability to translate complex engineering topics into actionable recommendations.

Responsibilities

  • Design solutions to visualize key production support metrics enabling Operational Readiness and Site Reliability Engineer teams to identify scenarios requiring intervention.
  • Develop software solutions and/or improved processes to address work identified as ‘toil’ by collaborating with key partners to identify, track and remediate processes to free time allocated to reliability.
  • Partner with Development and Infrastructure teams to create error-budget policies prioritizing reliability stories that fall below Service Level Objective (SLO) thresholds and suggest code optimizations, additional instrumentation and/or logging structures to gain service reliability visibility.
  • Identify and plan for capacity bottlenecks, vulnerabilities and opportunities for reliability improvement, such as low-level error rates and ‘noise’, and reduce manual support effort and/or improve system reliability.
  • Assess monitoring for new changes with development partners and work with monitoring tools team to monitor dashboards and enhance application and system monitoring designs.
  • Engage as a subject-matter expert in incident triage efforts, failure scenario modelling and work with the Problem Manager to diagnose root causes for complex/high-impact incident/problem-management investigations.
  • Collaborate with Development and Infrastructure teams to understand technical solutions and develop Service Level Indicators and SLOs to measure/improve the reliability of the services they support.
  • Lead complex platform reliability initiatives such as secondary-region readiness, egress/ingress observability, private DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation.
  • Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting for Azure platform services.
  • Develop reusable Terraform modules, automation frameworks, and CI/CD patterns that improve consistency, compliance, and operational quality.
  • Drive observability improvements using Azure Monitor, Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools.
  • Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, automation opportunities, and operational controls.
  • Partner with security and governance teams to integrate IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness into Azure platform operations.
  • Provide technical design input for new Azure services and workloads to ensure operational readiness before production adoption.
  • Mentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production support.
  • Create executive-ready technical summaries, reliability narratives, and recommendations for leadership review.

Skills

Architecture
Collaboration
Innovative Thinking
Result Orientation
Solution Design
Adaptability
Analytical Thinking
Influence
Stakeholder Management
Technical Strategy Development
Terraform
Python

Job description

Bank of America is seeking a Senior Azure Site Reliability Engineer to design and mature reliability capabilities across the enterprise Azure platform. You will build automation, define SLOs, mentor engineers, and drive production readiness and incident management initiatives.

You will collaborate with cross-functional teams to improve observability, implement error budgets, and deliver reliable cloud services with strong governance and security alignment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Azure SRE: Automation, Observability & Reliability
Senior Azure SRE: Automation, Observability & Reliability

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
Senior Azure SRE: Reliability, Automation & Observability
Senior Azure SRE: Reliability, Automation & Observability

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Senior Cloud Reliability Engineer, GCP & Azure
Senior Cloud Reliability Engineer, GCP & Azure

Bank of America • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Senior SRE Lead: Cloud Platform Reliability & Automation
Senior SRE Lead: Cloud Platform Reliability & Automation

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior SRE — Automation, Observability & Reliability
Senior SRE — Automation, Observability & Reliability

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Azure SRE Lead: Reliability Strategy & Incident Leadership
Azure SRE Lead: Reliability Strategy & Incident Leadership

Tata Consultancy Services • Bellevue (WA)

On-site
USD 38,000 - 121,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Maternal & Parental Leaves
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Senior Cloud Platform Engineering Lead
Senior Cloud Platform Engineering Lead

National Black MBA Association • Jersey City (NJ)

On-site
USD 136,000 - 220,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible