Senior Azure SRE: Automation, Observability & Reliability

Koitecc Solutions

Plano, Northern (TX, KY)

Hybrid

USD 153,000 - 192,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Discretionary incentive
Benefits package

Job summary

Bank of America is seeking a Senior Azure Site Reliability Engineer to partner with engineering and technology leaders to define reliability goals and build robust observability for services across the Azure platform.

The role focuses on automating services, improving production readiness, and driving incident response and problem management, mentoring engineers and shaping long‑term reliability improvements.

Qualifications

  • Advanced experience in Azure platform engineering, SRE, cloud infrastructure, or enterprise cloud operations.
  • Deep knowledge of Microsoft Azure architecture, including networking, identity, compute, PaaS, monitoring, security, governance, and resiliency patterns.
  • Strong experience designing and developing Terraform modules and infrastructure‑as‑code automation in enterprise environments.
  • Strong understanding of Azure landing zones, hub‑and‑spoke networking, ExpressRoute or enterprise connectivity, private endpoints, DNS, routing, firewalls, and workload isolation.
  • Advanced observability experience with Azure Monitor, Log Analytics, Dynatrace, dashboards, alerting, metrics, and platform telemetry.
  • Experience with SRE operating models, SLIs, SLOs, incident response, problem management, toil reduction, and production‑readiness reviews.
  • Strong scripting or programming experience with Python, PowerShell, Bash, Java, or similar languages.
  • Experience operating highly available Azure IaaS and PaaS services in enterprise‑scale environments.
  • Ability to influence architecture and engineering decisions across multiple technical teams.
  • Strong communication skills with the ability to translate complex engineering topics into actionable recommendations.

Responsibilities

  • Designs solutions to visualize key production support metrics enabling Operational Readiness and Site Reliability Engineer teams to identify scenarios requiring intervention.
  • Develops software solutions and/or improved processes to address toil by collaborating with partners to remediate processes and free time for reliability.
  • Partners with Development and Infrastructure teams to create error budget policies prioritizing reliability above SLO thresholds and suggests code optimizations or added instrumentation.
  • Identifies capacity bottlenecks, vulnerabilities, and opportunities for reliability improvement to reduce manual effort.
  • Assesses monitoring for new changes with development partners and enhances dashboards and monitoring designs.
  • Engages as a subject matter expert in incident triage, failure scenario modeling, and root cause analysis with Problem Management.
  • Collaborates to develop SLIs/SLOs to measure and improve reliability of services.
  • Leads complex platform reliability initiatives such as secondary-region readiness and enterprise dashboard automation.
  • Defines and matures SLIs/SLOs, alerting standards, and service health reporting for Azure services.
  • Develops reusable Terraform modules, automation frameworks, and CI/CD patterns for consistency and compliance.
  • Drives observability improvements using Azure Monitor, Log Analytics, Dynatrace, and enterprise monitoring tools.
  • Identifies systemic reliability risks and translates them into engineering roadmaps and automation opportunities.

Skills

Architecture
Collaboration
Innovative Thinking
Result Orientation
Solution Design
Adaptability
Analytical Thinking
Influence
Stakeholder Management
Technical Strategy Development
Terraform
Python

Tools

Terraform
Dynatrace
Azure Monitor
Log Analytics
CI/CD pipelines
Git
PowerShell
Bash
Jira

Job description

Bank of America is seeking a Senior Azure Site Reliability Engineer to partner with engineering and technology leaders to define reliability goals and build robust observability for services across the Azure platform.

The role focuses on automating services, improving production readiness, and driving incident response and problem management, mentoring engineers and shaping long‑term reliability improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — Automation, Observability & Reliability
Senior SRE — Automation, Observability & Reliability

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Senior Cloud Reliability Engineer, GCP & Azure
Senior Cloud Reliability Engineer, GCP & Azure

Bank of America • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
Remote Azure SRE – Observability, CI/CD & Resilience
Remote Azure SRE – Observability, CI/CD & Resilience

System Automation Corporation • United States

On-site
USD 120,000 - 140,000
Senior SRE Lead: Cloud Platform Reliability & Automation
Senior SRE Lead: Cloud Platform Reliability & Automation

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Azure SRE Architect: Scale, Automate & Observability
Azure SRE Architect: Scale, Automate & Observability

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Senior SRE, AI Production Reliability & Observability
Senior SRE, AI Production Reliability & Observability

EPAM Systems • Town of Poland (NY)

On-site
USD 140,000 - 210,000
Senior Azure SRE: Cloud Reliability & Automation
Senior Azure SRE: Cloud Reliability & Automation

TeamViewer • Austin (CO)

On-site
USD 130,000 - 190,000
Competitive compensation
Flexible PTO
401(k) with employer matching
+6
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior SRE: Azure Reliability & Observability Lead
Senior SRE: Azure Reliability & Observability Lead

Visa Hunt • United States

On-site
USD 140,000 - 200,000