SRE Lead: Incident Recovery & Observability

Bestegg

Wilmington (DE)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Retirement plans
Paid time off
Health insurance options
Flexible Spending Plans
Life Insurance
Wellness programs
Employee assistance programs
Discounted benefits

Job summary

Best Egg is seeking a Site Reliability Engineer to lead major incident recovery, improve observability, mentor engineers, and drive reliability improvements across our platform.

You will own incident response, optimize alerts and dashboards in Datadog, and collaborate with AWS/cloud teams to reduce outages and toil. Strong communication and leadership under pressure are essential to success in a fast‑paced financial services environment.

Qualifications

  • Hands‑on familiarity with production support, monitoring, alerting, and incident response practices.
  • Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
  • Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
  • Experience with AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools.
  • Clear communication skills during incidents, service requests, and post‑incident follow‑through.
  • Strong experience leading production incident recovery and cross‑system reliability investigations.
  • Ability to mentor engineers and influence technical decisions without direct authority.
  • Datadog, AWS, ITIL, Linux, or automation certification.
  • Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments.
  • Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
  • Familiarity with financial services controls, secure file transfer, or regulated operations.

Responsibilities

  • Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
  • Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
  • Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
  • Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
  • Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
  • Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards.
  • Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.

Skills

Incident response
Observability
AWS
Mentoring
Communication
CI/CD
Automation
SRE practices
Reliability engineering
AIOps

Tools

Datadog
JAMS
GoAnywhere
xMatters
ServiceNow/Jira
CI/CD tools
AWS tooling

Job description

Best Egg is seeking a Site Reliability Engineer to lead major incident recovery, improve observability, mentor engineers, and drive reliability improvements across our platform.

You will own incident response, optimize alerts and dashboards in Datadog, and collaborate with AWS/cloud teams to reduce outages and toil. Strong communication and leadership under pressure are essential to success in a fast‑paced financial services environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead: Reliability, Incident Command & Automation
SRE Lead: Reliability, Incident Command & Automation

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 112,000 - 131,000
Life insurance
Disability
Parental leave
+5
Site Reliability Engineer
Site Reliability Engineer

Bestegg • Wilmington (DE)

On-site
USD 120,000 - 160,000
Retirement plans
Paid time off
Health insurance options
+5
SRE Lead: Incident Commander & Reliability Champion
SRE Lead: Incident Commander & Reliability Champion

U.S. Bank • Northern (KY)

Hybrid
USD 112,000 - 131,000
Healthcare
Retirement plan
Paid vacation
+2
SRE Lead: Incident Commander & Reliability Architect
SRE Lead: Incident Commander & Reliability Architect

U.S. Bank • Chicago (IL)

On-site
USD 112,000 - 131,000
Healthcare benefits
401(k) retirement plan
Paid vacation
Senior Site Reliability Engineer — Scalable, Observability-Driven
Senior Site Reliability Engineer — Scalable, Observability-Driven

Origami Risk • Chicago (IL)

Hybrid
USD 100,000 - 120,000
Hybrid work arrangement
401(k) with company match
Medical, dental, and vision benefits
+1
SRE Lead: Production Reliability & Observability Architect
SRE Lead: Production Reliability & Observability Architect

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Hybrid Principal SRE - Reliability at Scale & Observability
Hybrid Principal SRE - Reliability at Scale & Observability

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 194,000 - 237,000
Healthcare Coverage
401(k) Plan
Paid Time Off
+2
Site Reliability Engineer - Incident & Observability Lead
Site Reliability Engineer - Incident & Observability Lead

Worky • Atlanta (GA)

Hybrid
USD 100,000 - 120,000
Medical & Dental
Hybrid/Remote work
Vision Insurance
+4
Senior SRE: Lead Reliability for AI Platform (Hybrid)
Senior SRE: Lead Reliability for AI Platform (Hybrid)

Docebo • United States

Hybrid
USD 140,000 - 210,000
Health benefits
Paid vacation days
Docebo Days
+3
Lead Observability Engineer (SRE) — Reliability & Dashboards
Lead Observability Engineer (SRE) — Reliability & Dashboards

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 86,000 - 102,000
401(k) retirement plan
Paid vacation
Up to 11 paid holidays
+2