Site Reliability Engineer

Bestegg

Wilmington (DE)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Retirement plans
Paid time off
Health insurance options
Flexible Spending Plans
Life Insurance
Wellness programs
Employee assistance programs
Discounted benefits

Job summary

Best Egg is seeking a Site Reliability Engineer to lead major incident recovery, improve observability, mentor engineers, and drive reliability improvements across our platform.

You will own incident response, optimize alerts and dashboards in Datadog, and collaborate with AWS/cloud teams to reduce outages and toil. Strong communication and leadership under pressure are essential to success in a fast‑paced financial services environment.

Qualifications

  • Hands‑on familiarity with production support, monitoring, alerting, and incident response practices.
  • Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
  • Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
  • Experience with AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools.
  • Clear communication skills during incidents, service requests, and post‑incident follow‑through.
  • Strong experience leading production incident recovery and cross‑system reliability investigations.
  • Ability to mentor engineers and influence technical decisions without direct authority.
  • Datadog, AWS, ITIL, Linux, or automation certification.
  • Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments.
  • Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
  • Familiarity with financial services controls, secure file transfer, or regulated operations.

Responsibilities

  • Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
  • Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
  • Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
  • Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
  • Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
  • Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards.
  • Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.

Skills

Incident response
Observability
AWS
Mentoring
Communication
CI/CD
Automation
SRE practices
Reliability engineering
AIOps

Tools

Datadog
JAMS
GoAnywhere
xMatters
ServiceNow/Jira
CI/CD tools
AWS tooling

Job description

Best Egg is a market‑leading financial platform offering installment lending solutions. We are looking for a Site Reliability Engineer to lead major incident recovery, improve observability, mentor engineers, and drive reliability improvements.

Responsibilities
  • Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
  • Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
  • Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
  • Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
  • Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
  • Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards.
  • Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.
Qualifications
  • Hands‑on familiarity with production support, monitoring, alerting, and incident response practices.
  • Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
  • Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
  • Experience with AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools.
  • Clear communication skills during incidents, service requests, and post‑incident follow‑through.
  • Strong experience leading production incident recovery and cross‑system reliability investigations.
  • Ability to mentor engineers and influence technical decisions without direct authority.
  • Datadog, AWS, ITIL, Linux, or automation certification.
  • Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments.
  • Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
  • Familiarity with financial services controls, secure file transfer, or regulated operations.
Benefits
  • Pre‑tax and post‑tax retirement savings plans with competitive company matching.
  • Generous paid time‑off plans including vacation, personal/sick time, paid short‑term and long‑term disability leaves, paid parental leave, and paid company holidays.
  • Multiple health care plans to choose from, including dental and vision options.
  • Flexible Spending Plans for Health Care, Dependent Care, and Health Reimbursement Accounts.
  • Company‑paid benefits such as life insurance, wellness platforms, employee assistance programs, and Health Advocate programs.
  • Other great discounted benefits including identity theft protection, pet insurance, fitness center reimbursements, and many more.

Please note this job description is not designed to cover every activity, duty, or responsibility required for the job. Duties, responsibilities, and activities may change at any time with or without notice.

Best Egg celebrates diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Best Egg • Wilmington (DE)

On-site
USD 120,000 - 140,000
Retirement plans
Paid time off
Health care plans
+2
Site Reliability Engineer
Site Reliability Engineer

Best Egg • Delaware

On-site
USD 120,000 - 140,000
Retirement plans
Generous PTO
Health care plans
+3
Site Reliability Engineer: Lead Incidents & Observability
Site Reliability Engineer: Lead Incidents & Observability

Best Egg • Delaware

On-site
USD 120,000 - 140,000
Retirement plans
Generous PTO
Health care plans
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
SRE Lead: Incident Recovery & Observability
SRE Lead: Incident Recovery & Observability

Bestegg • Wilmington (DE)

On-site
USD 120,000 - 160,000
Retirement plans
Paid time off
Health insurance options
+5
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Plano (TX)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Pleasanton (CA)

On-site
USD 81,000 - 187,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Austin (TX)

On-site
USD 81,000 - 187,000
Medical, dental, and vision insurance
401(k) savings and investment plan
Paid parental leave
+1
Site Reliability Engineer - Intermediate
Site Reliability Engineer - Intermediate

Equifax, Inc. • Alpharetta (GA)

Hybrid
USD 120,000 - 180,000