Lead Site Reliability Engineer

Docusign

United States

Remote

USD 180,000 - 240,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote work

Job summary

DocuSign is seeking an experienced incident-response specialist to lead high-severity outages for a large cloud platform. This individual-contributor role requires coordinating technical, product, and security teams to restore services quickly and reliably.

You will drive incident-management best practices, communicate with executives and stakeholders, and contribute to ongoing reliability improvements across a globally distributed environment.

Qualifications

  • 12+ years of experience in incident management or site reliability engineering.
  • Led major incidents and other high-severity operational situations.
  • Experience using incident-management tools and strong programming ability.
  • Advanced troubleshooting and problem-solving skills in a 24x7x365 environment.
  • Strong judgment, decision-making, and problem-identification capabilities.
  • Excellent written and verbal communication, including the ability to present effectively to business audiences.

Responsibilities

  • Direct incident commanders and coordinate response to critical technology, product, and security incidents.
  • Communicate clearly with technical teams, business stakeholders, and executive leadership during challenging events.
  • Act as a subject-matter expert for incident-management practices and help strengthen the wider service-excellence program.
  • Apply engineering discipline, automation, and operational best practices to improve availability, reliability, and scalability.
  • Review incident data for anomalies, correlations, and recurring trends, then recommend improvements.
  • Participate in a continuous 24x7x365 rotational coverage schedule.

Skills

Incident management
Site reliability engineering
Incident command
Leadership
Programming
Stakeholder communication
Troubleshooting

Tools

Incident-management tools
Monitoring/Observability tools

Job description

Role overview

Lead high-severity incident response for a large cloud software platform, coordinating technical, product, and security teams through complex service disruptions. This individual-contributor role combines incident command, stakeholder communication, operational analysis, and reliability improvement across a globally distributed environment.

Responsibilities
  • Direct incident commanders and coordinate response to critical technology, product, and security incidents.
  • Communicate clearly with technical teams, business stakeholders, and executive leadership during challenging events.
  • Act as a subject-matter expert for incident-management practices and help strengthen the wider service-excellence program.
  • Apply engineering discipline, automation, and operational best practices to improve availability, reliability, and scalability.
  • Review incident data for anomalies, correlations, and recurring trends, then recommend improvements.
  • Participate in a continuous 24x7x365 rotational coverage schedule.
Requirements
  • At least 12 years of experience in incident management or site reliability engineering.
  • Demonstrated leadership of major incidents and other high-severity operational situations.
  • Experience using incident-management tools and strong programming ability.
  • Advanced troubleshooting and problem-solving skills in a 24x7x365 environment.
  • Strong judgment, decision-making, and problem-identification capabilities.
  • Excellent written and verbal communication, including the ability to present effectively to business audiences.
Benefits and work setup
  • Remote position, with most work performed from a designated remote location.
  • Individual-contributor role focused on operational leadership and reliability expertise.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Boston (MA)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • United States

Remote
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Hidden Jobs • United States

On-site
USD 150,000 - 190,000
Healthcare
Retirement matching
Paid family leave
+3
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000