Senior Customer Reliability Engineer

United States Digital Space LLC

United States

Hybrid

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Flexible working hours
Opportunity for professional development

Job summary

United States Digital Space LLC seeks an engineer for the Customer Reliability Engineering team to ensure the reliability of critical services across various industries, including finance and media. This role requires handling high-severity incidents and building proactive solutions to prevent future issues. You will collaborate with engineers and design AI-assisted diagnostics to enhance operation efficiencies.

The ideal candidate should have strong incident response experience, knowledge of AI tooling, and excellent communication skills to work effectively with cross-functional teams. This position is essential to our mission of providing critical infrastructure for reliable services.

Qualifications

  • Experience managing high-severity incidents and resolving complex technical issues.
  • Strong analytical skills for diagnosing systemic risks and opportunities for improvement.
  • Ability to collaborate with cross-functional teams for incident resolution.

Responsibilities

  • Lead response of high-severity incidents and ensure customer satisfaction.
  • Build reliable systems and automate processes to improve operational efficiency.
  • Document and communicate technical procedures and insights from incidents.

Skills

Incident response
Root cause analysis
Proactive engineering
AI-assisted diagnostics

Education

B.Sc. in Computer Science or related field

Tools

Telemetry tools
Logging frameworks
AI tooling

Job description

Available Location: Singapore
Why This Role Exists

the company built its reputation helping build a better Internet, defending millions of sites, giving away SSL and DDoS mitigation when the industry charged premium prices. In an acceleratingly dangerous world, the scope of that mission has changed. We are becoming something more: critical infrastructure. Banks run their payment rails on us. Governments run public services on us. Media companies depend on us during live events. Health systems depend on us to provide care. Reliability for these customers is no longer a feature of our product. It is a mission.

Serving that customer base demands a different operating model. Traditional support organizations route tickets. Traditional engineering organizations ship features. Neither alone is enough when the stakes are this high. We are pivoting to something different: a customer‑facing engineering organization, directly engaged with our customers at scale. This is work a central dev team cannot do from the inside of the network.

The Customer Reliability Engineering function is the spine of that pivot. CRE is SRE applied outward, the same engineering discipline, applied to the reliability of the systems our customers run on the company. You are the engineer who owns the problems that matter most to the customers who matter most, and you contribute directly to our products and tooling, in partnership with Product Engineering, to hold that standard across the entire customer base.

The Role

CRE is a rapid response team and a proactive engineering team. You fix things at the edge as they come up, and you help build the product capabilities that identify customer issues before they become a crisis. Both modes are equally core.

Rapid response. When a customer issue surfaces that is high‑severity, cross‑layer and complex, you are the engineer who answers. You reproduce the defect, isolate the root cause across the company's infrastructure and the customer's stack, drive the fix with Product Engineering, and confirm resolution. You hold on‑call for high‑severity incidents as part of a global shift rotation.

Proactive engineering. When no fire is burning, you work with Product Engineering and our platform teams to build the capabilities that make the next fire cheaper or unnecessary: telemetry pipelines that correlate signals across the customer base, detectors that fire before a human notices, diagnostic tooling that scales across hundreds of customers, automation that reduces toil for Customer Support. Every incident you carry generates engineering output that reduces the cost of recurrence. The work compounds.

the company is building CRE as an AI‑native function. You will work with and help build agents and tooling that pre‑diagnose incidents, surface relevant logs and configuration, and propose fixes with cited evidence. Engineers who ship AI‑assisted diagnostics are the ones defining this discipline.

What You Might Work On

Rapid response:

  • Own a Sev‑1 incident where a large financial services customer sees asymmetric latency from a single POP. Trace it through BGP routing and origin configuration. Produce the fix upstream.
  • Diagnose a recurring WebSocket disconnect that a media customer has been fighting for weeks. Isolate it to a specific interaction between WAF and their origin load balancer. Drive the fix with Product Engineering.
  • Partner with a government customer's SRE team during an active DDoS event. Help them shape their Magic Transit and WAF configuration in real time.

Proactive engineering:

  • Build, with Product Engineering, a distributed tracing capability that correlates the company edge signals with customer origin metrics so a single query tells the story of a failing request end‑to‑end.
  • Ship a detector for a class of WAF false positives silently degrading several customers. Get it into production before the next renewal cycle.
  • Prototype an AI agent that takes a new customer case, pulls relevant logs and config, and proposes a root cause with linked evidence. Deploy it internally. Measure whether it makes engineers faster. Iterate.
Responsibilities

Rapid incident response and root cause analysis. Own the most complex, high‑severity customer issues end‑to‑end, from first signal through confirmed resolution. Lead deep‑digging debugging across the full stack: edge, network, DNS, transport, APIs, application, customer‑side configuration. Reproduce defects, validate fixes with Engineering, and confirm customer‑side resolution. Produce post‑mortems other engineers rely on. Hold on‑call for high‑severity incidents as part of a global rotation that includes weekends.

Proactive reliability engineering. Analyze support and telemetry signals across the customer base to find systemic risks before they become incidents. Contribute monitoring, detection, and diagnostic capability to the core product and the engineering systems that give Customer Support early visibility into customer‑affecting issues. Define customer‑facing reliability metrics (error rates, resolution times, repeat‑contact rates) and drive measurable improvement. Write automation that reduces mean‑time‑to‑detect and mean‑time‑to‑resolve.

Cross‑functional partnership. Manage the technical escalation lifecycle with clear ownership and timely communication. Partner with Product Engineering to drive fixes, workarounds, and configuration changes that address underlying gaps. Represent the customer reliability perspective in engineering syncs, incident reviews, and post‑mortem processes.

Technical leadership and enablement. Raise the technical floor of Customer Support through pair‑debugging, structured knowledge transfer, and shared tooling. Document diagnostic procedures and resolution patterns in runbooks, internal knowledge bases, and AI skills. Share insights from customer‑facing incidents to improve product documentation and operational readiness.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Customer Reliability Engineer (CRE)
Customer Reliability Engineer (CRE)

Arista Networks • Santa Clara (CA)

On-site
USD 100,000 - 130,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering

Bank of America • United States

On-site
USD 110,000 - 140,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
SRE Support Engineer - Observability
SRE Support Engineer - Observability

Gigster • Austin (TX)

Remote
USD 80,000 - 100,000
High autonomy in a remote-first environment
Real technical problem solving
Opportunity for scaling support
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Technical Customer Engineer Remote (United States)
Senior Technical Customer Engineer Remote (United States)

S27a • Northern (KY)

Hybrid
USD 120,000 - 170,000