SRE - Incident RCA & Observability (Remote)

Socket.dev

Georgia

Hybrid

USD 100,000 - 120,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical and Dental coverage
Vision coverage
401(k) match
Flexible time off
Hybrid/Remote work
Vacation and sick leave
Wellness reimbursement
Life insurance
Education assistance
Pre-Tax Savings Accounts
EAP

Job summary

Origami Risk is hiring a Site Reliability Engineer to help improve time-to-resolution and scale our SaaS platform. You will lead post-incident investigations, craft RCAs, and collaborate across Cloud Operations and Engineering to prevent incidents and enhance observability.

The role emphasizes deep observability tool experience, cloud familiarity (AWS/Azure), and strong coding skills to support automation and incident response. Hybrid or remote work options are available where permitted.

Qualifications

  • Bachelor's degree in Computer Science or related field (or equivalent experience)
  • 5+ years of proven experience in a Site Reliability Engineering role
  • Strong knowledge of SRE best practices and incident management protocols
  • Deep experience using and/or configuring New Relic, Data Dog, SumoLogic or similar observability tools
  • Proficiency in reading and writing code (e.g., JavaScript, .NET, SQL)
  • Familiarity with cloud platforms (e.g., AWS, Azure) and architectural patterns
  • Excellent problem-solving skills and a data-driven approach to incident analysis
  • Prior experience operating within a Public Cloud environment (AWS strongly preferred)
  • Experience troubleshooting C#/.Net based web applications to identify bugs/performance challenges.
  • Solid knowledge of SaaS operations
  • Advanced written and verbal communication skills
  • Knowledge of CI/CD pipelines
  • Experience working in an IaC environment

Responsibilities

  • Leads post-incident investigations for the Site Reliability team.
  • Conducts in-depth post-incident analyses to identify root causes and develops preventive strategies.
  • Drafts clear and insightful RCAs for customer delivery.
  • Cross trains colleagues on how to best leverage observability tools during incident and performance investigations.
  • Provides visibility to all stakeholders throughout the entire Site Reliability process.
  • Collaborates with cross-functional teams to implement system enhancements that enhance scalability and stability.
  • Develops client-focused dashboards/alerts to proactively identify performance challenges.
  • Monitors and continuously improves our time to resolution metrics.
  • Maintains and configures core observability tools to ensure optimum performance and key metrics/data are available for incident response and performance investigations.
  • Provides an actionable feedback loop to Observability and Engineering teams toward improving MELT and development patterns.
  • Contributes to the development of automation tools to streamline incident response.
  • Works proactively to prevent incidents and reduce their impact on our platform.
  • Partners with the larger Cloud Operations, SRE, Engineering teams, and the business-at-large to advance our SaaS platforms.
  • Other duties as assigned.

Skills

SRE best practices
Incident management
Observability
Code reading/writing
Cloud platforms
Public cloud experience
C#/.NET troubleshooting
SaaS operations
Communication skills
CI/CD pipelines
IaC

Education

Bachelor's degree in CS or related field

Tools

New Relic
DataDog
Sumo Logic

Job description

Origami Risk is hiring a Site Reliability Engineer to help improve time-to-resolution and scale our SaaS platform. You will lead post-incident investigations, craft RCAs, and collaborate across Cloud Operations and Engineering to prevent incidents and enhance observability.

The role emphasizes deep observability tool experience, cloud familiarity (AWS/Azure), and strong coding skills to support automation and incident response. Hybrid or remote work options are available where permitted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Incident & Observability Lead
Site Reliability Engineer - Incident & Observability Lead

Worky • Atlanta (GA)

Hybrid
USD 100,000 - 120,000
Medical & Dental
Hybrid/Remote work
Vision Insurance
+4
Remote SRE Lead - Incident, Reliability & Observability
Remote SRE Lead - Incident, Reliability & Observability

NightDragon Acquisition Corp. • United States

On-site
USD 260,000 - 280,000
Hybrid & Remote Work
Competitive Compensation
Equity package
Site Reliability Engineer — Cloud, Incidents & RCA
Site Reliability Engineer — Cloud, Incidents & RCA

Resolve Tech Solutions • Irving (TX)

On-site
USD 120,000 - 160,000
Remote SRE II: Cloud, Data Ops & Incident Response
Remote SRE II: Cloud, Data Ops & Incident Response

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 100,000 - 110,000
Fully remote
5% travel
Medical insurance
+7
Remote Evening SRE: ROSA & AWS Observability Expert
Remote Evening SRE: ROSA & AWS Observability Expert

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
Senior SRE: Platform Reliability & Incident Lead (Remote)
Senior SRE: Platform Reliability & Incident Lead (Remote)

Affirm, Inc. • Town of Poland (NY)

On-site
USD 32,000 - 48,000
Health insurance
Equity rewards
Flexible Spending Wallets
+1
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE - Remote Observability & Reliability Leader
Senior SRE - Remote Observability & Reliability Leader

DOMA Technologies • Leesburg (VA)

Remote
USD 120,000 - 150,000
SRE Lead — AI-Driven Reliability & Observability (Remote)
SRE Lead — AI-Driven Reliability & Observability (Remote)

Domino Data Lab • United States

On-site
USD 120,000 - 150,000
Evening SRE – Remote Cloud & OpenShift Reliability Engineer
Evening SRE – Remote Cloud & OpenShift Reliability Engineer

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000