Site Reliability Engineer - Incident & Observability Lead

Worky

Atlanta (GA)

Hybrid

USD 100,000 - 120,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical & Dental
Hybrid/Remote work
Vision Insurance
Life Insurance
401(k) Match
PTO & Holidays
Education Assistance

Job summary

Origami Risk is seeking a Site Reliability Engineer to improve time-to-resolution and overall platform reliability for our SaaS solutions. You will lead post-incident investigations, identify root causes, and implement preventive measures while supporting client performance challenges and observability initiatives.

You will configure monitoring tools, develop dashboards, and collaborate with Cloud Operations, SRE, and Engineering to scale and stabilize services.

Qualifications

  • Bachelor's degree in Computer Science or related field or equivalent experience.
  • 5+ years of proven Site Reliability Engineering experience.
  • Strong knowledge of SRE best practices and incident management protocols.
  • Proficiency in reading and writing code (JavaScript, .NET, SQL).
  • Experience with cloud platforms (AWS, Azure) and architectural patterns.
  • Experience in Public Cloud environments, AWS preferred.
  • Experience troubleshooting C#/.Net web apps for bugs/performance.
  • Solid knowledge of SaaS operations.
  • Knowledge of CI/CD pipelines.
  • Experience with IaC environments.
  • Excellent written and verbal communication skills.

Responsibilities

  • Leads post-incident investigations for the Site Reliability team.
  • Performs in-depth post-incident analyses to identify root causes and preventive strategies.
  • Drafts clear RCAs for customer delivery.
  • Cross-trains colleagues on observability tools during incidents and performance investigations.
  • Provides visibility to stakeholders throughout the SRE process.
  • Collaborates to implement system enhancements for scalability and stability.
  • Develops client dashboards/alerts to identify performance challenges.
  • Monitors and improves time-to-resolution metrics.
  • Maintains core observability tools and ensures data for incident response.
  • Provides feedback to Observability and Engineering to improve MELT and patterns.
  • Contributes to automation to streamline incident response.
  • Works to prevent incidents and reduce impact on our platform.
  • Partners with Cloud Operations, SRE, Engineering and the business.

Skills

SRE best practices
Coding: JavaScript/.NET/SQL
Cloud platforms: AWS/Azure
Public Cloud experience (AWS)
C#/.NET debugging
SaaS operations
CI/CD pipelines
IaC tooling
Communication skills

Education

Bachelor's degree in Computer Science or related field

Tools

New Relic
DataDog
SumoLogic

Job description

Origami Risk is seeking a Site Reliability Engineer to improve time-to-resolution and overall platform reliability for our SaaS solutions. You will lead post-incident investigations, identify root causes, and implement preventive measures while supporting client performance challenges and observability initiatives.

You will configure monitoring tools, develop dashboards, and collaborate with Cloud Operations, SRE, and Engineering to scale and stabilize services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE - Incident RCA & Observability (Remote)
SRE - Incident RCA & Observability (Remote)

Socket.dev • Georgia

Hybrid
USD 100,000 - 120,000
Medical and Dental coverage
Vision coverage
401(k) match
+8
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Georgia

Hybrid
USD 100,000 - 120,000
Medical and Dental coverage
Vision coverage
401(k) match
+8
Site Reliability Engineer
Site Reliability Engineer

Worky • Atlanta (GA)

Hybrid
USD 100,000 - 120,000
Medical & Dental
Hybrid/Remote work
Vision Insurance
+4
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Boston (MA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • United States

On-site
USD 120,000 - 180,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Competitive salary
Equity
Wellness benefits
+2
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Site Reliability Engineer — Incident & Deployment Expert
Site Reliability Engineer — Incident & Deployment Expert

Re Focus LLC • O’Fallon (MO)

On-site
USD 75,000 - 105,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000