Sr. Site Reliability Engineer

MeridianLink, Inc.

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

MeridianLink, Inc. is seeking a Senior Site Reliability Engineer to own reliability and observability of critical financial SaaS applications. You will design SLOs, build and maintain runbooks, and lead incident response across AWS/Azure environments.

You will implement IaC, automate workflows in Python, and drive chaos engineering to prevent outages. The role demands strong scripting, deep cloud knowledge, and collaboration with cross-functional teams.

Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, or related roles with production systems responsibility.
  • Expert-level experience with AWS and/or Azure; deep knowledge of compute, networking, storage, and managed services.
  • Hands-on in observability: monitoring, alerting, logging, and distributed tracing; experience with full-stack observability tools.
  • Strong background in SLOs/SLIs/SLAs; ability to define objectives and build resilient systems.
  • Proven experience designing highly available, scalable systems and distributed architectures.
  • Proficiency in Python and shell scripting for automation and tooling; maintainable code.
  • Experience with AIOps practices and automated remediation is a plus.
  • Experience with IaC tools (Terraform, CloudFormation) and CI/CD pipelines.
  • Incident management and on-call ownership; ability to stay composed under pressure.
  • Excellent cross-functional communication; mentoring junior engineers.

Responsibilities

  • Design, implement, and maintain SLOs and SLIs for critical systems.
  • Lead observability strategy with monitoring, logging, and tracing across platforms.
  • Build runbooks, incident response procedures, and post-incident reviews; mentor team.
  • Architect and deploy cloud infrastructure on AWS or Azure; codify with IaC.
  • Develop automation and AIOps to reduce toil and enable self-healing systems.
  • Drive reliability improvements via load testing, chaos engineering, failure analysis.
  • Collaborate with application and backend teams on reliable system design.
  • Write production-grade Python tooling for automation and metrics.
  • Champion security and compliance in fintech infrastructure.

Skills

AWS
Azure
Cloud architecture
Observability tooling
Python scripting
IaC (Terraform, CloudFormation)
CI/CD pipelines
Incident management
On-call ownership

Tools

Prometheus
Grafana
ELK
Datadog
New Relic
Gremlin
Terraform
CloudFormation
Ansible

Job description

About the Role

We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.

Key Responsibilities
  • Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently

  • Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior

  • Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems

  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity

  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives

  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure

  • Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments

  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows

  • Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment

Required Qualifications
  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems

  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale

  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)

  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization

  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes

  • Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows

  • Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design

  • Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure

  • Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers

Preferred Qualifications
  • Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)

  • Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management

  • Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically

  • Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms

  • Contribution to open-source observability or infrastructure projects

  • Expertise in network security, application security, or infrastructure hardening

  • Experience with database optimization, query performance tuning, and backup/recovery strategies

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

On-site
USD 130,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000