Senior SRE: Cloud Reliability & Observability Leader

MeridianLink, Inc.

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MeridianLink, Inc. is seeking a Senior Site Reliability Engineer to join the cloud engineering team. You will own reliability, scalability, and observability of financial SaaS applications across cloud platforms, ensuring secure and performant services.

You will design SLOs/SLIs, lead observability initiatives, and architect cloud infra with IaC. The role emphasizes incident management, automation (Python), and strong collaboration with cross-functional teams.

Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems.
  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale.
  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (Prometheus, Grafana, ELK, Datadog, New Relic, or similar).
  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization.
  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes.
  • Proficiency in scripting (Python, PowerShell, Bash) for automation and tooling; ability to write clean, maintainable code for operational workflows.
  • Hands-on experience with AIOps: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus.
  • Experience with infrastructure-as-code tools (Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design.
  • Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure.
  • Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers.

Responsibilities

  • Design, implement, and maintain SLOs and SLIs across all critical systems; ensure targets are met.
  • Lead observability strategy with comprehensive monitoring, logging, and tracing architectures; deploy deep-visibility tools.
  • Build and own runbooks, incident response procedures, and post-incident reviews; mentor on incident management and blameless postmortems.
  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code and ensure high availability, DR, and business continuity.
  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems.
  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis.
  • Collaborate with application/backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments.
  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows.
  • Champion security and compliance in infrastructure; implement defense-in-depth for regulated fintech environments.

Skills

Site Reliability Engineering
DevOps
Platform engineering
Observability
SLOs/SLIs/SLAs
Python
PowerShell
Bash
Automation
AIOps
Infrastructure as code
CI/CD
Incident management
On-call ownership
Mentoring
Communication

Tools

Prometheus
Grafana
ELK
Datadog
New Relic
Kubernetes
Terraform
CloudFormation
Ansible

Job description

MeridianLink, Inc. is seeking a Senior Site Reliability Engineer to join the cloud engineering team. You will own reliability, scalability, and observability of financial SaaS applications across cloud platforms, ensuring secure and performant services.

You will design SLOs/SLIs, lead observability initiatives, and architect cloud infra with IaC. The role emphasizes incident management, automation (Python), and strong collaboration with cross-functional teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE: Observability, Automation & Resilience
Senior Cloud SRE: Observability, Automation & Resilience

MeridianLink • United States

Remote
USD 140,000 - 190,000
Remote Senior Site Reliability Engineer-Scale & Resilience
Remote Senior Site Reliability Engineer-Scale & Resilience

Far Coder • Northern (KY)

Hybrid
USD 25,000 - 40,000
Senior Cloud SRE: AWS Serverless & Reliability
Senior Cloud SRE: AWS Serverless & Reliability

MeridianLink • United States

Remote
USD 140,000 - 180,000
Senior Cloud SRE: AWS Serverless & Reliability Leader
Senior Cloud SRE: AWS Serverless & Reliability Leader

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE Manager: Cloud Reliability & Observability
Senior SRE Manager: Cloud Reliability & Observability

Peraton • Reston (VA)

On-site
USD 135,000 - 216,000
Senior SRE: Cloud Reliability & Automation Lead
Senior SRE: Cloud Reliability & Automation Lead

Ridgeline • New York (NY)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Educational reimbursements
$0 cost employee insurance plans
Senior Site Reliability Engineer — Cloud Observability
Senior Site Reliability Engineer — Cloud Observability

Guidehouse • San Antonio (TX)

On-site
USD 106,000 - 176,000
Medical, Rx, Dental & Vision Insurance
401(k) Retirement Plan
Tuition Reimbursement
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior SRE: Cloud Reliability & Incidents Lead
Senior SRE: Cloud Reliability & Incidents Lead

Illumio • San Jose (CA)

On-site
USD 120,000 - 150,000