Sr. Site Reliability Engineer

MeridianLink

United States

Remote

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MeridianLink is seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own reliability, scalability, and observability of our financial SaaS applications and infrastructure across cloud platforms, ensuring secure and high-performing services.

This high-impact role involves shaping SLOs/SLIs, building runbooks, implementing IaC, and advancing AIOps. You will mentor teams and drive incident response improvements in a regulated fintech environment.

Qualifications

  • 7+ years of SRE/DevOps with production systems responsibility.
  • Expert-level AWS and/or Azure experience.
  • Strong observability expertise with monitoring, logging, and tracing.
  • Proficiency in Python and scripting for automation.

Responsibilities

  • Define SLOs and SLIs for critical systems.
  • Design and deploy observability architectures.
  • Develop runbooks and post-incident reviews.
  • Architect and deploy cloud infrastructure on AWS or Azure.
  • Implement infrastructure-as-code and DR strategies.
  • Build automation and AIOps capabilities.
  • Conduct reliability reviews with cross-functional teams.
  • Write production-grade tooling in Python.
  • Ensure security and compliance in fintech environment.

Skills

SRE experience
AWS & Azure
Observability design
Python tooling
Infrastructure as Code
Incident management
AIOps
Monitoring & tracing

Tools

Prometheus
Grafana
ELK
Datadog
New Relic

Job description

About the Role

We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.

Key Responsibilities
  • Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently
  • Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior
  • Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems
  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity
  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives
  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure
  • Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments
  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows
  • Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment
Required Qualifications
  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems
  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale
  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)
  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization
  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes
  • Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operat
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler-Technologies-29572f8 • Lakewood (CO)

On-site
USD 93,547 - 150,000