Senior Site Reliability Engineer

MeridianLink

United States

Remote

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MeridianLink is seeking a Senior Site Reliability Engineer to own the day-to-day operations of our cloud platform. You will manage AWS accounts, databases, backups, and production monitoring to keep our fully serverless stack healthy and scalable.

You will drive reliability improvements across deployment, incident response, and runbooks, partnering with engineering to automate, optimize costs, and ensure rapid restoration after outages while maintaining secure, compliant infrastructure.

Qualifications

  • Bachelor's degree or equivalent in a related field.
  • 5+ years in DevOps/SRE/operations with production systems.
  • 3+ years hands-on AWS experience with serverless services.
  • Strong database administration and backup/recovery experience.
  • Scripting in TypeScript, Python, or Bash for automation.
  • Familiarity with Linux, networking, TLS, and IaC (SST, Pulumi, Terraform).

Responsibilities

  • Own day-to-day AWS accounts, services, and access management.
  • Manage backups across databases and S3, and disaster recovery.
  • Monitor production with CloudWatch and alerting; address issues proactively.
  • Lead incident response, runbooks, and on-call rotation.
  • Improve infrastructure as code for deployability, cost visibility, and scalability.
  • Share knowledge to foster learning and growth.

Skills

DevOps
SRE
AWS
Serverless
Automation
Incident Response
Linux
Monitoring

Education

Bachelor's degree

Tools

SST
Pulumi
Terraform
GitHub Actions
CloudWatch
PostgreSQL
S3
Docker

Job description

As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.

Responsibilities
  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.
Qualifications:
Knowledge, Skills, Abilities
  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.
  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring and alerting, incident response, and on-call ownership.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Senior CloudOps & Site Reliability Engineer
Senior CloudOps & Site Reliability Engineer

MangoApps • Seattle (WA)

On-site
USD 110,000 - 150,000
DevOps / SRE Cloud Engineer
DevOps / SRE Cloud Engineer

Compunnel, Inc. • Irving (TX)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 190,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2