Senior Cloud SRE: AWS, Serverless & Incident Response

Apply

Northern (KY)

Hybrid

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apply is seeking a Senior Site Reliability Engineer to own the day-to-day operations of our AWS-backed, fully serverless platform. You will manage accounts, databases, backups, and production monitoring to keep the system healthy and scalable.

You will lead incident response, build runbooks, and drive improvements in deployability and cost visibility, collaborating with the team to share knowledge and best practices.

Qualifications

  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.
  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring and alerting, incident response, and on-call ownership.

Responsibilities

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.

Skills

DevOps
Site reliability
AWS
Database administration
Scripting
Linux & networking
Monitoring & incident response

Education

Bachelor's degree

Tools

AWS
PostgreSQL
SST
Pulumi
Terraform
Terraform
CloudWatch
SQS
EventBridge
Lambda
S3
Docker
GitHub Actions

Job description

Apply is seeking a Senior Site Reliability Engineer to own the day-to-day operations of our AWS-backed, fully serverless platform. You will manage accounts, databases, backups, and production monitoring to keep the system healthy and scalable.

You will lead incident response, build runbooks, and drive improvements in deployability and cost visibility, collaborating with the team to share knowledge and best practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE: AWS Serverless & Reliability
Senior Cloud SRE: AWS Serverless & Reliability

MeridianLink • United States

Remote
USD 140,000 - 180,000
Senior Cloud SRE: AWS Serverless & Reliability Leader
Senior Cloud SRE: AWS Serverless & Reliability Leader

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Lead Site Reliability Engineer: AWS Cloud & Automation
Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Remote Senior SRE – AWS Cloud Reliability & Automation
Remote Senior SRE – AWS Cloud Reliability & Automation

IMPACT Technology Recruiting • New York (NY)

On-site
USD 120,000 - 160,000
Senior Cloud SRE & Reliability Engineer
Senior Cloud SRE & Reliability Engineer

Verygoodsecurity • United States

Hybrid
USD 110,000 - 140,000
Flexible work hours
Competitive health benefits
VGS stock options
+1
Remote SRE II: Cloud, Data Ops & Incident Response
Remote SRE II: Cloud, Data Ops & Incident Response

Cohere Health, Inc. • Boston (MA)

Hybrid
USD 100,000 - 110,000
Fully remote
5% travel
Medical insurance
+7
Senior Cloud & Reliability Engineer
Senior Cloud & Reliability Engineer

Very Good Security • United States

Hybrid
USD 100,000 - 130,000
Flexible PTO
Health benefits
401k with matching
+2