Senior Site Reliability Engineer

Linuxconfig

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Linuxconfig is seeking a Senior Site Reliability Engineer to join our cloud engineering team. This operations-focused role owns day-to-day administration of AWS accounts and databases, maintaining backup posture across data stores and providing production monitoring and debugging for a serverless platform.

You will lead incident response, manage on-call rotation, and continuously refine infrastructure as code (SST/Pulumi/Terraform) to keep deployments safe, scalable, and cost-visible, while

Qualifications

  • Bachelor's degree and 4–6 years of related experience or equivalent work experience.
  • 5+ years of DevOps, site reliability, or platform operations with significant responsibility for production systems.
  • 3+ years hands-on experience with AWS, emphasizing serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup/recovery, and query performance.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and tooling.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring, incident response, and on-call ownership.

Responsibilities

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and other data stores.
  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting.
  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures.
  • Continuously refine infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code accurate, retire unused infrastructure, and keep cost visible.
  • Share knowledge of production operations with the team, fostering a culture of learning and growth.

Skills

DevOps
Site Reliability
AWS
Monitoring
Incident response
Scripting
Linux
Docker
IaC
On-call

Education

Bachelor's degree

Tools

SST
Pulumi
Terraform
PostgreSQL
S3
CloudWatch
EventBridge
SQS

Job description

As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.

Responsibilities
  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.

  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.

  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.

  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.

  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.

  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.

Qualifications: Knowledge, Skills, & Abilities
  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.

  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.

  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).

  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.

  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.

  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).

  • Experience with production monitoring and alerting, incident response, and on-call ownership.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Senior CloudOps & Site Reliability Engineer
Senior CloudOps & Site Reliability Engineer

MangoApps • Seattle (WA)

On-site
USD 110,000 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler-Technologies-29572f8 • Lakewood (CO)

On-site
USD 93,547 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Cloud SRE — Serverless Ops & Reliability Lead
Senior Cloud SRE — Serverless Ops & Reliability Lead

Linuxconfig • Northern (KY)

Hybrid
USD 120,000 - 180,000