Senior AWS Site Reliability Engineer (SRE)

System One

Birmingham (AL)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
401(k) plan
Life insurance
Voluntary plans

Job summary

System One is seeking an experienced Senior AWS Site Reliability Engineer (SRE) to join our Birmingham, Alabama operations onsite on a contract-to-hire basis. You will define and monitor SLIs/SLOs, build resilient infrastructure, and lead reliability improvements across AWS hosted applications.

Ideal candidates have 8+ years of AWS experience, extensive Terraform IaC usage, and strong Python/Java automation.

Qualifications

  • 8+ years of hands on AWS engineering experience in enterprise production environments.
  • Deep knowledge of ECS, EC2, RDS, Lambda, Route 53, Step Functions, Redshift, EMR, DynamoDB, S3, CloudWatch.
  • Proven SRE practices: SLIs/SLOs, error budgets, monitoring, alerting, incident response, resiliency.
  • Strong Terraform / IaC experience for infrastructure provisioning.
  • Experience building CI/CD pipelines with Jenkins, GitLab, or similar tools.
  • Programming/automation skills in Python or Java.

Responsibilities

  • Implement and improve SRE practices across AWS hosted enterprise applications and services.
  • Define and monitor SLIs/SLOs, error budgets, alarms, availability, and reliability metrics.
  • Automate infrastructure and operational processes using Python or Java.
  • Develop and maintain infrastructure as code using Terraform and AWS native tooling.
  • Build and enhance CI/CD pipelines using Jenkins, GitLab, and related technologies.
  • Implement monitoring, APM, distributed tracing, dashboards, and alerting across logs, metrics, traces, and dashboards.
  • Improve application resilience, failover readiness, recovery automation, and production stability.
  • Lead technical discussions and coordinate resolution across teams.

Skills

AWS
SRE practices
Terraform IaC
CI/CD pipelines
Python/Java
Observability
CloudWatch
Troubleshooting

Tools

Jenkins
GitLab
OpenTelemetry
Splunk

Job description

Job Title: Senior AWS Site Reliability Engineer (SRE)

Location: Birmingham, Alabama

Type: Contract To Hire

Work Model: Onsite – onsite

Hours: 40.0

Security Clearance:

Overview

Responsibilities
  • Implement and improve SRE practices across AWS hosted enterprise applications and services.
  • Define and monitor SLIs/SLOs, error budgets, alarms, availability, and reliability metrics.
  • Automate infrastructure and operational processes using Python or Java.
  • Develop and maintain infrastructure as code using Terraform and AWS native tooling.
  • Build and enhance CI/CD pipelines using Jenkins, GitLab, and related technologies.
  • Implement monitoring, APM, distributed tracing, dashboards, and alerting using Splunk, SignalFx, OpenTelemetry, and CloudWatch.
  • Improve application resilience, failover readiness, recovery automation, and production stability.
  • Analyze application and infrastructure performance, identify recurring reliability issues, and implement sustainable improvements.
  • Troubleshoot AWS workloads, application performance, data pipelines, networking/DNS, CI/CD, and infrastructure automation issues.
  • Support production incidents, releases, platform upgrades, vulnerability remediation, and controlled infrastructure changes.
  • Establish performance baselines and help improve monitoring and operational readiness.
  • Lead technical discussions and coordinate resolution across application, cloud, DevOps, security, performance, and production support teams.
Requirements
  • 8+ years of hands on AWS engineering experience supporting enterprise production environments.
  • Experience across AWS services such as ECS, EC2, RDS, Lambda, Route 53, Step Functions, Redshift, EMR, DynamoDB, S3, and CloudWatch.
  • Proven implementation of core Site Reliability Engineering (SRE) practices, including SLIs/SLOs, error budgets, monitoring, alerting, incident response, reliability, and resiliency.
  • Strong Terraform / Infrastructure as Code (IaC) experience.
  • Experience building and supporting CI/CD pipelines with Jenkins, GitLab, or comparable tooling.
  • Strong programming and automation skills using Python or Java.
  • Hands on observability experience using Splunk, SignalFx, OpenTelemetry, and CloudWatch across logs, metrics, traces, dashboards, and alerts.
  • Experience with APM and distributed tracing in enterprise applications.
  • Strong production troubleshooting skills across application, cloud infrastructure, performance, deployment, and reliability issues.
  • Experience implementing resilience, recovery, failover, and production stability improvements.
  • Ability to troubleshoot complex AWS environments, including unhealthy workloads, failed serverless workflows, database/data pipeline issues, routing/DNS problems, capacity constraints, and automation failures.
  • Strong ownership, analytical troubleshooting, communication, and cross team collaboration skills.

System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.

#M-#LI-

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AWS SRE: Reliability, IaC & CI/CD Champion
Senior AWS SRE: Reliability, IaC & CI/CD Champion

System One • Birmingham (AL)

On-site
USD 120,000 - 180,000
Health insurance
Dental insurance
Vision insurance
+3
AWS Application Production Support Engineer
AWS Application Production Support Engineer

System One • Columbia (SC)

On-site
USD 90,000 - 130,000
Health and welfare benefits
401(k) plan
Sr. Software Engineer (SRE)
Sr. Software Engineer (SRE)

Flexton Inc. • Atlanta (GA)

Hybrid
USD 110,000 - 140,000
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Software Integration Engineer
Software Integration Engineer

System One • Corridor North (MD)

On-site
USD 120,000 - 180,000
L2 Production Support Specialist with AWS
L2 Production Support Specialist with AWS

System One • Lafayette (LA)

On-site
USD 75,000 - 110,000
AWS Production Support Engineer
AWS Production Support Engineer

System One • Birmingham (AL)

On-site
USD 68,000 - 86,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Randstad Digital Americas • Plano (TX)

On-site
USD 115,000 - 125,000
Medical insurance
401K plan