SRE Engineer

SOMERSET STAFFING

Atlanta (GA)

On-site

USD 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SOMERSET STAFFING is seeking a production support engineer to manage AWS-based systems and incident response in Atlanta, GA. The role emphasizes reliability, on-call readiness, and rapid incident containment across EC2, VPC, RDS, Lambda, and EKS.

You will use CloudWatch, Dynatrace and other tools to monitor health, perform root-cause analysis, and drive CI/CD improvements across Linux environments. Strong communication with development teams is essential for timely resolutions.

Qualifications

  • Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.
  • Hands-on experience with incident management and 24/7 production support models.
  • Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.
  • Experience building and maintaining monitoring dashboards.
  • Strong troubleshooting skills across infrastructure, networking, and application layers.
  • Working knowledge of CI/CD pipelines and AWS deployment processes.
  • Experience working with databases and Unix/Linux environments.

Responsibilities

  • Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.
  • Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
  • Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
  • Participate in on-call rotations, major incident bridges, and post-incident reviews.
  • Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.
  • Monitor system health across applications, infrastructure, and AWS services.
  • Respond proactively to issues related to resource utilization, latency, errors, and availability.
  • Maintain and improve monitoring and observability dashboards.

Skills

CloudWatch
Dynatrace
Git
Observability
Reliability Patterns
Chaos Testing
Shell Scripting

Tools

EC2
VPC
ALB/NLB
RDS
Lambda
EKS

Job description

Job description


Qualifications

Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.


Hands-on experience with incident management and 24/7 production support models.


Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.


Experience building and maintaining monitoring dashboards.


Strong troubleshooting skills across infrastructure, networking, and application layers.


Working knowledge of CI/CD pipelines and AWS deployment processes.


Experience working with databases and Unix/Linux environments.


Key Responsibilities

Incident Management and Production Support

Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.


Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.


Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.


Participate in on-call rotations, major incident bridges, and post-incident reviews.


Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.


Monitoring and Operational Health

Perform regular health checks across applications, infrastructure, and AWS services.


Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes.


Respond proactively to s related to resource utilization, latency, errors, and availability.


Maintain and improve monitoring and observability dashboards.


Skills

Mandatory Skills : CloudWatch, Dynatrace, Git, Observability, Reliability Patterns


Good to Have Skills : Chaos Testing, Shell Scripting


Required Skills :

Basic Qualification :

Additional Skills :

Background Check : No


Drug Screen : No

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer
SRE Engineer

Blue Ribbon Global Technologies • Atlanta (GA)

On-site
USD 90,000 - 120,000
SRE Lead engineer
SRE Lead engineer

TechDigital Group • Bellevue (WA)

On-site
USD 100,000 - 130,000
AWS DevOps / SRE Engineer
AWS DevOps / SRE Engineer

TechDigital Group • Eagan (MN)

On-site
USD 90,000 - 130,000
SRE Specialist
SRE Specialist

TechDigital Group • Dallas (TX)

On-site
USD 100,000 - 130,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
SRE Engineer
SRE Engineer

ALLTECH CONSULTING SVC INC • Oregon (WI)

On-site
USD 90,000 - 120,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
SRE Production Support
SRE Production Support

SelectMinds LLC • Livonia (MI)

On-site
USD 100,000 - 140,000
SRE Engineer: AWS Incident & Observability Lead
SRE Engineer: AWS Incident & Observability Lead

SOMERSET STAFFING • Atlanta (GA)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Town of Florida (NY)

Hybrid
USD 150,000 - 190,000