Senior SRE / DevOps Engineer (Atlanta)

Broad Reach Partners

Alpharetta (GA)

Hybrid

USD 140,000 - 145,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Vision insurance
401(k)

Job summary

A staffing and recruiting firm is seeking a Site Reliability Engineer to improve the stability and reliability of production systems in Alpharetta, GA. In this full-time role, you will work with development, DevOps, and security teams on automating deployments and enhancing cloud infrastructure. Ideal candidates should possess over 8 years of SRE experience, with a strong proficiency in AWS and Kubernetes. The position requires troubleshooting skills and the ability to prioritize efficiently while working independently.

Qualifications

  • 8+ years of experience as a Site Reliability Engineer, or equivalent.
  • 3+ years of experience with AWS or Azure.
  • Proficiency with public cloud environments (AWS preferred).

Responsibilities

  • Maintain and enhance monitoring tools for service health and performance metrics.
  • Implement and maintain high-availability systems with capacity planning and fault tolerance.
  • Define and monitor Service Level Indicators and Objectives.

Skills

Site Reliability Engineering
Monitoring tools (New Relic, Graylog)
Kubernetes
Amazon Web Services (AWS)
Scripting languages (Bash, Groovy, Python)
Debugging and troubleshooting
CI/CD pipelines

Tools

New Relic
AWS EKS
Terraform
Ansible

Job description

Base pay range

$140,000.00/yr - $145,000.00/yr

We are seeking a Site Reliability Engineer to join our team in Atlanta and play a key role in enhancing the stability, performance, and reliability of our production systems. You’ll work closely with development, DevOps, and security teams to improve observability, optimize system performance, and ensure production readiness. From monitoring to automation, you’ll make a direct impact on our cloud infrastructure and service reliability.

In this role, you will work hand‑in‑hand with our development, operations, and security teams worldwide to implement best practices, automate deployments, and ensure our platforms are reliable, secure, and scalable. Troubleshooting in Kubernetes is required, which will involve a deep understanding of pods, nodes, networking, scaling, logs, and service‑to‑service communication.

This role requires a deep understanding of SRE best practices and a strong ability to troubleshoot complex issues.

Responsibilities
  • Maintain and enhance monitoring tools (New Relic, Graylog) for service health and performance metrics.
  • Implement and maintain high‑availability systems with capacity planning, performance optimization, and fault tolerance.
  • Define and monitor Service Level Indicators, Objectives, and Agreements with teams.
  • Deploy and manage Kubernetes workloads to AWS EKS using Helm, ArgoCD.
  • Automate operational processes to reduce manual interventions.
  • Manage Kubernetes workloads on AWS EKS for secure and stable deployments.
  • Participate in on‑call rotation, troubleshoot production issues, and implement permanent fixes.
  • Work with DevOps to improve CI/CD pipelines and with development teams to embed resilience and observability.
  • Document operational runbooks, escalation procedures, and production playbooks.
Required Skills and Experience
  • 8+ years of experience as a Site Reliability Engineer, or equivalent.
  • Experience with tools like New Relic for monitoring and Graylog for logging.
  • 3+ years of experience with Amazon Web Services (AWS) or Microsoft Azure.
  • 3+ years of experience with Kubernetes clusters – performance monitoring in Kubernetes.
  • Proficiency with public cloud environments (AWS preferred).
  • Proficiency in a scripting language, like Bash, Groovy, Python.
  • Excellent debugging and troubleshooting skills.
  • Ability to prioritize tasks efficiently and independently under minimal supervision.
Nice to Have
  • Familiarity with .NET applications.
  • Knowledge in Terraform, Ansible, monitoring tools.

This is a full‑time role and unfortunately we can’t sponsor, so you must be a US citizen or a green‑card holder. You must currently live in the Atlanta area as you will need to come into our Atlanta office one or two times each month for key meetings with our team.

If you thrive on solving complex technical challenges, have a passion for automation, and want to influence how enterprise platforms evolve and modernize, this is an ideal opportunity for you.

Seniority Level

Mid‑Senior level

Employment Type

Full‑time

Job Function

Information Technology

Industries

Staffing and Recruiting

Referrals increase your chances of interviewing at Broad Reach Partners by 2x

Benefits
  • Medical insurance
  • Vision insurance
  • 401(k)
Contact

Get notified about new Site Reliability Engineer jobs in Alpharetta, GA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment • Alpharetta (GA)

Hybrid
Medical Insurance
Health Savings Account
401(k)
+1
AWS Site Reliability Engineer (SRE)
AWS Site Reliability Engineer (SRE)

STAFFWORXS • Atlanta (GA)

Hybrid
USD 72,000 - 195,000
Site Reliability Engineer
Site Reliability Engineer

Prestige Staffing • Atlanta (GA)

On-site
USD 130,000 - 150,000
Vision insurance
401(k)
Paid maternity leave
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

STAFFWORXS • Atlanta (GA)

Hybrid
Sr. Software Engineer (SRE)
Sr. Software Engineer (SRE)

Flexton Inc. • Atlanta (GA)

Hybrid
USD 110,000 - 140,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment • Atlanta (GA)

On-site
USD 93,000 - 158,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Robotics Prcocess Automation, LLC • Atlanta (GA)

On-site
USD 100,000 - 150,000