Reliability Engineer (onsite)

System One

Atlanta (GA)

On-site

USD 110,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
401(k) plan

Job summary

System One seeks a Reliability Engineer to ensure the reliability, scalability, and operational health of Azure-based infrastructure. You will design, build, and troubleshoot Terraform IaC for mission-critical environments, collaborating with platform engineers, developers, and security teams to automate cloud operations and improve AKS reliability.

The role requires 4+ years in cloud infrastructure and hands-on Azure and Terraform expertise, with on-site work in Atlanta, GA or the Washington DC

Qualifications

  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles.
  • 2+ years of hands on Microsoft Azure experience.
  • 2+ years of hands on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure.
  • Experience administering/supporting Kubernetes (preferably AKS).
  • Experience supporting cloud hosted systems and automating cloud operations.
  • Knowledge of cloud networking concepts and troubleshooting.
  • Possession of strong analytical and problem-solving abilities.
  • Ability to work on-site in Atlanta, GA or the Washington DC metro-area.
  • Ability to obtain/maintain Public Trust/Suitability clearance.
  • Bachelor’s degree or equivalent experience.
  • Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder).

Responsibilities

  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
  • Automate cloud operations to reduce manual work and improve consistency.
  • Collaborate with development and operations teams to enhance deployment and incident response practices.
  • Implement and refine monitoring, alerting, dashboards, and operational reporting.
  • Identify reliability risks and recommend improvements to cloud architecture and processes.
  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.

Skills

Cloud infrastructure
DevOps
SRE
Azure
Terraform
Kubernetes
Networking
Analytical thinking
Security clearance

Education

Bachelor’s degree or equivalent experience

Tools

Terraform
Kubernetes
Azure

Job description

Reliability Engineer

Washington, DC or Atlanta, GA - ONSITE

$130,000.00

Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder) per government contract

Security Clearance: Public Trust

In this role, the engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.

Job Responsibilities:
  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
  • Automate cloud operations to reduce manual work and improve consistency.
  • Collaborate with development and operations teams to enhance deployment and incident response practices.
  • Implement and refine monitoring, alerting, dashboards, and operational reporting.
  • Identify reliability risks and recommend improvements to cloud architecture and processes.
  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.
Job Requirements:
  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
  • 2+ years of hands on Microsoft Azure experience
  • 2+ years of hands on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure
  • Experience administering/supporting Kubernetes (preferably AKS)
  • Experience supporting cloud hosted systems and automating cloud operations
  • Knowledge of cloud networking concepts and troubleshooting
  • Possession of strong analytical and problem-solving abilities
  • Ability to work on-site in Atlanta, GA or the Washington DC metro-area
  • Ability to obtain/maintain Public Trust/Suitability clearance
  • Bachelor’s degree or equivalent experience
  • Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder)
Job Desirables:
  • Experience with certificate lifecycle management (maintaining, renewing, rotating certs)
  • Experience creating operational dashboards or reports (Power BI preferred)
  • Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)
  • Experience integrating Azure resources with Active Directory
  • Experience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automation
  • Azure, Kubernetes, or Terraform certifications

System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.

#M-1

#LI-VH1

Ref: #851-Rockville-S1

System One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Professional Services Strategic Business Unit • Atlanta (GA)

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Teamwork culture
Azure Reliability Engineer — Terraform & AKS Automation (Onsite)
Azure Reliability Engineer — Terraform & AKS Automation (Onsite)

System One • Atlanta (GA)

On-site
USD 110,000 - 150,000
Health benefits
401(k) plan
Azure Cloud Engineer
Azure Cloud Engineer

System One • Washington

On-site
USD 77,000 - 87,000
Site Reliability Engineer (DevOps)
Site Reliability Engineer (DevOps)

Accenture Federal Services • Reston (VA)

On-site
USD 111,000 - 222,000
Sr. Cloud Engineer (Azure)
Sr. Cloud Engineer (Azure)

Eliassen Group • Hagåtña (GU)

On-site
USD 140,000 - 160,000
Azure Cloud Engineer
Azure Cloud Engineer

World Wide Technology, Inc. • Washington, Northern (KY)

Hybrid
USD 145,000 - 165,000
Health, Dental, Vision
Onsite Health Centers
Employee Assistance Program
+10
Cloud Engineer - Secret Clearance Required
Cloud Engineer - Secret Clearance Required

System One • Washington

On-site
USD 120,000 - 160,000
Health and welfare benefits
401(k) plan
Enterprise - DevOps Engineer - Terraform, Kubernetes, AWS
Enterprise - DevOps Engineer - Terraform, Kubernetes, AWS

Erias Ventures, LLC • Maryland

On-site
USD 175,000 - 265,000
Above market pay
11% Roth/Traditional 401k with vesting
Spot bonuses
+8