Reliability Engineer

GAP SOLUTIONS INC

Atlanta (GA)

On-site

USD 110,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

On-site in Atlanta
Disability accommodations
Public Trust clearance support

Job summary

GAP Solutions INC seeks a Reliability Engineer to design, build, and troubleshoot Azure infrastructure using Terraform and to improve Kubernetes operations in AKS. The role collaborates with platform engineers, developers, security teams, and stakeholders to automate cloud infrastructure and enhance incident response.

You will implement monitoring, dashboards, and runbooks while identifying reliability risks and applying architectural improvements.

Qualifications

  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
  • 2+ years of hands‑on Microsoft Azure experience
  • 2+ years of hands‑on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure
  • Experience administering/supporting Kubernetes (preferably AKS)
  • Experience supporting cloud‑hosted systems and automating cloud operations
  • Knowledge of cloud networking concepts and troubleshooting
  • Possession of strong analytical and problem‑solving abilities
  • Ability to work on-site in Atlanta, GA or the Washington D.C. metro-area
  • Bachelor’s degree or equivalent experience
  • Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder)
  • Experience with certificate lifecycle management (maintaining, renewing, rotating certs)
  • Experience creating operational dashboards or reports (Power BI preferred)
  • Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)
  • Experience integrating Azure resources with Active Directory
  • Experience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automation
  • Azure, Kubernetes, or Terraform certifications

Responsibilities

  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform
  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS)
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues
  • Automate cloud operations to reduce manual work and improve consistency
  • Collaborate with development and operations teams to enhance deployment and incident response practices
  • Implement and refine monitoring, alerting, dashboards, and operational reporting
  • Identify reliability risks and recommend improvements to cloud architecture and processes
  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks

Skills

Azure
Terraform
Kubernetes
AKS
Automation
CI/CD
Monitoring
Power BI
Grafana
Prometheus
Active Directory
GitHub Copilot

Education

Bachelor's degree or equivalent experience

Tools

Grafana
Prometheus
Elastic
Splunk
Power BI
Active Directory
GitHub Copilot
Azure DevOps

Job description

Position Objective: In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.Duties and Responsibilities:Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.Automate cloud operations to reduce manual work and improve consistency.Collaborate with development and operations teams to enhance deployment and incident response practices.Implement and refine monitoring, alerting, dashboards, and operational reporting.Identify reliability risks and recommend improvements to cloud architecture and processes.Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.Basic Qualifications:4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles2+ years of hands‑on Microsoft Azure experience2+ years of hands‑on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructureExperience administering/supporting Kubernetes (preferably AKS)Experience supporting cloud‑hosted systems and automating cloud operationsKnowledge of cloud networking concepts and troubleshootingPossession of strong analytical and problem‑solving abilitiesAbility to work on-site in Atlanta, GA or the Washington D.C. metro-areaAbility to obtain/maintain Public Trust/Suitability clearanceBachelor’s degree or equivalent experienceMust be U.S. Citizen or Lawful Permanent Resident (Green Card Holder)Preferred Qualifications:Experience with certificate lifecycle management (maintaining, renewing, rotating certs)Experience creating operational dashboards or reports (Power BI preferred)Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)Experience integrating Azure resources with Active DirectoryExperience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automationAzure, Kubernetes, or Terraform certifications*This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required by this position.To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed above are representative of the knowledge, skill, and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.GAP Solutions provides reasonable accommodations to qualified individuals with disabilities. If you need an accommodation to apply for a job, email us at recruiting@gapsi.com. You will need to reference the requisition number of the position in which you are interested. Your message will be routed to the appropriate recruiter who will assist you. Please note, this email address is only to be used for those individuals who need an accommodation to apply for a job. Emails for any other reason or those that do not include a requisition number will not be returned.Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status or other characteristics protected by law.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer (52718)
Reliability Engineer (52718)

GAP Solutions, Inc. • Atlanta (GA)

Hybrid
USD 90,000 - 140,000
Reliability Engineer
Reliability Engineer

Chenega Professional Services Strategic Business Unit • Atlanta (GA)

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Teamwork culture
Reliability Engineer (onsite)
Reliability Engineer (onsite)

System One • Atlanta (GA)

On-site
USD 110,000 - 150,000
Health benefits
401(k) plan
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Washington

On-site
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Professional Services Strategic Business Unit • Washington

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Team-oriented culture
Azure Reliability Engineer — Terraform & AKS Specialist
Azure Reliability Engineer — Terraform & AKS Specialist

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Azure Reliability Engineer & Terraform Automation Lead
Azure Reliability Engineer & Terraform Automation Lead

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
Azure Reliability Engineer: Terraform & AKS
Azure Reliability Engineer: Terraform & AKS

Chenega Professional Services Strategic Business Unit • Atlanta (GA)

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Teamwork culture