Reliability Engineer

Chenega Professional Services Strategic Business Unit

Washington (District of Columbia)

On-site

USD 95,000 - 105,000

Full time

2 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Benefits program
Promotion opportunities
Team-oriented culture

Job summary

Chenega Services & Federal Solutions, LLC, a Chenega Professional Services company, seeks a Reliability Engineer to ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this role; you will design, build, review, and troubleshoot infrastructure-as-code for mission-critical environments.

You will collaborate with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve

Qualifications

  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
  • 2+ years of hands-on Microsoft Azure experience
  • 2+ years of hands-on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure
  • Experience administering/supporting Kubernetes (preferably AKS)
  • Experience supporting cloud-hosted systems and automating cloud operations
  • Knowledge of cloud networking concepts and troubleshooting
  • Possession of strong analytical and problem-solving abilities
  • Ability to work on-site in the Washington DC metro-area
  • Ability to obtain/maintain Public Trust/Suitability clearance
  • Bachelor’s degree or equivalent experience

Responsibilities

  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
  • Automate cloud operations to reduce manual work and improve consistency.
  • Collaborate with development and operations teams to enhance deployment and incident‑response practices.
  • Implement and refine monitoring, alerting, dashboards, and operational reporting.
  • Identify reliability risks and recommend improvements to cloud architecture and processes.
  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.

Skills

Azure cloud
Terraform
Kubernetes AKS
DevOps/SRE
Cloud monitoring
Networking concepts
Analytical thinking
On-site Washington DC

Education

Bachelor’s degree or equivalent experience

Tools

AKS

Job description

Summary


etails below are subject to change based on final award. Come join a company that strives for


Summary


Position contingent on contract award - d
etails below are subject to change based on final award. Come join a company that strives for Extraordinary People and Exceptional Performance! Chenega Services & Federal Solutions, LLC, a Chenega Professional Services’ company, is looking for a Reliability Engineer. In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission‑critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.

Our company offers employees the opportunity to join a team where there is a robust employee benefits program, management engagement, quality leadership, an atmosphere of teamwork, recognition for performance, and promotion opportunities. We actively strive to channel our highly engaged employee’s knowledge, critical thinking, innovative solutions for our clients.


Responsibilities


  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.

  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).

  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.

  • Automate cloud operations to reduce manual work and improve consistency.

  • Collaborate with development and operations teams to enhance deployment and incident‑response practices.

  • Implement and refine monitoring, alerting, dashboards, and operational reporting.

  • Identify reliability risks and recommend improvements to cloud architecture and processes.

  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.


Qualifications


  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles

  • 2+ years of hands‑on Microsoft Azure experience

  • 2+ years of hands‑on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure

  • Experience administering/supporting Kubernetes (preferably AKS)

  • Experience supporting cloud‑hosted systems and automating cloud operations

  • Knowledge of cloud networking concepts and troubleshooting

  • Possession of strong analytical and problem‑solving abilities

  • Ability to work on-site in the Washington DC metro‑area

  • Ability to obtain/maintain Public Trust/Suitability clearance

  • Bachelor’s degree or equivalent experience


Nice To Haves


  • Experience with certificate lifecycle management (maintaining, renewing, rotating certs)

  • Experience creating operational dashboards or reports (Power BI preferred)

  • Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)

  • Experience integrating Azure resources with Active Directory

  • Experience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automation

  • Azure, Kubernetes, or Terraform certifications


Estimated Salary/Wage

USD $95,000.00/Yr. Up to USD $105,000.00/Yr.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer
Reliability Engineer

Chenega Professional Services Strategic Business Unit • Atlanta (GA)

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Teamwork culture
Reliability Engineer
Reliability Engineer

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Washington

On-site
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
Azure Reliability Engineer & Terraform Automation Lead
Azure Reliability Engineer & Terraform Automation Lead

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
Azure Reliability Engineer — Terraform & AKS Specialist
Azure Reliability Engineer — Terraform & AKS Specialist

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

GAP SOLUTIONS INC • Atlanta (GA)

On-site
USD 110,000 - 160,000
On-site in Atlanta
Disability accommodations
Public Trust clearance support
Reliability Engineer (onsite)
Reliability Engineer (onsite)

System One • Atlanta (GA)

On-site
USD 110,000 - 150,000
Health benefits
401(k) plan
Azure Reliability Engineer: Terraform & AKS
Azure Reliability Engineer: Terraform & AKS

Chenega Professional Services Strategic Business Unit • Atlanta (GA)

On-site
USD 95,000 - 105,000
Benefits program
Promotion opportunities
Teamwork culture
Reliability Engineer (52718)
Reliability Engineer (52718)

GAP Solutions, Inc. • Atlanta (GA)

Hybrid
USD 90,000 - 140,000