Site Reliability Engineer

The Amatriot Group

Norfolk (VA)

On-site

USD 145,000 - 165,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

The Amatriot Group is seeking a Site Reliability Engineer to develop and execute tests focused on resilience, performance, and failure scenarios in a Navy-related environment. You will work with development and operations teams to automate testing and ensure services meet SLIs/SLOs.

You will support, migrate, automate, and optimize CI/CD, IaC, and identity management across the Navy Enterprise Network, contributing to the maturity of the SRE program and platform reliability in a 100% onsite

Qualifications

  • Active DoD Secret security clearance required.
  • Experience with Active Directory, DNS, DHCP, and WINS.
  • CI/CD toolsets such as Jenkins or GitLab.
  • On-site 100% work in Norfolk, VA.
  • Scripting in PowerShell or Python.

Responsibilities

  • Develop and execute test strategies that simulate real-world failure scenarios.
  • Create, script, and run automated test suites for infrastructure and application components.
  • Ensure testing is integrated into the CI/CD pipeline to validate reliability with each release.
  • Test, patch, and upgrade Active Directory, Azure, Delinea, Ansible, and related services.
  • Coordinate with project teams to support modernization of the Navy Enterprise Network.
  • Maintain and patch systems to meet STIG requirements and security baselines.
  • Validate monitoring, logging, and alerting mechanisms under failure conditions.
  • Support ongoing activities to mature the SRE program and platform reliability.

Skills

Agile methodologies
Linux/Unix command line
SRE/DevOps concepts
CI/CD tooling
Jira/Confluence
PowerShell/Python scripting

Education

Bachelor's degree
Master’s degree (with <2 years experience)

Tools

Jira
GitLab
Ansible
Terraform
CloudFormation
OpenShift/Kubernetes

Job description

Clearance: Secret Clearance

Location: Norfolk, VA

Job Type: Full-Time (on-call support as needed)

Target Salary Range*: $145,000 - 165,000

  • This represents the potential salary range for this position depending on education level, years of experience and/or certifications in addition to other position specific requirements which may impact salary
Position Overview

The Site Reliability Engineer develops and executes tests focused on system resilience, performance under load, and failure scenarios. This role works with Site Reliability Engineers, development teams, and operations teams to create automated testing frameworks that simulate real-world conditions, validate system behavior under normal and stress conditions, and ensure services are resilient and meet established service level objectives.

The Site Reliability Engineer supports, migrates, automates, and optimizes software development and deployment processes, Infrastructure as Code, and identity and access management capabilities. This role contributes to the maturity of the Site Reliability Engineering program and supports the creation, maintenance, modernization, and refresh of capabilities and components for the Navy Enterprise Network.

Key Responsibilities
SRE Testing and Reliability Engineering
  • Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
  • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic.
  • Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components.
  • Ensure testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with Site Reliability Engineers, developers, and operations teams to define reliability goals and develop testing strategies to validate those goals.
  • Ensure new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
  • Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
  • Ensure Service Level Indicators and Service Level Objectives are accurately measured and tracked through automated testing frameworks.
Platform Administration and Support
  • Test, maintain, patch, STIG, upgrade, troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager, Active Directory Federation Services, DHCP, DNS, WINS, Group Policy Objects, and PKI.
  • Support the creation, maintenance, update, modernization, and refresh of capabilities and components of the Navy Enterprise Network.
  • Provide technical leadership and related task knowledge.
  • Coordinate with project managers, customers, stakeholders, and engineers to support ongoing activities and new projects that maintain, transform, and modernize the Navy Enterprise Network.
Automation, CI/CD, and Infrastructure as Code
  • Write code to automate software releases, monitor systems, and detect and resolve problems before users are impacted.
  • Develop features using AI coding tools and script repositories to automate, scale, test, and secure cloud infrastructure and pipelines.
  • Develop and code high-quality pipeline automation workflows to support environments inside and outside the cloud platform.
  • Ensure automation workflows align with business and technology strategies.
  • Contribute to ongoing SRE maturity by recommending improvements to engineering build, maintenance, automation, and reliability across the platform using SRE/DevOps tools and Infrastructure as Code.
Monitoring, Performance, and Platform Optimization
  • Work with development and operations teams to support fast and reliable software deployments.
  • Monitor systems and improve overall platform reliability.
  • Discover, document, and resolve system bugs.
  • Enhance performance monitoring of systems using Splunk or other dashboard reporting tools.
  • Identify performance bottlenecks and optimize cloud infrastructure performance.
  • Maintain complex computer systems through automation, monitoring, and proactive issue resolution.
Operational Support and Escalation
  • Resolve most conflicts between timeline, budget, and scope independently.
  • Escalate sophisticated or consequential issues to senior management.
  • Work nights, weekends, and provide on-call support as needed.
Qualifications
Education
  • B.S. degree and 2–4 years of prior relevant experience; or
  • Master’s degree with less than 2 years of relevant experience.
Experience
  • Experience designing, configuring, and managing Active Directory, Group Policy Objects, DNS, DHCP, and WINS.
  • Experience with Windows Server 2016, 2019, 2022, or 2025.
  • Experience with SQL Server 2019, 2022, or 2025.
  • Experience with CI/CD toolsets, such as Jenkins or GitLab.
  • Experience in application administration, configuration, and integration.
  • Experience creating Jira and/or Azure DevOps workflows, projects, and custom configurations.
  • Experience administering and maintaining an SRE platform using Ansible playbooks, such as upgrading Jenkins.
  • Experience automating tasks with scripting languages such as PowerShell or Python.
  • Experience integrating and maintaining third-party CI/CD tools such as Jenkins and GitLab.
  • Experience with PaaS using Red Hat OpenShift, Kubernetes, and Docker containers.
  • Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
  • Experience with automated provisioning and configuration tools such as Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
  • Working knowledge of the Risk Management Framework and DISA STIGs.
Skills
  • Familiarity with Agile development methodologies.
  • Good command of Linux/Unix and command-line concepts.
  • Automated script design, coding, debugging, and maintenance skills using Bash, Python, or similar languages.
  • Knowledge of Agile, DevSecOps, and SRE concepts and best practices, with a desire to grow that knowledge.
  • Hands-on experience with Atlassian products, including Jira, Confluence, and Bitbucket.
  • Ability to work with a distributed team.
  • Ability to work in a highly collaborative, forward-thinking, and innovation-driven environment.
Certifications
  • DoD 8570.01 IAT Level II certification required prior to onboarding and must be maintained while supporting the SMIT Contract. [Required]
Clearance
  • Active DoD Secret security clearance, with ability to maintain the clearance. [Required]
Other Requirements
  • Must be able to support program execution in classified environments and access SIPRNet from an NMCI location. [Required]
  • 100% onsite work is required. [Required]
  • Must be willing to work nights, weekends, and provide on-call support as needed. [Required]
Preferred Qualifications
  • Previous work experience providing support to the NGEN-NMCI program.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation for automating test environments.
  • ITILv4, Scrum Master, or Agile SAFe certification, or applicable experience.
  • Familiarity with designing, configuring, and managing FIM 2010 R2/MIM 2016 synchronization.
  • Familiarity with VB.NET.
  • Familiarity with designing, configuring, and managing Delinea.
  • Familiarity with Microsoft Active Directory Federation Services structure.
  • Familiarity with cloud engineering.
  • Familiarity with Public Key Infrastructure certificates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Amatriot Group, LLC • Norfolk (VA)

On-site
USD 145,000 - 165,000
Systems Engineer – DevSecOps/SRE (w/ active Secret)
Systems Engineer – DevSecOps/SRE (w/ active Secret)

CriticalSolutions, LLC • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical coverage
Dental coverage
401K matching
+1
Systems Engineer – DevSecOps/SRE (w/ active Secret)
Systems Engineer – DevSecOps/SRE (w/ active Secret)

Critical Solutions • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical coverage
Dental coverage
Vision coverage
+4
Network Site Reliability Engineer (w/ active Secret)
Network Site Reliability Engineer (w/ active Secret)

CriticalSolutions, LLC • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical Insurance
Dental Insurance
Vision Insurance
+4
Network Site Reliability Engineer (w/ active Secret)
Network Site Reliability Engineer (w/ active Secret)

Critical Solutions • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical and dental insurance
Vision insurance
Life Insurance
+2
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

The Amatriot Group • San Diego (CA)

On-site
USD 145,000 - 165,000
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

The Amatriot Group • Honolulu (HI)

On-site
USD 145,000 - 165,000
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

The Amatriot Group • Norfolk (VA)

On-site
USD 145,000 - 165,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Eliassen Group • Norfolk (VA)

On-site
Confidential
Medical, Dental, Vision
401k with company matching
Life insurance