Site Reliability Engineer

Amatriot Group, LLC

Norfolk (VA)

On-site

USD 145,000 - 165,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Amatriot Group, LLC in Norfolk, VA seeks a Site Reliability Engineer to design and implement resilient systems, automated testing, and robust CI/CD pipelines. You will work with development and operations teams to improve reliability and performance, and help migrate and modernize Navy Enterprise Network capabilities.

The role requires hands-on experience with AD/Azure/Ansible, cloud platforms, IaC, and monitoring tools.

Qualifications

  • Experience designing, configuring, and managing Active Directory, Group Policy Objects, DNS, DHCP, and WINS.
  • Experience with Windows Server 2016, 2019, 2022, or 2025.
  • Experience with CI/CD toolsets, such as Jenkins or GitLab.
  • Experience in application administration, configuration, and integration.
  • Experience creating Jira and/or Azure DevOps workflows, projects, and custom configurations.
  • Experience administering and maintaining an SRE platform using Ansible playbooks, such as upgrading Jenkins.
  • Experience automating tasks with scripting languages such as PowerShell or Python.
  • Experience integrating and maintaining third-party CI/CD tools such as Jenkins and GitLab.
  • Experience with PaaS using Red Hat OpenShift, Kubernetes, and Docker containers.
  • Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
  • Experience with automated provisioning and configuration tools such as Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
  • Working knowledge of the Risk Management Framework and DISA STIGs.

Responsibilities

  • Develop and execute test strategies that simulate real-world failure scenarios.
  • Create, script, and run performance tests to measure system behavior under varying levels of load.
  • Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components.
  • Ensure testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with Site Reliability Engineers, developers, and operations teams to define reliability goals and develop testing strategies.
  • Ensure new services and features undergo thorough testing for performance, reliability, and failure recovery before production deployment.
  • Validate that monitoring, logging, and alerting mechanisms function correctly by testing systems under failure conditions.
  • Ensure Service Level Indicators and Service Level Objectives are accurately measured and tracked through automated testing frameworks.
  • Test, maintain, patch, STIG, upgrade, troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager, Active Directory Federation Services, DHCP, DNS, WINS, Group Policy Objects, and PKI.
  • Coordinate with project managers, customers, stakeholders, and engineers to support ongoing activities and new projects that maintain, transform, and modernize the Navy Enterprise Network.
  • Provide technical leadership and related task knowledge.

Skills

Agile methodologies
Linux/Unix
Automation scripting
SRE concepts
DevSecOps
Distributed team
CI/CD concepts

Education

B.S. degree in a related field
Master’s degree with less than 2 years of relevant experience

Tools

Jenkins
GitLab
Jira
Confluence
Bitbucket
Ansible
Terraform
CloudFormation
Kubernetes
Docker
OpenShift
PowerShell
Python
VB.NET

Job description

Career Opportunities with Amatriot Group, LLC

A great place to work.

Share with friends or Subscribe!

Are you ready for new challenges and new opportunities?

Join our team!

Current job opportunities are posted here as they become available.

Subscribe to our RSS feeds to receive instant updates as new positions become available.

Clearance: Secret Clearance
Location: Norfolk, VA
Job Type: Full-Time (on-call support as needed)
Target Salary Range*:$145,000 - 165,000
*This represents the potential salary range for this position depending on education level, years of experience and/or certifications in addition to other position specific requirements which may impact salary

Position Overview

The Site Reliability Engineer develops and executes tests focused on system resilience, performance under load, and failure scenarios. This role works with Site Reliability Engineers, development teams, and operations teams to create automated testing frameworks that simulate real-world conditions, validate system behavior under normal and stress conditions, and ensure services are resilient and meet established service level objectives.

The Site Reliability Engineer supports, migrates, automates, and optimizes software development and deployment processes, Infrastructure as Code, and identity and access management capabilities. This role contributes to the maturity of the Site Reliability Engineering program and supports the creation, maintenance, modernization, and refresh of capabilities and components for the Navy Enterprise Network.

Key Responsibilities
SRE Testing and Reliability Engineering
  • Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
  • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic.
  • Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components.
  • Ensure testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with Site Reliability Engineers, developers, and operations teams to define reliability goals and develop testing strategies to validate those goals.
  • Ensure new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
  • Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
  • Ensure Service Level Indicators and Service Level Objectives are accurately measured and tracked through automated testing frameworks.
Platform Administration and Support
  • Test, maintain, patch, STIG, upgrade, troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager, Active Directory Federation Services, DHCP, DNS, WINS, Group Policy Objects, and PKI.
  • Support the creation, maintenance, update, modernization, and refresh of capabilities and components of the Navy Enterprise Network.
  • Provide technical leadership and related task knowledge.
  • Coordinate with project managers, customers, stakeholders, and engineers to support ongoing activities and new projects that maintain, transform, and modernize the Navy Enterprise Network.
Automation, CI/CD, and Infrastructure as Code
  • Write code to automate software releases, monitor systems, and detect and resolve problems before users are impacted.
  • Develop features using AI coding tools and script repositories to automate, scale, test, and secure cloud infrastructure and pipelines.
  • Develop and code high-quality pipeline automation workflows to support environments inside and outside the cloud platform.
  • Ensure automation workflows align with business and technology strategies.
  • Contribute to ongoing SRE maturity by recommending improvements to engineering build, maintenance, automation, and reliability across the platform using SRE/DevOps tools and Infrastructure as Code.
Monitoring, Performance, and Platform Optimization
  • Work with development and operations teams to support fast and reliable software deployments.
  • Monitor systems and improve overall platform reliability.
  • Discover, document, and resolve system bugs.
  • Enhance performance monitoring of systems using Splunk or other dashboard reporting tools.
  • Identify performance bottlenecks and optimize cloud infrastructure performance.
  • Maintain complex computer systems through automation, monitoring, and proactive issue resolution.
Operational Support and Escalation
  • Resolve most conflicts between timeline, budget, and scope independently.
  • Escalate sophisticated or consequential issues to senior management.
  • Work nights, weekends, and provide on-call support as needed.
Qualifications
Education
  • B.S. degree and 2–4 years of prior relevant experience; or
  • Master’s degree with less than 2 years of relevant experience.
Experience
  • Experience designing, configuring, and managing Active Directory, Group Policy Objects, DNS, DHCP, and WINS.
  • Experience with Windows Server 2016, 2019, 2022, or 2025.
  • Experience with SQL Server 2019, 2022, or 2025.
  • Experience with CI/CD toolsets, such as Jenkins or GitLab.
  • Experience in application administration, configuration, and integration.
  • Experience creating Jira and/or Azure DevOps workflows, projects, and custom configurations.
  • Experience administering and maintaining an SRE platform using Ansible playbooks, such as upgrading Jenkins.
  • Experience automating tasks with scripting languages such as PowerShell or Python.
  • Experience integrating and maintaining third-party CI/CD tools such as Jenkins and GitLab.
  • Experience with PaaS using Red Hat OpenShift, Kubernetes, and Docker containers.
  • Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
  • Experience with automated provisioning and configuration tools such as Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
  • Working knowledge of the Risk Management Framework and DISA STIGs.
Skills
  • Familiarity with Agile development methodologies.
  • Good command of Linux/Unix and command-line concepts.
  • Automated script design, coding, debugging, and maintenance skills using Bash, Python, or similar languages.
  • Knowledge of Agile, DevSecOps, and SRE concepts and best practices, with a desire to grow that knowledge.
  • Hands-on experience with Atlassian products, including Jira, Confluence, and Bitbucket.
  • Ability to work with a distributed team.
  • Ability to work in a highly collaborative, forward-thinking, and innovation-driven environment.
Certifications
  • DoD 8570.01 IAT Level II certification required prior to onboarding and must be maintained while supporting the SMIT Contract. [Required]
Clearance
  • Active DoD Secret security clearance, with ability to maintain the clearance. [Required]
Other Requirements
  • Must be able to support program execution in classified environments and access SIPRNet from an NMCI location. [Required]
  • 100% onsite work is required. [Required]
  • Must be willing to work nights, weekends, and provide on-call support as needed. [Required]
Preferred Qualifications
  • Previous work experience providing support to the NGEN-NMCI program.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation for automating test environments.
  • ITILv4, Scrum Master, or Agile SAFe certification, or applicable experience.
  • Familiarity with designing, configuring, and managing FIM 2010 R2/MIM 2016 synchronization.
  • Familiarity with VB.NET.
  • Familiarity with designing, configuring, and managing Delinea.
  • Familiarity with Microsoft Active Directory Federation Services structure.
  • Familiarity with cloud engineering.
  • Familiarity with Public Key Infrastructure certificates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

The Amatriot Group • Norfolk (VA)

On-site
USD 145,000 - 165,000
Systems Engineer – DevSecOps/SRE (w/ active Secret)
Systems Engineer – DevSecOps/SRE (w/ active Secret)

CriticalSolutions, LLC • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical coverage
Dental coverage
401K matching
+1
Systems Engineer – DevSecOps/SRE (w/ active Secret)
Systems Engineer – DevSecOps/SRE (w/ active Secret)

Critical Solutions • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical coverage
Dental coverage
Vision coverage
+4
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

Amatriot Group, LLC • Norfolk (VA)

On-site
USD 145,000 - 165,000
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

Amatriot Group, LLC • San Diego (CA)

On-site
USD 145,000 - 165,000
Site Reliability Engineer (Night Shift)
Site Reliability Engineer (Night Shift)

Amatriot Group, LLC • Honolulu (HI)

On-site
USD 145,000 - 165,000
Senior Network Engineer
Senior Network Engineer

Amatriot Group, LLC • Honolulu (HI)

On-site
USD 150,000 - 170,000
Technical Project Analyst
Technical Project Analyst

Amatriot Group, LLC • Arlington (VA)

On-site
USD 130,000 - 160,000
IT Support Analyst
IT Support Analyst

Amatriot Group, LLC • Maryland

On-site
USD 65,000 - 75,000
Information Assurance Engineer
Information Assurance Engineer

Agile IT Synergy, LLC • Tampa (FL)

On-site
USD 90,000 - 130,000