Site Reliability Engineer II/Infrastructure/Automation

Astreya

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading IT infrastructure firm is seeking an experienced Infra Automation Engineer to enhance the reliability and performance of critical physical security infrastructure. This role focuses on combining software engineering with operations to automate processes supporting server, network, and security device management. Ideal candidates should have 6+ years in Infra Automation Engineering, strong scripting skills in Python and Bash, and hands-on experience with Infrastructure-as-Code tools like Terraform and Ansible. The role involves 24x5 availability with on-call rotations and collaboration with cross-functional teams.

Qualifications

  • 6+ years of experience in Infra Automation Engineering or Infrastructure Engineering.
  • Strong proficiency in Python, Bash, and PowerShell for automation scripting.
  • Hands-on experience with Infrastructure-as-Code tools like Terraform or Ansible.
  • Advanced knowledge of Linux and Windows server environments.
  • Solid understanding of enterprise networking concepts and Cisco device administration.

Responsibilities

  • Establish, monitor, and enforce SLIs and SLOs for infrastructure tooling.
  • Provide Level 3 expertise for tooling-specific incidents.
  • Identify and automate repetitive manual tasks across managed infrastructure.
  • Conduct thorough root cause analysis and lead blameless postmortems.
  • Engineer and maintain automation processes for asset management platforms.

Skills

Infra Automation Engineering
Python
Bash
PowerShell
Go
Infrastructure-as-Code
Linux
Windows Server
Networking
Monitoring Solutions

Education

6+ years of experience in Infra Automation Engineering

Tools

Terraform
Ansible
Chef
Puppet
Prometheus
Grafana
Datadog
ELK Stack

Job description

We are seeking an experienced Infra Automation Engineer to join our IT Infrastructure Support team,

responsible for ensuring the reliability, scalability, and performance of critical physical security

infrastructure and supporting systems. In this role, you will combine software engineering expertise with

operations knowledge to build and maintain automation tools, monitoring systems, and processes that

support enterprise-grade server, network, and security device management. You will work closely with

cross-functional teams to define and enforce service level objectives, reduce operational toil through

automation, and drive continuous improvement in system resilience. This position requires 24x5

availability with on-call rotation to ensure uninterrupted support for mission-critical infrastructure.

Key Responsibilities
  • Partner with leadership to establish, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tooling, including configuration compliance rates, patch success rates, and deployment latency metrics.
  • Provide Level 3 expertise for tooling-specific incidents, focusing on automating incident remediation workflows and reducing Mean Time To Repair (MTTR) through intelligent automation and runbook development.
  • Identify and automate repetitive manual tasks across managed infrastructure, targeting measurable reductions in operational overhead (e.g., 50% reduction in manual server build time) through scripting and workflow automation.
  • Conduct thorough root cause analysis and lead blameless postmortems for all major service-impacting incidents, driving systemic improvements in tooling reliability and infrastructure resilience.
  • Engineer and maintain automated processes and scripts to populate, update, and synchronize asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems for internal and external stakeholders.
  • Design, develop, and deploy full-stack applications, custom plugins, and automation scripts to extend functionality of management and monitoring systems, enabling direct device interaction for configuration management.
  • Develop and maintain fully automated Infrastructure-as-Code configurations for Windows and Linux server roles using tools such as Ansible, Terraform, or Puppet, including drift detection and
  • Build end-to-end automation pipelines for vulnerability patching, security baseline enforcement (CIS benchmarks), and continuous compliance auditing against internal and regulatory standards for physical security devices.
  • Develop API-driven tools for network configuration management, automated firmware updates, pre/post-change validation, and real-time network health monitoring across the device fleet.
  • Deploy and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error rate, saturation, traffic) for servers and
  • Build automation scripts for intelligent ticket handling, problem validation, and escalation workflows within enterprise ticketing systems, ensuring 2-hour initial response SLAs are consistently met.
  • Participate in 24x5 on-call rotation to provide timely support for infrastructure systems, security devices, and related tooling, ensuring service continuity and rapid incident response.
Required Skills
  • 6+ years of experience in Infra Automation Engineering, or Infrastructure Engineering
  • Strong proficiency in Python, Bash, and PowerShell for automation scripting, with experience in Go for building high-performance backend services and APIs.
  • Hands-on experience with Infrastructure-as-Code tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including drift detection, version control, and automated remediation.
  • Advanced knowledge of Linux and Windows server environments, including Tier 3 troubleshooting capabilities, system hardening, and enterprise-scale server management.
  • Solid understanding of enterprise networking concepts, Cisco device administration, network automation protocols (NETCONF/RESTCONF), and experience with network monitoring and flow analysis tools.
  • Experience implementing and managing monitoring solutions (Prometheus, Grafana, Datadog) and centralized logging platforms (ELK Stack), with ability to create custom dashboards and
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Automation Engineer
Infrastructure Automation Engineer

Meriton • Irving (TX), Phoenix (AZ), Tampa (FL)

On-site
USD 110,000 - 165,000
Senior Infrastructure Automation Engineer
Senior Infrastructure Automation Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 100,000 - 130,000
Infrastructure Engineer
Infrastructure Engineer

Murtech Staffing & Solutions • Pittsburgh

On-site
USD 90,000 - 120,000
Automation Engineer
Automation Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 90,000 - 120,000
IT Engineer - Linux and Ansible Automation
IT Engineer - Linux and Ansible Automation

TechWish • Walnut Creek (CA)

On-site
USD 120,000 - 180,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Insight Global • Billerica (MA)

On-site
USD 100,000 - 130,000
Systems Engineer
Systems Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 80,000 - 110,000
Infrastructure Engineer
Infrastructure Engineer

ISA Consulting Group • Tampa (FL)

On-site
USD 90,000 - 130,000
Senior Infrastructure Automation Engineer
Senior Infrastructure Automation Engineer

Prestige Staffing • Atlanta (GA)

On-site
USD 115,000 - 125,000
Medical insurance
Vision insurance
401(k)
Associate II Infrastructure Engineer
Associate II Infrastructure Engineer

ISA Consulting Group • Tampa (FL)

On-site
USD 85,000 - 110,000