IT Infrastructure Support Site Reliability Engineer II

astreya

Ireland

On-site

EUR 80,000 - 110,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

astreya is seeking an experienced Infra Automation Engineer to join our IT Infrastructure Support team. You will ensure reliability and performance of critical security infrastructure, including IP camera fleets, access control, Cisco switch fleet, servers, and cloud environments.

The role blends software engineering with operations to build automation, CMDB, monitoring, and tooling. You will implement IaC with Terraform/Ansible, automate remediation, and support 24x5 on-call, driving resilience

Qualifications

  • 6+ years of experience in Infra Automation Engineering, or Infrastructure Engineering.
  • Strong proficiency in Python, Bash, and PowerShell; Go for backend services.
  • Hands-on experience with IaC tools Terraform, Ansible, Chef, or Puppet, with drift detection and remediation.
  • Advanced knowledge of Linux and Windows server environments; Tier 3 troubleshooting.
  • Understanding of enterprise networking, Cisco device administration, NETCONF/RESTCONF.

Responsibilities

  • Establish and monitor SLIs/SLOs for infrastructure tooling and deployment latency.
  • Provide Level 3 expertise for tooling incidents and automate remediation workflows.
  • Automate repetitive tasks across managed infrastructure to reduce manual effort.
  • Perform root cause analyses and blameless postmortems to improve tooling reliability.
  • Develop automated processes to manage NetBox, CMDB, and monitoring data.
  • Design and deploy full-stack apps, plugins, and automation scripts for device interaction.
  • Create IaC configurations for Windows and Linux servers with drift detection.
  • Build end-to-end automation for patching, CIS baselines, and compliance.
  • Develop API-driven tools for network config, firmware updates, zero-touch provisioning, and validation.
  • Deploy monitoring agents, log collection, and dashboards with SLIs-based alerts.
  • Create monitoring exporters for device fleet and clock drift diagnostics.
  • Develop automation scripts for intelligent ticket handling and 2-hour SLAs.
  • Support security improvements like credential controls and automated backups.
  • Participate in 24x5 on-call rotation for infrastructure and security devices.

Skills

Python
Bash
PowerShell
Go
Terraform
Ansible
Chef
Puppet
Linux
Windows Server
Cisco device administration
NETCONF/RESTCONF
Network monitoring
NetBox

Tools

NetBox

Job description

About the Job

We are seeking an experienced Infra Automation Engineer to join our IT Infrastructure Support team, responsible for ensuring the reliability, scalability, and performance of critical physical security infrastructure, including IP camera fleets, access control systems, and a large-scale Cisco switch fleet, and the servers, network, and cloud environment that support them. In this role, you will combine software engineering expertise with operations knowledge to build and maintain automation tools, a centralized CMDB, monitoring systems, and processes that support enterprise-grade server, network, and security device management within a large-scale, cloud-hosted enterprise environment. You will work closely with cross-functional teams to define and enforce service level objectives, reduce operational toil through automation, and drive continuous improvement in system resilience. This position requires 24x5 availability with on-call rotation to ensure uninterrupted support for mission-critical infrastructure.

Key Responsibilities
  • Partner with leadership to establish, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tooling, including configuration compliance rates, patch success rates, and deployment latency metrics.
  • Provide Level 3 expertise for tooling-specific incidents, focusing on automating incident remediation workflows and reducing Mean Time To Repair (MTTR) through intelligent automation and runbook development.
  • Identify and automate repetitive manual tasks across managed infrastructure, targeting measurable reductions in operational overhead (e.g., 50% reduction in manual server build time) through scripting and workflow automation.
  • Conduct thorough root cause analysis and lead blameless postmortems for all major service-impacting incidents, driving systemic improvements in tooling reliability and infrastructure resilience.
  • Engineer and maintain automated processes and scripts to populate, update, and synchronize asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems for internal and external stakeholders.
  • Design, develop, and deploy full-stack applications, custom plugins, and automation scripts to extend functionality of management and monitoring systems, enabling direct device interaction for configuration management.
  • Develop and maintain fully automated Infrastructure-as-Code configurations for Windows and Linux server roles using tools such as Ansible, Terraform, or Puppet, including drift detection and auto-remediation capabilities.
  • Build end-to-end automation pipelines for vulnerability patching, security baseline enforcement (CIS benchmarks), and continuous compliance auditing against internal and regulatory standards for physical security devices.
  • Develop API-driven tools for network configuration management, automated firmware updates, zero-touch provisioning, pre/post-change validation, and real-time network health monitoring across the device fleet.
  • Deploy and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error rate, saturation, traffic) for servers and edge devices.
  • Build custom monitoring exporters for the physical security device fleet, including camera systems, ensuring accurate metrics and structured, multi-severity logging output.
  • Build diagnostic tooling to correlate timestamps across distributed log streams and detect clock drift or NTP desync, a recurring root cause of false-positive outages across the device fleet.
  • Build automation scripts for intelligent ticket handling, problem validation, and escalation workflows within enterprise ticketing systems, ensuring 2-hour initial response SLAs are consistently met.
  • Support foundational security improvements across the device fleet, including managed credential/access controls and automated configuration backup.
  • Participate in 24x5 on-call rotation to provide timely support for infrastructure systems, security devices, and related tooling, ensuring service continuity and rapid incident response.
Required Skills
  • 6+ years of experience in Infra Automation Engineering, or Infrastructure Engineering.
  • Strong proficiency in Python, Bash, and PowerShell for automation scripting, with experience in Go for building high-performance backend services and APIs.
  • Hands-on experience with Infrastructure-as-Code tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including drift detection, version control, and automated remediation.
  • Advanced knowledge of Linux and Windows server environments, including Tier 3 troubleshooting capabilities, system hardening, and enterprise-scale server management.
  • Solid understanding of enterprise networking concepts, Cisco device administration, network automation protocols (NETCONF/RESTCONF), and experience with network monitoring and flow analysis tools.
  • Experience i
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

IT Infrastructure Support Site Reliability Engineer II
IT Infrastructure Support Site Reliability Engineer II

Astreya • Cork

On-site
EUR 50,000 - 64,000
IT Infrastructure Support Site Reliability Engineer II
IT Infrastructure Support Site Reliability Engineer II

Astreya • Dublin

On-site
EUR 50,000 - 64,000
IT Infrastructure Support Site Reliability Engineer II
IT Infrastructure Support Site Reliability Engineer II

Astreya Inc. • Dublin

On-site
EUR 50,000 - 64,000
Site Reliability Engineer II — Infra Automation & Cisco
Site Reliability Engineer II — Infra Automation & Cisco

astreya • Ireland

On-site
EUR 80,000 - 110,000
Site Reliability Engineer III - Eng
Site Reliability Engineer III - Eng

UKG • Leinster

On-site
EUR 60,000 - 80,000
Systems Administrator
Systems Administrator

Cpl • Dublin

On-site
EUR 90,000 - 120,000
IT Infrastructure Installation Services II
IT Infrastructure Installation Services II

Astreya • Ireland

On-site
EUR 51,000 - 63,000
Technology Architect Analyst
Technology Architect Analyst

ELLIOTT MOSS CONSULTING PTE. LTD. • Dublin

On-site
EUR 60,000 - 90,000
Datacenter Technician
Datacenter Technician

Apex Systems • Dublin

On-site
EUR 42,000 - 60,000
Senior DevOps Automation Engineer
Senior DevOps Automation Engineer

Realtime Recruitment • Dublin

On-site
EUR 90,000 - 140,000