We're looking for an experienced Senior Site Reliability / Infrastructure Engineer to operate, maintain, and improve production infrastructure across complex and highly governed environments.
This role combines Site Reliability Engineering, Infrastructure as Code, Kubernetes, CI/CD, Security, and Observability, working closely with engineering, security, operations, and global technology teams to ensure reliable, secure, and resilient production systems.
- Operate and maintain production infrastructure with a strong focus on reliability, availability, and performance.
- Define and apply SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems.
- Design and manage infrastructure using Infrastructure as Code, treating environments as declarative configurations rather than individual machines.
- Operate and troubleshoot Kubernetes workloads, including scheduling, storage, networking, and deployment issues.
- Build and maintain CI/CD pipelines, with a preference for Azure DevOps.
- Manage secrets, credentials, and certificates using platforms such as HashiCorp Vault, CyberArk, or equivalent.
- Support PKI operations including certificate issuance, renewal, credential rotation, and recovery of administrative access.
- Implement and maintain observability solutions across metrics, logs, traces, alerting, and system correlation.
- Automate infrastructure and operational processes using Python and/or PowerShell.
- Collaborate with engineering, security, and operations teams across global environments to improve system reliability and operational maturity.
- 5+ years of experience operating production systems, or 3+ years with demonstrable end-to-end ownership of production infrastructure.
- Hands-on experience with Site Reliability Engineering practices, including SLIs, SLOs, error budgets, incident management, and postmortems.
- Strong experience with Infrastructure as Code, using Terraform, Ansible, or equivalent technologies.
- Hands-on experience operating Kubernetes workloads in production.
- Practical experience with secrets management platforms such as HashiCorp Vault, CyberArk, or equivalent.
- Experience with PKI, certificate management, credential rotation, and recovery procedures.
- Experience building and maintaining CI/CD pipelines, preferably using Azure DevOps.
- Strong scripting skills in Python and/or PowerShell.
- Solid understanding of observability fundamentals, including metrics, logs, traces, alerting, and correlation.
- Experience working in complex, highly governed, or regulated environments, preferably OT/PCN.
- Strong communication skills and ability to collaborate effectively across global teams.
- Advanced English (mandatory).
Nice to Have
- Experience working in air-gapped or egress-restricted environments.
- Experience with governed artifact distribution and repositories such as Pulp, Artifactory, Nexus, or Harbor.
- Background in Industrial, OT, Oil & Gas, Utilities, Manufacturing, or similar environments.
- Experience with ServiceNow Service Mapping, CMDB, or automation integrations.
- Knowledge of security and compliance frameworks such as CIS, NIST, or ISA/IEC 62443.
- Familiarity with SolarWinds or similar observability and compliance platforms.
- Experience working with Agile, backlog-driven development environments.
What We Offer
- Full time employment.
- 100% remote.
- Competitive salary in ARS.
- Opportunity to work with complex production infrastructure and global technology teams.