Production Support / Site Reliability Engineer (SRE)
Role Overview
We are looking for an experienced Production Support / Site Reliability Engineer (SRE) with experience in Agentic AI to support mission-critical, business-facing applications. This is a hands‑on, techno-functional role covering production operations, incident resolution, system reliability, and stakeholder support across modern cloud and containerised environments.
Key Responsibilities
- Provide day‑to‑day production and application support for mission‑critical, business‑facing systems.
- Investigate and resolve complex application and infrastructure issues across multiple technology layers.
- Manage incident triage, incident management, problem management, and root cause analysis.
- Monitor application health, system performance, batch processes, and scheduled workloads to maintain service reliability.
- Troubleshoot issues across Linux/Unix, cloud, containerised, and application environments.
- Develop and maintain Bash/Shell scripts to support operational activities and automation.
- Work closely with business users, engineering teams, and other stakeholders to resolve production issues and minimise service disruption.
- Identify opportunities to improve system reliability, monitoring, automation, and operational processes.
Requirements
- Degree in Computer Science, Information Technology, or a related discipline.
- At least 5 years of experience in Production/Application Support or Site Reliability Engineering (SRE).
- Strong hands‑on experience supporting business‑facing applications and users.
- Proficiency in Control‑M, Unix/Linux, Bash, and Shell scripting.
- Experience with AWS and/or Azure and cloud‑native environments.
- Hands‑on experience with Kubernetes and containerised applications.
- Familiarity with operational and infrastructure tools such as AutoSys, Datadog, and Terraform.
- Strong troubleshooting skills across application and infrastructure layers.
- Experience with incident and problem management.
- Strong analytical, communication, and stakeholder management skills.
- Proactive, collaborative, and adaptable approach to working in fast‑paced environments
- Exposure to Agentic AI technologies and AI‑enabled operational use cases.
Key Technologies
Control-M | Unix/Linux | Bash/Shell | AWS | Azure | Kubernetes | AutoSys | Datadog | Terraform | Cloud‑Native | Agentic AI