To lead the next phase of our AI evolution, we’ve launched a new business unit AIDA – Artificial Intelligence & Data Analytics – a strategic engine driving our transformation designed to scale our AI ambitions with precision and purpose. This marks a pivotal shift in how we operate, innovate, and serve to embed intelligence into every layer of our business.
Make An Impact By
- Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform
- Ensures continuous availability, resilience, and efficiency of the Red Hat OpenShift–powered RE:AI cloud platform.
- Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity
- Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement
- Support security audits, compliance reporting, and ensure alignment with Singtel policies, regulatory frameworks and industry best practices
- Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows
- Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives
Skills For Success
- Bachelor’s degree in Computer Science, Engineering, or a related field
- 4-6 years of experience in cloud administration and/or operations
- Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights
- Hands‑on experience in observability and performance tuning for Red Hat OpenShift clusters, leveraging built-in monitoring and logging tools.
- Strong background in incident management, SRE practices, and disaster recovery design
- Hands‑on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection
- Proficiency in infrastructure‑as‑code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python)
- Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands
- Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)
- Excellent problem‑solving, communication, and leadership skills, especially in high‑pressure incident scenarios
- Forward thinking ability to identify possible failure scenarios and formulate effective response plans