Get more replies from employers
Send a job-specific resume in minutes.
itcan pte. limited seeks an experienced Cloud Platform SRE to ensure availability and performance of our Azure AI cloud platform built on Red Hat OpenShift.
You will monitor, detect outages, optimize performance, and drive automation and resilience across the service. Responsibilities include incident response, root cause analysis, disaster recovery design, and overseeing security operations including IAM, SIEM/SOAR.
Make an Impact by:
Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform
Ensures continuous availability, resilience, and efficiency of the Red Hat OpenShift–powered RE:AI cloud platform.
Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity
Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement
Support security audits, compliance reporting, and ensure alignment withpolicies, regulatory frameworks and industry best practices
Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows
Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives
Skills for Success:
Bachelor’s degree in Computer Science, Engineering, or a related field
4-6 years of experience in cloud administration and/or operations
Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights
Hands‑on experience in observability and performance tuning for Red Hat OpenShift clusters, leveraging built‑in monitoring and logging tools.
Strong background in incident management, SRE practices, and disaster recovery design
Hands‑on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection
Proficiency in infrastructure-as-code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python)
Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands
Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)
Excellent problem-solving, communication, and leadership skills, especially in high‑pressure incident scenarios
Forward thinking ability to identify possible failure scenarios and formulate effective response plans