Senior AI Platform Operations Engineer

itcan pte. limited

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

itcan pte. limited seeks an experienced Cloud Platform SRE to ensure availability and performance of our Azure AI cloud platform built on Red Hat OpenShift.

You will monitor, detect outages, optimize performance, and drive automation and resilience across the service. Responsibilities include incident response, root cause analysis, disaster recovery design, and overseeing security operations including IAM, SIEM/SOAR.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • 4–6 years of experience in cloud administration and/or operations.
  • Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights.
  • Hands-on experience in observability and performance tuning for OpenShift clusters, leveraging built-in monitoring and logging tools.
  • Strong background in incident management, SRE practices, and disaster recovery design.
  • Hands-on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection.
  • Proficiency in infrastructure-as-code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python).
  • Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands.
  • Knowledge of security compliance frameworks (ISO 27001, CIS, NIST).
  • Excellent problem-solving, communication, and leadership skills, especially in high‑pressure incident scenarios.

Responsibilities

  • Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform.
  • Ensures continuous availability, resilience, and efficiency of the OpenShift–powered RE:AI cloud platform.
  • Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity.
  • Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement.
  • Support security audits, compliance reporting, and ensure alignment with policies, regulatory frameworks and industry best practices.
  • Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows.
  • Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives.

Skills

Azure operations
OpenShift
Incident management
Disaster recovery
Automation scripting
Terraform
Python
PowerShell
Security monitoring
Observability

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

Azure Monitor
Log Analytics
Application Insights
Red Hat OpenShift
Terraform
Bicep
ARM
PowerShell
Python

Job description

Make an Impact by:

Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform

Ensures continuous availability, resilience, and efficiency of the Red Hat OpenShift–powered RE:AI cloud platform.

Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity

Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement

Support security audits, compliance reporting, and ensure alignment withpolicies, regulatory frameworks and industry best practices

Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows

Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives

Skills for Success:

Bachelor’s degree in Computer Science, Engineering, or a related field

4-6 years of experience in cloud administration and/or operations

Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights

Hands‑on experience in observability and performance tuning for Red Hat OpenShift clusters, leveraging built‑in monitoring and logging tools.

Strong background in incident management, SRE practices, and disaster recovery design

Hands‑on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection

Proficiency in infrastructure-as-code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python)

Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands

Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)

Excellent problem-solving, communication, and leadership skills, especially in high‑pressure incident scenarios

Forward thinking ability to identify possible failure scenarios and formulate effective response plans

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Operations Engineer- #AIDA
Senior AI Platform Operations Engineer- #AIDA

Singtel • Singapore

On-site
SGD 70,000 - 90,000
AI Platform Engineer (Network, Operations & Security)
AI Platform Engineer (Network, Operations & Security)

PEOPLESEARCH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Employment Opportunity - AI Platform Operations Engineer
Employment Opportunity - AI Platform Operations Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 60,000 - 100,000
Senior Azure AI Platform SRE & Resilience Engineer
Senior Azure AI Platform SRE & Resilience Engineer

itcan pte. limited • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Reliability Engineer
Senior AI Platform Reliability Engineer

Singapore Telecommunications Limited • Singapore

On-site
SGD 90,000 - 130,000
Senior AI Platform Operations Engineer
Senior AI Platform Operations Engineer

Singapore Telecommunications Limited • Singapore

On-site
SGD 90,000 - 130,000
AI Platform Engineer #AIDA
AI Platform Engineer #AIDA

Singtel • Singapore

On-site
SGD 75,000 - 95,000
Lead AI Platform Network & Security Engineer
Lead AI Platform Network & Security Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 180,000 - 240,000
Deployment & Operations Engineer
Deployment & Operations Engineer

GTS Consulting • Singapore

On-site
SGD 90,000 - 150,000
AI DevOps Engineer
AI DevOps Engineer

GTS Consulting • Singapore

On-site
SGD 100,000 - 180,000