Employment Opportunity - AI Platform Operations Engineer

ITCAN PTE. LIMITED

Singapore

On-site

SGD 60,000 - 100,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ITCAN PTE. LIMITED is seeking a cloud operations professional to manage Azure AI platform availability, security, and continuity. You will collaborate with MLOps and engineering teams to integrate automation and observability into AI/ML workflows, driving continuous improvements through IaC and scripting.

The role emphasizes incident response, disaster recovery planning, and compliance support, with 1–2 years of cloud operations experience and a Bachelor’s degree in CS or engineering.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 1-2 years of experience in cloud administration and/or operations
  • Expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, and Application Insights

Responsibilities

  • Monitor Azure AI cloud platform availability and detect outages to minimize downtime and operational risks
  • Respond to incidents and perform root cause analysis to support rapid recovery and business continuity
  • Implement disaster recovery strategies to safeguard critical AI workloads
  • Support security audits and compliance reporting to align with policies, regulatory frameworks, and industry best practices
  • Collaborate with MLOps, LLMOps, and engineering teams to integrate automation, observability, and security best practices into platform operations and AI/ML workflows
  • Drive continuous improvement initiatives in platform operations through automation and enhanced observability
  • Use infrastructure-as-code tools (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python) to maintain and optimize cloud infrastructure
  • Apply knowledge of AI/ML infrastructure components such as AKS, GPU VMs, data pipelines, and model hosting to meet operational demands
  • Demonstrate problem-solving and communication skills to lead effectively during high-pressure incident scenarios
  • Anticipate potential failure scenarios and develop effective response plans to mitigate risks

Skills

Azure cloud operations
Incident response
Automation & observability
PowerShell
Python
Terraform
AKS
GPU VMs
Data pipelines
Communication

Education

Bachelors in CS/Engineering

Tools

Terraform
Bicep
ARM
PowerShell
Python
Azure Monitor
Log Analytics
Application Insights

Job description

Job Summary

You will operate and optimize Azure AI cloud platform operations, ensuring high availability, security, and business continuity. Collaborate with cross-functional teams to integrate automation and observability, driving continuous operational excellence.

Responsibilities
  • Monitor Azure AI cloud platform availability and detect outages to minimize downtime and operational risks
  • Respond to incidents and perform root cause analysis to support rapid recovery and business continuity
  • Implement disaster recovery strategies to safeguard critical AI workloads
  • Support security audits and compliance reporting to align with policies, regulatory frameworks, and industry best practices
  • Collaborate with MLOps, LLMOps, and engineering teams to integrate automation, observability, and security best practices into platform operations and AI/ML workflows
  • Drive continuous improvement initiatives in platform operations through automation and enhanced observability
  • Use infrastructure-as-code tools (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python) to maintain and optimize cloud infrastructure
  • Apply knowledge of AI/ML infrastructure components such as AKS, GPU VMs, data pipelines, and model hosting to meet operational demands
  • Demonstrate problem-solving and communication skills to lead effectively during high-pressure incident scenarios
  • Anticipate potential failure scenarios and develop effective response plans to mitigate risks
Preferred competencies and qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 1-2 years of experience in cloud administration and/or operations
  • Expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, and Application Insights
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Operations Engineer
AI Platform Operations Engineer

Krisvconsulting Services Pte Ltd • Singapore

On-site
SGD 50,000 - 80,000
Azure AI Platform Reliability Engineer
Azure AI Platform Reliability Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 60,000 - 100,000
AI Platform Operations Engineer- #AIDA (Singapore, Singapore)
AI Platform Operations Engineer- #AIDA (Singapore, Singapore)

wimatec MATTES GmbH • Singapore

On-site
SGD 60,000 - 80,000
Deployment & Operations Engineer
Deployment & Operations Engineer

GTS Consulting • Singapore

On-site
SGD 90,000 - 150,000
DevOps Engineer (AI Platform)DevOps Engineer
DevOps Engineer (AI Platform)DevOps Engineer

GOLDTECH RESOURCES PTE LTD • Singapore

On-site
SGD 100,000 - 160,000
Devops Engineer
Devops Engineer

TOSS-EX PR PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (Azure AI Solutions)
AI Engineer (Azure AI Solutions)

Unison Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Operations Engineer- #AIDA
Senior AI Platform Operations Engineer- #AIDA

Singtel • Singapore

On-site
SGD 70,000 - 90,000
Senior AI Platform Operations Engineer — Azure & Automation
Senior AI Platform Operations Engineer — Azure & Automation

Singtel Group • Singapore

On-site
SGD 80,000 - 100,000
AI DevOps Engineer
AI DevOps Engineer

GTS Consulting • Singapore

On-site
SGD 100,000 - 180,000