AIOps, LLMOps, MLOps (Machine Learning Operations) DevOps SRE Engineer

Avensys Consulting

Singapore

On-site

SGD 120,000 - 180,000

Full time

22 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Avensys Consulting in Singapore seeks a specialist to implement automated AI/ML pipelines, observability, and model lifecycle automation across training, deployment, and evaluation. You will drive AIOps-based reliability, trigger retraining, and guardrail enforcement for AI and LLM operations, while monitoring health, latency, drift, and GPU utilization.

You’ll collaborate with safety and security teams to enforce guardrails, version control, audit readiness, and governance in CI/CD workflows

Qualifications

  • Hands-on experience in AI/ML pipeline orchestration and model lifecycle automation.
  • Strong knowledge of MLOps, LLMOps, and AIOps for CI/CD and model retraining.
  • Experience with AI observability, monitoring, and telemetry for AI systems.

Responsibilities

  • Hands-on experience in AI/ML pipeline orchestration (Kubeflow, MLflow, Airflow, Azure ML) and model lifecycle automation.
  • Deep understanding of MLOps, LLMOps, and AIOps frameworks for CI/CD, model retraining, evaluation, and deployment at scale.
  • Proficient in observability and telemetry integration for AI systems — monitoring model health, drift, inference latency, and GPU utilization.
  • Skilled in automated incident detection, correlation, and remediation using AI-driven observability and alerting systems.
  • Knowledge of data governance, evaluation frameworks (Promptfoo, Portkey, RAG validation) and DevSecOps principles for AI pipelines.
  • Design and operationalize the AI reliability stack, covering training, inference, and evaluation pipelines across Central Kitchen shared services.
  • Implement automated health checks, anomaly detection, and retraining triggers for deployed LLM and RAG models.
  • Collaborate with AI Safety and Security teams to enforce model guardrails, version control, and change governance in CI/CD workflows.
  • Establish monitoring baselines and SLOs for AI workloads (throughput, latency, accuracy, drift) across all environments.
  • Integrate AIOps-driven diagnostics to reduce mean time to detect (MTTD) and mean time to recover (MTTR) in AI platform operations.
  • Support SRB/ARB reviews by maintaining clear operational documentation, evaluation evidence, and audit readiness of AI components

Skills

AI/ML pipelines
MLOps
CI/CD
Kubeflow
MLflow
Airflow
Azure ML

Tools

Kubeflow
MLflow
Airflow
Azure ML

Job description

Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success, we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail and supply chain,

Job Summary

Implements automated pipelines and observability for AI model training, deployment, and evaluation.

Drives AIOps-based reliability, model retraining triggers, and guardrail enforcement for AI and LLM operations.

Responsibilities | CORE
  • Hands-on experience in AI/ML pipeline orchestration (Kubeflow, MLflow, Airflow, Azure ML) and model lifecycle automation.
  • Deep understanding of MLOps, LLMOps, and AIOps frameworks for CI/CD, model retraining, evaluation, and deployment at scale.
  • Proficient in observability and telemetry integration for AI systems — monitoring model health, drift, inference latency, and GPU utilization.
  • Skilled in automated incident detection, correlation, and remediation using AI-driven observability and alerting systems.
  • Knowledge of data governance, evaluation frameworks (Promptfoo, Portkey, RAG validation) and DevSecOps principles for AI pipelines.
  • Design and operationalize the AI reliability stack, covering training, inference, and evaluation pipelines across Central Kitchen shared services.
  • Implement automated health checks, anomaly detection, and retraining triggers for deployed LLM and RAG models.
  • Collaborate with AI Safety and Security teams to enforce model guardrails, version control, and change governance in CI/CD workflows.
  • Establish monitoring baselines and SLOs for AI workloads (throughput, latency, accuracy, drift) across all environments.
  • Integrate AIOps-driven diagnostics to reduce mean time to detect (MTTD) and mean time to recover (MTTR) in AI platform operations.
  • Support SRB/ARB reviews by maintaining clear operational documentation, evaluation evidence, and audit readiness of AI components
WHAT’S ON OFFER

You will be remunerated with an excellent base salary and entitled to attractive company benefits. Additionally, you will get the opportunity to enjoy a fun and collaborative work environment, alongside a strong career progression.

Your interest will be treated with strict confidentiality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - AIOps, LLMOps, MLOps
Site Reliability Engineer - AIOps, LLMOps, MLOps

AVENSYS CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
SRE: AIOps, LLMOps & MLOps for AI Reliability
SRE: AIOps, LLMOps & MLOps for AI Reliability

AVENSYS CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Manager ML Operation
Senior Manager ML Operation

AVENSYS CONSULTING PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
Company benefits
Fun and collaborative environment
Strong career progression
Lead AI / ML Operations #AIDA
Lead AI / ML Operations #AIDA

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Lead AI / ML Operations #AIDA
Lead AI / ML Operations #AIDA

Singtel Group • Singapore

On-site
SGD 180,000 - 260,000
AI Engineer
AI Engineer

Seatrium Ltd • Singapore

On-site
SGD 90,000 - 130,000
AI Engineer
AI Engineer

Seatrium • Singapore

On-site
SGD 60,000 - 100,000
MLOps Platform Engineer: Scale AI Workloads
MLOps Platform Engineer: Scale AI Workloads

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 130,000
Platforms Engineer, MLOps
Platforms Engineer, MLOps

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 130,000
Senior Data Scientist - AI, MLOps & Cloud
Senior Data Scientist - AI, MLOps & Cloud

avensys consulting pte. ltd. • Singapore

On-site
SGD 95,000 - 130,000
Competitive salary
Company benefits