AI Service Delivery Manager

DSTA

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

DSTA is seeking an AI Service Delivery Manager to own a portfolio of live AI deployments running in on-premise and private cloud environments. You will ensure trusted, available and effective AI systems, taking accountability for service health, adoption, and escalation when needed.

Collaborate with AI product teams, infrastructure and engineering to bring new environments into service, coordinate upgrades, and drive post-incident reviews.

Qualifications

  • Hands-on experience supporting software deployed on infrastructure outside your direct control.
  • Working knowledge of Linux, containers, networking fundamentals, and at least one private cloud or virtualization platform.
  • Familiarity with Kubernetes and MLSecOps is advantageous.

Responsibilities

  • Own end-to-end service delivery for on-premise and private cloud AI deployments, meeting service levels for availability, latency, throughput and support.
  • Coordinate upgrades, patches, and configuration changes across environments and manage release processes.
  • Act as incident manager for major incidents, drive RCA, publish post-incident reviews with corrective actions.
  • Monitor production performance, identify degradation early, ensure auditability and traceability for changes.
  • Collaborate with AI products, infrastructure and engineering teams to stabilise new environments and transfer to service.
  • Drive continual service improvement and maintain runbooks for restricted environments.

Education

Degree in Computer Science, Computer Engineering and AI/ML

Tools

Prometheus
Grafana
OpenTelemetry

Job description

We are seeking an AI Service Delivery Manager to own a portfolio of live AI deployments running in on-premise and private cloud environments. We are looking for someone who believes that keeping AI systems trusted, available and effective in operational use is engineering work in its own right. As the single accountable owner for service health and adoption, you will hold service levels, drive incidents through to resolution, and coordinate upgrades across live operational environments. Working with infrastructure, engineering and user teams, you will ensure AI systems continue to perform reliably in highly controlled environments where internet connectivity is often unavailable, and change must be carefully managed.

What You'll Do
  • Service Ownership: Own end-to-end service delivery for a portfolio of on-premise and private cloud AI deployments, meeting agreed service levels for availability, inference latency, throughput, and support responsiveness. Maintain an accurate configuration baseline for every environment, and lead regular service reporting and reviews through to senior stakeholders.
  • Deployment & Transition to Service: Work with AI products, infrastructure and engineering teams to bring new environments into supported service, from site readiness through cutover, stabilisation, and formal handover. Own the service acceptance process and drive the site-side prerequisites that determine the success of on-premise deployments, including GPU and server availability, power and cooling, network rules, directory integration, and security sign-off.
  • Incident & Problem Management: Act as incident manager for major incidents, coordinating engineering, infrastructure, security, and user teams through to resolution. Drive root cause analysis, publish post-incident reviews with tracked corrective actions, and convert recurring failure patterns into permanent fixes or product improvements.
  • Change & Release Management: Plan and coordinate AI product and platform upgrades, model version rollouts, patching, and configuration changes across environments with differing maintenance windows and approval forums. Manage the end-to-end release process for on-premise deployments, including offline and air-gapped update bundles, registry mirrors, signed artefacts, alongside hardware dependencies such as GPU drivers, firmware, and version compatibility.
  • AI Service Assurance: Monitor model and application performance in production, including latency, throughput, evaluation and regression results, output quality issues, and drift indicators. Identify and elevate service degradation before it impacts users. Ensure evaluation suites are executed before every model or prompt change reaches production, with all changes documented, versioned, and traceable for accreditation and audit.
  • Adoption & Continuous Improvement: Monitor adoption and utilisation across the portfolio, identify deployments that stall after go-live, and determine whether the underlying issue relate to training, documentation, capacity or reliability. Own the continual service improvement plan, maintain operational runbooks for restricted environments, and feed structured operational insights back into engineering teams.
We Are Looking For Someone Who
  • Is organised and meticulous, able to own multiple production environments without sight of the details
  • Takes initiative and is comfortable building processes where none yet exist
  • Remains calm under pressure and becoming more methodical, not less, during major incidents
  • Enjoys collaborating across engineering, security and user organisations
  • Communicates clearly with both engineers and senior stakeholders, building trusted working relationships
  • Adapts quickly to ambiguity in a fast-moving environment
  • Takes pride in delivering reliable services that others depend on
  • Degree in Computer Science, Computer Engineering and AI/ML, or a related discipline.
  • Operationally-grounded problem solvers, from backgrounds including but not limited to service delivery, technical account management, implementation, technical programme management, senior technical support, or other operationally focused engineering roles.
  • Hands-on experience supporting software deployed on infrastructure outside your direct control, with working knowledge of Linux, containers, networking fundamentals, and at least one private cloud or virtualisation platform.
  • Able to remain clam under pressure, communicate confidently with senior stakeholders, and drive complex incidents through to resolution.
  • Resourceful self-starters comfortable with ambiguity and enjoy building processes where none exists.
  • Strong collaborator who values clear documentation, honest reporting, and effective communication across technical and operational teams.
  • Curious about how AI systems behave in production, with a willingness to continuously learn and influence outcomes across teams without direct authority.
  • Familiarity with Kubernetes, GPU infrastructure, MLSecOps, evaluation tooling, observability platforms (Prometheus, Grafana, OpenTelemetry), infrastructure-as-code, or ITIL v4 will be advantageous.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI DevOps Engineer
AI DevOps Engineer

GTS Consulting • Singapore

On-site
SGD 100,000 - 180,000
AI Engineer
AI Engineer

UNIZEN TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 180,000 - 320,000
Senior IT Infrastructure Engineer (Cloud, AI) - #1586
Senior IT Infrastructure Engineer (Cloud, AI) - #1586

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 90,000 - 130,000
AI Services Delivery Lead - On-Prem & Private Cloud
AI Services Delivery Lead - On-Prem & Private Cloud

DSTA • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Engineer - Banking
Senior AI Engineer - Banking

UNISON Group • Singapore

On-site
SGD 150,000 - 190,000
Head of AI Engineering & Platforms
Head of AI Engineering & Platforms

Gravitas Recruitment Group (Global) Ltd • Singapore

On-site
SGD 300,000 - 460,000
Site Reliability Engineer
Site Reliability Engineer

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Engineer
Senior AI Engineer

Zurich Services (Hong Kong) Limited • Singapore

On-site
SGD 120,000 - 180,000
AI Software Engineer
AI Software Engineer

CODEX SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 60,000 - 90,000
AI Solutions Specialist
AI Solutions Specialist

Chan Brothers Travel Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000