DevOps Engineer

NCS Group

Singapore

On-site

SGD 70,000 - 110,000

Full time

39 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career development opportunities
Collaborative culture
Continuous learning programs

Job summary

NCS Group is seeking a DevOps Engineer to join NCS AI Central's Forward Deployed Engineering team. You will own deployment pipelines for LLM/agentic apps, build repeatable CI/CD on Kubernetes, and harden production-grade deployments for POC/POV to scale.

You will instrument observability, track SLAs, and optimize token costs while collaborating with AI Engineers and Cloud Architects. The role blends rapid POC work with disciplined production operations, mentoring on LLMOps and contributing to

Qualifications

  • 2+ years in DevOps/MLOps/platform engineering, with hands-on exposure to LLM or ML systems in production; 5+ years with end-to-end ownership expected at Senior level.
  • Hands-on with containers, Kubernetes, CI/CD (GitHub Actions/GitLab/Jenkins), and Infrastructure-as-Code (Terraform).
  • Practical experience deploying and operating LLM applications (RAG/agentic systems) at scale, including model gateway/routing patterns.
  • Strong scripting/programming ability (Python and/or Go), comfortable working across cloud platforms (AWS/Azure/GCP); exposure to GCC/HCC a plus.
  • Working knowledge of observability tooling (OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch) applied to AI-specific signals.
  • Comfortable operating in fast-paced POC/POV settings and disciplined SLA-driven production support environments.
  • Working knowledge of the China AI model/tech stack (DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.

Responsibilities

  • Own the deployment pipeline for LLM and agentic applications — versioning models, prompts, and configurations as code, with safe rollout and rollback paths.
  • Build repeatable CI/CD pipelines for AI services (containerised, on Kubernetes) to ship models and prompts without manual intervention.
  • Support rapid environment spin-up for POC/POV work, then harden pipelines into production-grade deployments for scaling engagements.
  • Instrument LLM applications with observability for latency, error rates, drift, and hallucination signals; define SLAs/SLOs with AI Architects.
  • Set up alerting and runbooks for production AI systems to catch issues before clients notice.
  • Monitor token spend and inference cost; tune routing between model tiers for cost-performance at scale.
  • Stand up reusable deployment scaffolding during FDE engagements and own steady-state production operations; contribute to internal asset library.
  • Mentor engineers on LLMOps practices; participate in Production Readiness Review gate reviews.

Skills

LLMOps
Kubernetes
CI/CD
Terraform
Python/Go
Cloud platforms
Observability
Model gateways
Vector databases
China AI stack

Tools

GitHub Actions
GitLab
Jenkins
OpenTelemetry
Prometheus
Grafana
ELK/OpenSearch
Terraform
vLLM/TGI

Job description

NCS is a leading AI Tech Services company. With a 15,000-strong team across the Asia Pacific, NCS scales its platforms and capabilities to provide clients with greater agility and AI expertise across a range of Industries. Embracing a strong ecosystem of global partners, NCS transforms technology services delivery combining AI with digital resilience to drive real business impact. NCS is a subsidiary of the Singtel Group.

This will be a DevOps Engineer role that sits within NCS AI Central's (AIC) Forward Deployed Engineering (FDE) model — the combined capability that takes AI solutions from proof-of-concept through to hardened production systems. You will operate across both fast-moving FDE engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work — bringing the same rigor and a reusable, asset-fed approach to both.

What will you do?
1. Model & Deployment Operations
  • Own the deployment pipeline for LLM and agentic applications — versioning models, prompts, and configurations as code, with safe rollout and rollback paths.
  • Build repeatable CI/CD pipelines for AI services (containerised, on Kubernetes) so new model or prompt versions ship without manual intervention.
  • Support rapid, disposable environment spin-up for FDE POC/POV work, then harden the same pipeline into a production-grade deployment when an engagement scales.
2. Monitoring & Reliability
  • Instrument LLM applications with observability for latency, error rates, output drift, and hallucination signals — not just infrastructure uptime.
  • Define and track SLAs/SLOs for production AI services in partnership with AI Architects and Solution Architects.
  • Set up alerting and runbooks so issues in live AI systems are caught and triaged before clients notice.
3. Cost & Performance Management
  • Monitor token spend and inference cost per model/engagement; flag anomalies and right-sizing opportunities in partnership with AI FinOps.
  • Tune routing between model tiers (frontier vs. smaller/fine-tuned models) for cost-performance balance at production scale.
4. FDE & Development/Maintenance Coverage
  • During FDE engagements: stand up lightweight, reusable deployment scaffolding that lets AI Engineers iterate quickly on POC/POV without operational overhead.
  • During system development & maintenance engagements: take ownership of steady-state production operations, patching, upgrades, and incident response for live AI systems.
  • Contribute reusable deployment patterns back into the shared internal asset library so future engagements start from a hardened baseline, not from zero.
  • Work closely with AI Engineers, Cloud Architects, and AI Site-facing leads to ensure a smooth handoff from POC to Scale to Operate.
  • Mentor engineers on LLMOps practices and participate in PRR (Production Readiness Review) gate reviews.

We are hiring at two levels for this role. All responsibilities above apply to both; the distinction is in scope of ownership, years of experience, and seniority of judgement expected.

  • 2–4 years of relevant experience. Executes deployment pipelines, monitoring setup, and cost tracking for individual engagements, under guidance from a Senior LLMOps Engineer or AI Architect.
  • Builds and maintains CI/CD and observability for one or two engagements at a time; escalates novel production incidents to senior team members.
  • 5+ years of relevant experience, including prior ownership of production AI/ML systems end-to-end. Sets LLMOps standards and reusable deployment patterns across multiple engagements.
  • Leads incident response for critical production issues, mentors junior LLMOps and AI Engineers, and engages directly with client technical stakeholders on production-readiness and reliability.
The ideal candidate should possess:
  • 2+ years in DevOps/MLOps/platform engineering, with hands-on exposure to LLM or ML systems in production (not just prototypes); 5+ years with end-to-end ownership expected at Senior level.
  • Hands-on with containers, Kubernetes, CI/CD (GitHub Actions/GitLab/Jenkins), and Infrastructure-as-Code (Terraform).
  • Practical experience deploying and operating LLM applications (RAG/agentic systems) at scale, including model gateway/routing patterns.
  • Strong scripting/programming ability (Python and/or Go), comfortable working across cloud platforms (AWS/Azure/GCP; GCC/HCC exposure a plus).
  • Working knowledge of observability tooling (OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch) applied to AI-specific signals (drift, hallucination rate, token cost).
  • Comfortable operating in both fast-paced, ambiguous POC/POV settings and disciplined, SLA-driven production support environments.
  • Working knowledge of the China AI model/tech stack (e.g., DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.
Preferred Qualifications
  • Experience with model gateways/routers (LiteLLM, Bedrock, Vertex, Azure OpenAI) and vector databases (pgvector, Pinecone, Weaviate).
  • Exposure to regulated or government environments (IM8/VAPT, PDPA) and multi-tenant/data-residency patterns.
  • Familiarity with agent orchestration frameworks (LangGraph, Semantic Kernel) and prompt/version control tooling.
  • Experience contributing to or maintaining a reusable internal platform/asset library.
  • Hands-on deployment/self-hosting experience with Chinese open-weight models (DeepSeek, Qwen, GLM) via vLLM/TGI or similar inference runtimes.
Tech Stack (Illustrative)
  • Languages: Python, Go/TypeScript
Why Join NCS?
Grow with Us
  • Work on cutting-edge AI products that shape the future of technology
  • Collaborate with talented, passionate teams across research, engineering, and design
  • Access continuous learning opportunities and career development pathways
Make an Impact
  • Transform AI research into products that solve real problems for clients and users
  • Drive innovation in a leading Technology Services Firm with regional presence
  • Contribute to building a better future through responsible, human-centred AI
Thrive in Our Culture
  • Experience a human-to-human approach where relationships and collaboration matter
  • Be part of Team NCS, where bold ideas meet practical execution
  • Enjoy a supportive environment that values diversity, inclusion, and respect

We are driven by our AEIOU beliefs— Adventure, Excellence, Integrity, Ownership, and Unity—and we seek individuals who embody these values in both their professional and personal lives. We are committed to our Impact: Valuing our clients, Growing our people, and Creating our future.

Together, we make the extraordinary happen.

Learn more about us at ncs.co and visit our LinkedIn career site.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

#EG Senior / LLMOps Engineer
#EG Senior / LLMOps Engineer

NCS Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
Learning & development
Artificial Intelligence Specialist
Artificial Intelligence Specialist

NCS Group • Singapore

On-site
SGD 100,000 - 160,000
#EG Fullstack Developer – AI Applications
#EG Fullstack Developer – AI Applications

NCS Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
#EG Senior / AI FinOps Engineer
#EG Senior / AI FinOps Engineer

NCS Group • Singapore

On-site
SGD 120,000 - 180,000
#EG AI / LLM Specialist
#EG AI / LLM Specialist

NCS Pte Ltd • Singapore

On-site
SGD 120,000 - 190,000
#EG AI Engineer
#EG AI Engineer

NCS Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
Cost Engineer
Cost Engineer

NCS Group • Singapore

On-site
SGD 120,000 - 210,000
#EG Cloud Engineer / Architect – AI Infrastructure
#EG Cloud Engineer / Architect – AI Infrastructure

NCS Group • Singapore

On-site
SGD 120,000 - 180,000
Competitive compensation
Professional development support
#EG Data Scientist
#EG Data Scientist

NCS Pte Ltd • Singapore

On-site
SGD 90,000 - 140,000
#EG Fullstack Developer – AI Applications
#EG Fullstack Developer – AI Applications

NCS Group • Singapore

On-site
SGD 70,000 - 110,000