JOB OVERVIEW
We are looking for a mid-senior MLOps Engineer to join the Artificial Intelligence Development Office - Singapore General Hospital's in-house AI team. AIDO functions as an end-to-end product team, building and deploying AI solutions that directly support clinical decision-making.
You will work closely with our technical leads to build out MLOps and platform engineering functions from the ground up. You will be involved in tooling decisions, pipeline setup, deployment workflows, and engineering standards as we scale our clinical AI projects into production.
KEY RESPONSIBILITIES
Platform & Infrastructure
- Support the design and build-out of AIDO's MLOps platform - contributing to tooling choices, deployment standards, and infrastructure decisions
- Build and maintain CI/CD pipelines for model training, validation, containerisation, and deployment
- Set up and maintain monitoring, alerting, and observability solutions covering model performance, data drift, latency, and infrastructure health
- Maintain clear documentation for infrastructure and deployment processes
Cloud Strategy & Architecture
- Support and drive the deployment of clinical AI products to cloud infrastructure including AWS
- Implement and configure cloud-native tooling suitable for AI workloads in a regulated healthcare environment
- Support the team's progression toward hybrid and cloud-ready infrastructure
Engineering Standards & Team Enablement
- Help the department to establish and adopt engineering best practices - code review processes, branching strategies, deployment checklists, etc.
- Work with data scientists to package models as reproducible, production-ready services (REST APIs, batch inference)
- Translate ambiguous requirements into clear technical plans and execute with minimal direction and approach problems systematically
REQUIRED SKILLS & EXPERIENCE
- STEM Bachelor degree with 3-5 years of experience in MLOps, DevOps, software engineering, or cloud infrastructure roles. Relevant post-graduate degree would be an advantage.
- Hands-on experience with AWS services — EKS, ECR, S3, IAM, CloudWatch
- Proficiency in Python for scripting, automation, and service development
- Working experience with Docker and Kubernetes in real-world deployment environments
- Experience with building or maintaining CI/CD pipelines — GitHub Actions, Jenkins, GitLab CI, or equivalent
- Familiarity with ML lifecycle tooling such as MLflow for experiment tracking and model versioning
- Familiarity with AI/ML platform concepts — model registries, RAG pipelines, and LLM frameworks such as LangChain or LlamaIndex
- Basic understanding of model monitoring and observability practices
- Comfortable working in highly regulated or on-premise environments — not exclusively cloud-native
- Ability to translate ambiguous requirements into clear technical plans and execute with minimal direction
- Curious and self-directed learner — comfortable picking up new tools and technologies without a formal training path
- Good communication skills — able to work alongside data scientists, engineers, IT professionals and non-technical stakeholders
PREFERRED QUALIFICATIONS
- Exposure to LLM serving infrastructure such as vLLM
- Background in healthcare, government, or regulated environments
- Familiarity with Prefect, Airflow, or similar orchestration tools
- Experience with Prometheus, Grafana, or ELK stack
- Experience with OpenShift or Red Hat environments
- Prior experience at GovTech, AISG, A*STAR, DSO, or similar Singapore public sector organisations
- AWS certification is a plus — DevOps Engineer, Machine Learning Specialty, or Solutions Architect Associate
TECHNICAL ENVIRONMENT
Python, FastAPI, Docker, Kubernetes, AWS (EKS / ECR / S3 / IAM / CloudWatch), MLflow, Prefect / Airflow, vLLM, CI/CD (GitHub Actions / Jenkins / GitLab CI), Prometheus, Grafana, OpenShift / Red Hat AI, Nginx