Senior Engineer - Compute, Container & AI Infrastructure

Singtel Group

Singapore

On-site

SGD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Singtel Group in Singapore seeks an experienced Infrastructure Engineer to design, operate and enhance enterprise AI infrastructure powered by Red Hat OpenShift AI, Kubernetes, and DevSecOps practices. You will enable container and VM workloads at scale while ensuring platform reliability, security, and automation.

The role covers OpenShift platform engineering, GitOps, and automation, with focus on AI/ML workloads and NVIDIA GPU infrastructure.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 10+ years of experience in platform engineering, infrastructure engineering, cloud platforms, Kubernetes, or enterprise IT environments.
  • 5+ years of hands-on experience managing Red Hat OpenShift Container Platform in production environments.
  • Strong background in running Kubernetes, container technologies, Linux/RHEL, and cloud-native infrastructure at scale, with ability to troubleshoot complex container and infrastructure issues.
  • Experience with platform operations, troubleshooting, patching, upgrades, automation, and capacity planning.
  • Experience with enterprise DevOps, GitOps, CI/CD, and service management practices.
  • Hands-on experience supporting x86 servers and NVIDIA GPU infrastructure.
  • Good understanding of AI/ML workloads and concepts, including model serving, inferencing, MIG, RAG.

Responsibilities

  • Administer and support OpenShift Container Platform environments, including OpenShift Virtualization.
  • Manage platform operations: cluster health, node management, networking, storage, monitoring, logging, security, upgrades.
  • Troubleshoot Kubernetes and OpenShift issues across large clusters and workloads.
  • Support Hosted Control Plane environments and lifecycle management.
  • Manage multi-cluster OpenShift environments and enterprise integrations (DNS, identity, storage, certificates).
  • Support OpenShift Service Mesh, OpenShift Virtualization, and other cloud-native capabilities.
  • Maintain runbooks, documentation and SLA compliance.
  • Drive GitOps-based deployment and CI/CD pipelines.

Skills

Extensive platform engineering
Kubernetes
OpenShift
GitOps
CI/CD
DevSecOps
Automation
Linux / RHEL
NVIDIA GPU infrastructure

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Red Hat OpenShift Container Platform (OCP)
Red Hat OpenShift AI (RHOAI)
Kubernetes
NVIDIA GPU Operator
NVIDIA MIG

Job description

Select how often (in days) to receive an alert:

We are building an AI-first telco of the future — connecting people, businesses and communities through trusted networks, digital services and next-generation technology. From everyday consumer experiences across mobile, broadband, entertainment and lifestyle services, to enterprise solutions in 5G+, cloud, cybersecurity, AI, data centres and digital platforms, we create the infrastructure and experiences that power how people live, work, play and do business.

You will be part of teams shaping intelligent networks, secure digital ecosystems, sovereign AI cloud platforms, sustainable data centres and seamless customer experiences at scale. Together, we combine technology, innovation and human potential to empower every generation to thrive in a more connected, resilient and AI-enabled future.

Be part of something BIG!

As part of our Infrastructure Engineering team, you will design, operate, and enhance Singtel's enterprise AI infrastructure powered by Red Hat OpenShift AI, Kubernetes, and DevSecOps practices. You will play a critical role in enabling container and VM workloads at scale while ensuring platform reliability, performance, security, and automation.

How You Will Make an Impact:
OpenShift Container, Virtualization and RHOAI Platform Engineering
  • Administer and support enterprise Red Hat OpenShift Container Platform (OCP) environments, including Openshift Virtualization, Containerization and AI variations
  • Manage platform operations including cluster health, node management, networking, storage, monitoring, logging, security, and upgrades.
  • Troubleshoot Kubernetes and OpenShift issues across large clusters, workloads, networking, and authentication layers, as well as Operators necessary to build, develop and operate the platform efficiently and reliably.
  • Support Hosted Control Plane (HyperShift) environments, including HostedCluster and NodePool lifecycle management.
  • Manage multi-cluster OpenShift environments and enterprise integrations such as DNS, identity, storage, certificate management, and security services.
  • Support OpenShift Service Mesh (Istio), OpenShift Virtualization, and other cloud-native platform capabilities.
  • Maintain operational standards, runbooks, and documentation while ensuring SLA compliance.
GitOps, DevOps & Platform Automation
  • Manage and support OpenShift GitOps (Argo CD) for platform and application deployments.
  • Implement GitOps-based deployment and configuration management practices.
  • Support CI/CD pipelines spanning source control, security scanning, container registries, and OpenShift deployments.
  • Manage Helm-based application deployments and lifecycle management.
  • Drive DevSecOps practices including vulnerability scanning, access controls, secrets management, and policy enforcement.
  • Automate platform operations using Infrastructure-as-Code and Configuration-as-Code principles using Ansible scripts, Powershell or Unix Bash scripts as appropriate.
  • Develop reusable deployment templates, standards, and automation frameworks.
  • Manage enterprise x86 server infrastructure supporting OpenShift, OpenShift AI, and private cloud platforms.
  • Support NVIDIA GPU servers for AI/ML and accelerated computing workloads.
  • Perform infrastructure lifecycle management including installation, upgrades, firmware updates, capacity planning, health monitoring, maintenance, and hardware replacement.
  • Monitor and optimize server, storage, networking, and GPU performance.
  • Troubleshoot hardware and infrastructure issues across compute, storage, networking, firmware, and GPU layers.
  • Work with vendors on support, maintenance, and infrastructure refresh initiatives.
  • Support data center operations and future infrastructure expansion projects.
Skills for Success
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 10+ years of experience in platform engineering, infrastructure engineering, cloud platforms, Kubernetes, or enterprise IT environments.
  • 5+ years of hands-on experience managing Red Hat OpenShift Container Platform in production environments.
  • Strong background in running Kubernetes, container technologies, Linux/RHEL, and cloud-native infrastructure at scale, with a proven track record being able to troubleshoot complex container and infrastructure issues.
  • Experience with platform operations, troubleshooting, patching, upgrades, automation, and capacity planning.
  • Experience with enterprise DevOps, GitOps, CI/CD, and service management practices.
  • Hands-on experience supporting x86 servers and NVIDIA GPU infrastructure.
  • Good understanding of AI/ML workloads and concepts, including model serving, inferencing, MIG, RAG.
Technical Expertise
  • Red Hat OpenShift Container Platform (OCP) and RedHat Openshift Virtualization
  • Red Hat OpenShift AI (RHOAI)
  • Kubernetes and container technologies
  • NVIDIA GPU Operator and GPU-accelerated workloads
  • NVIDIA Multi-Instance GPU (MIG)
  • OpenShift Networking, Security, RBAC, and Storage
  • Linux / RHEL administration and troubleshooting
  • Bare metal server and storage design, build and operations.
  • Automation using Ansible, Python, Bash, or similar technologies
  • Monitoring, logging, observability, and platform performance management
Nice to Have
  • DevSecOps and AI Security experience
  • Private Cloud Platform software
  • Red Hat certifications (RHCSA, RHCE, OpenShift)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineer - Compute, Container & AI Infrastructure
Senior Engineer - Compute, Container & AI Infrastructure

Singtel • Singapore

On-site
Confidential
Senior Engineer - Compute, Container & AI Infrastructure
Senior Engineer - Compute, Container & AI Infrastructure

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Senior AI Infra Engineer – OpenShift, Kubernetes & GPU
Senior AI Infra Engineer – OpenShift, Kubernetes & GPU

Singtel Group • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infra Engineer: OpenShift, Kubernetes & GPU
Senior AI Infra Engineer: OpenShift, Kubernetes & GPU

Singtel • Singapore

On-site
Confidential
Lead AI Platform Operations Engineer #AIDA
Lead AI Platform Operations Engineer #AIDA

Singtel • Singapore

On-site
Confidential
Senior AI Infra Engineer — OpenShift, Kubernetes & GPU
Senior AI Infra Engineer — OpenShift, Kubernetes & GPU

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Senior DevOps Engineer
Senior DevOps Engineer

Singtel • Singapore

On-site
Confidential
Senior Engineer, Cloud & Container Infra
Senior Engineer, Cloud & Container Infra

Singtel • Singapore

On-site
Confidential
DevOps Engineer (DevLabs by AngelHack)
DevOps Engineer (DevLabs by AngelHack)

STACKTRIBE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Lead AI Platform Engineer #AIDA
Lead AI Platform Engineer #AIDA

Singapore Telecommunications Limited • Singapore

Hybrid
SGD 180,000 - 260,000