Senior Engineer - Compute, Container & AI Infrastructure

Singtel

Singapore

On-site

Confidential

Full time

20 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Singtel is building an AI-first telco and seeks an experienced platform engineer to design, operate, and scale enterprise AI infrastructure. You will manage Red Hat OpenShift, Kubernetes, and GPU-enabled workloads in a secure, automated environment.

The role focuses on building reliable, scalable platform services, automating operations, and collaborating across teams to deliver high availability AI capabilities for customers.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 10+ years of experience in platform engineering, infrastructure engineering, cloud platforms, Kubernetes, or enterprise IT environments.
  • 5+ years of hands-on experience managing Red Hat OpenShift Container Platform in production environments.
  • Strong background in Kubernetes, container technologies, Linux/RHEL, and cloud-native infrastructure at scale.
  • Experience with DevOps, GitOps, CI/CD, and service management practices.

Responsibilities

  • Design, operate, and enhance Singtel's AI infrastructure using OpenShift AI, Kubernetes, and DevSecOps practices.
  • Manage cluster health, upgrades, networking, storage, monitoring, security, and automation.
  • Support OpenShift GitOps (Argo CD) and CI/CD pipelines across teams.
  • Maintain operational standards, runbooks, and documentation while ensuring SLA compliance.

Skills

OpenShift expertise
Kubernetes administration
Linux system administration
Automation & Scripting
GPU infrastructure experience
DevSecOps

Education

Bachelor's degree in CS/IT/Engineering

Tools

Argo CD
Python
Bash
Ansible
NVIDIA GPU Operator

Job description

At Singtel, your work unlocks BIG Possibilities.
We are building an AI-first telco of the future - connecting people, businesses and communities through trusted networks, digital services and next-generation technology. From everyday consumer experiences across mobile, broadband, entertainment and lifestyle services, to enterprise solutions in 5G+, cloud, cybersecurity, AI, data centres and digital platforms, we create the infrastructure and experiences that power how people live, work, play and do business.

At Singtel, your work unlocks BIG Possibilities.
We are building an AI-first telco of the future - connecting people, businesses and communities through trusted networks, digital services and next-generation technology. From everyday consumer experiences across mobile, broadband, entertainment and lifestyle services, to enterprise solutions in 5G+, cloud, cybersecurity, AI, data centres and digital platforms, we create the infrastructure and experiences that power how people live, work, play and do business.

You will be part of teams shaping intelligent networks, secure digital ecosystems, sovereign AI cloud platforms, sustainable data centres and seamless customer experiences at scale. Together, we combine technology, innovation and human potential to empower every generation to thrive in a more connected, resilient and AI-enabled future.

Be part of something BIG!

As part of our Infrastructure Engineering team, you will design, operate, and enhance Singtel's enterprise AI infrastructure powered by Red Hat OpenShift AI, Kubernetes, and DevSecOps practices. You will play a critical role in enabling container and VM workloads at scale while ensuring platform reliability, performance, security, and automation.

How You Will Make An Impact

OpenShift Container, Virtualization and RHOAI Platform Engineering

  • Administer and support enterprise Red Hat OpenShift Container Platform (OCP) environments, including Openshift Virtualization, Containerization and AI variations
  • Manage platform operations including cluster health, node management, networking, storage, monitoring, logging, security, and upgrades.
  • Troubleshoot Kubernetes and OpenShift issues across large clusters, workloads, networking, and authentication layers, as well as Operators necessary to build, develop and operate the platform efficiently and reliably.
  • Support Hosted Control Plane (HyperShift) environments, including HostedCluster and NodePool lifecycle management.
  • Manage multi-cluster OpenShift environments and enterprise integrations such as DNS, identity, storage, certificate management, and security services.
  • Support OpenShift Service Mesh (Istio), OpenShift Virtualization, and other cloud-native platform capabilities.
  • Maintain operational standards, runbooks, and documentation while ensuring SLA compliance.

GitOps, DevOps & Platform Automation

  • Manage and support OpenShift GitOps (Argo CD) for platform and application deployments.
  • Implement GitOps-based deployment and configuration management practices.
  • Support CI/CD pipelines spanning source control, security scanning, container registries, and OpenShift deployments.
  • Manage Helm-based application deployments and lifecycle management.
  • Drive DevSecOps practices including vulnerability scanning, access controls, secrets management, and policy enforcement.
  • Automate platform operations using Infrastructure-as-Code and Configuration-as-Code principles using Ansible scripts, Powershell or Unix Bash scripts as appropriate.
  • Develop reusable deployment templates, standards, and automation frameworks.

Infrastructure & GPU Server Engineering

  • Manage enterprise x86 server infrastructure supporting OpenShift, OpenShift AI, and private cloud platforms.
  • Support NVIDIA GPU servers for AI/ML and accelerated computing workloads.
  • Perform infrastructure lifecycle management including installation, upgrades, firmware updates, capacity planning, health monitoring, maintenance, and hardware replacement.
  • Monitor and optimize server, storage, networking, and GPU performance.
  • Troubleshoot hardware and infrastructure issues across compute, storage, networking, firmware, and GPU layers.
  • Work with vendors on support, maintenance, and infrastructure refresh initiatives.
  • Support data center operations and future infrastructure expansion projects.
Skills For Success
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 10+ years of experience in platform engineering, infrastructure engineering, cloud platforms, Kubernetes, or enterprise IT environments.
  • 5+ years of hands- on experience managing Red Hat OpenShift Container Platform in production environments.
  • Strong background in running Kubernetes, container technologies, Linux/RHEL, and cloud-native infrastructure at scale, with a proven track record being able to troubleshoot complex container and infrastructure issues.
  • Experience with platform operations, troubleshooting, patching, upgrades, automation, and capacity planning.
  • Experience with enterprise DevOps, GitOps, CI/CD, and service management practices.
  • Hands- on experience supporting x86 servers and NVIDIA GPU infrastructure.
  • Good understanding of AI/ML workloads and concepts, including model serving, inferencing, MIG, RAG.
Technical Expertise
  • Red Hat OpenShift Container Platform (OCP) and RedHat Openshift Virtualization
  • Red Hat OpenShift AI (RHOAI)
  • Kubernetes and container technologies
  • OpenShift GitOps / Argo CD
  • NVIDIA GPU Operator and GPU-accelerated workloads
  • NVIDIA Multi-Instance GPU (MIG)
  • OpenShift Networking, Security, RBAC, and Storage
  • Linux / RHEL administration and troubleshooting
  • Bare metal server and storage design, build and operations.
  • Automation using Ansible, Python, Bash, or similar technologies
  • Monitoring, logging, observability, and platform performance management
Nice to Have
  • Hosted Control Plane (HyperShift)
  • OpenShift Service Mesh / Istio
  • OpenShift Virtualization (KubeVirt)
  • DevSecOps and AI Security experience
  • Disconnected/OpenShift private environments
  • Private Cloud Platform software
  • Red Hat certifications (RHCSA, RHCE, OpenShift)
  • Kubernetes certifications (CKA, CKAD)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineer - Compute, Container & AI Infrastructure
Senior Engineer - Compute, Container & AI Infrastructure

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Senior Engineer - Compute, Container & AI Infrastructure
Senior Engineer - Compute, Container & AI Infrastructure

Singtel Group • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infra Engineer – OpenShift, Kubernetes & GPU
Senior AI Infra Engineer – OpenShift, Kubernetes & GPU

Singtel Group • Singapore

On-site
SGD 180,000 - 240,000
Lead AI Platform Operations Engineer #AIDA
Lead AI Platform Operations Engineer #AIDA

Singtel • Singapore

On-site
Confidential
Senior AI Infra Engineer: OpenShift, Kubernetes & GPU
Senior AI Infra Engineer: OpenShift, Kubernetes & GPU

Singtel • Singapore

On-site
Confidential
Senior DevOps Engineer
Senior DevOps Engineer

Singtel • Singapore

On-site
Confidential
Senior AI Infra Engineer — OpenShift, Kubernetes & GPU
Senior AI Infra Engineer — OpenShift, Kubernetes & GPU

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Lead AI Platform Engineer #AIDA
Lead AI Platform Engineer #AIDA

Singtel Group • Singapore

Hybrid
SGD 180,000 - 260,000
Data Platform Infrastructure Engineer
Data Platform Infrastructure Engineer

Singapore Telecommunications Limited • Singapore

On-site
SGD 120,000 - 180,000
Data Platform Infrastructure Engineer
Data Platform Infrastructure Engineer

Singtel • Singapore

On-site
Confidential