Cloud Native Computing Platform SRE Engineer

Tencent

Singapore

On-site

SGD 90,000 - 150,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Tencent Technology Engineering Group (TEG) seeks an experienced infrastructure engineer to optimize GPU/CPU computing resources and manage ML platforms. You will operate Kubernetes clusters, perform monitoring, upgrades, and disaster recovery, ensuring high availability.

Strong Linux skills, Docker/Kubernetes experience, and programming proficiency (Go/Python/Java) are essential. The role emphasizes automation, self-healing, and cross-team collaboration.

Qualifications

  • GPU/ML principles with cloud platforms (e.g., AWS)
  • Hands-on GPU hardware/drivers, CUDA, NCCL, Mellanox network operations/optimisation
  • Docker/Kubernetes operations experience required; familiar with cloud-native container tech
  • Linux/Shell environments; Go/Python/Java proficiency; automation/AI-driven methods to improve stability and efficiency

Responsibilities

  • Responsible for daily operations, hardware/software troubleshooting, and optimization of GPU/CPU computing infrastructure to enhance resource efficiency and service reliability.
  • Manage and operate Kubernetes clusters and ML platforms, including monitoring/alerting, version upgrades, disaster recovery optimization, and security drills to ensure system high availability and maintainability.
  • Drive automation of operational workflows covering resource management, change control, self-healing solutions, and user tools

Skills

GPU/ML principles
Cloud platforms (AWS)
Go
Python
Java
Linux/Shell
Automation/AI methods
Accountability & communication

Tools

Docker
Kubernetes

Job description

Technology Engineering Group (TEG) is responsible for supporting the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia, TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.

What The Role Entails
  • Responsible for daily operations, hardware/software troubleshooting, and optimization of GPU/CPU computing infrastructure to enhance resource efficiency and service reliability.
  • Manage and operate Kubernetes clusters and ML platforms, including monitoring/alerting, version upgrades, disaster recovery optimization, and security drills to ensure system high availability and maintainability.
  • Drive automation of operational workflows covering resource management, change control, self-healing solutions, and user tools.
Who We Look For
  • Proficient in GPU/ML principles and cloud platforms (eg. AWS) ; Hands-on experience in GPU hardware/drivers, CUDA, NCCL, and Mellanox network operations/optimization; Data center experience preferred.
  • Familiar with cloud native container technologies and disaster recovery solutions ; Practical Docker/Kubernetes operations experience required.
  • Skilled in Linux/Shell environments; Proficient in ≥1 language ( Go/Python/Java ); Adept at leveraging automation/AI-driven methods to further enhance service stability and efficiency.
  • Strong accountability and self-motivation ; Excellent learning/communication skills with demonstrated logical analysis, abstraction capabilities, and teamwork spirit.
Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Native Computing Platform Site Reliability Engineer
Cloud Native Computing Platform Site Reliability Engineer

Tencent • Singapore

On-site
SGD 60,000 - 90,000
Cloud Native SRE: GPU/ML & Kubernetes Engineer
Cloud Native SRE: GPU/ML & Kubernetes Engineer

Tencent • Singapore

On-site
SGD 90,000 - 150,000
Cloud Native SRE: Kubernetes, GPUs & ML Platforms
Cloud Native SRE: Kubernetes, GPUs & ML Platforms

Tencent • Singapore

On-site
SGD 60,000 - 90,000
Cloud Engineer
Cloud Engineer

Lightspeed Studios • Singapore

On-site
SGD 60,000 - 80,000
Cloud SRE Intern: Build Reliable Big Data Platforms
Cloud SRE Intern: Build Reliable Big Data Platforms

Tencent • Singapore

On-site
SGD 60,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

Lightspeed Studios • Singapore

On-site
SGD 50,000 - 70,000
Senior Big Data Platform SRE Engineer
Senior Big Data Platform SRE Engineer

Tencent • Singapore

On-site
SGD 90,000 - 150,000
Research Engineer
Research Engineer

Lightspeed Studios • Singapore

On-site
SGD 80,000 - 120,000
Senior Big Data Platform SRE: Scale, Optimize & Automate
Senior Big Data Platform SRE: Scale, Optimize & Automate

Tencent • Singapore

On-site
SGD 90,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Lightspeed Studios • Singapore

On-site
SGD 60,000 - 80,000