Cloud Native Computing Platform Site Reliability Engineer

Tencent

Singapore

On-site

SGD 60,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tencent is seeking a technology specialist in Singapore to optimize GPU/CPU infrastructure and manage Kubernetes clusters. You will be responsible for daily operations, troubleshooting, and enhancing resource efficiency and service reliability.

The ideal candidate should have hands-on experience in GPU hardware and cloud platforms like AWS, with solid skills in Docker and Kubernetes. Tencent promotes diversity and equal opportunity, ensuring a supportive environment for all employees.

Qualifications

  • Hands-on experience in GPU hardware/drivers, CUDA, NCCL, and Mellanox operations.
  • Familiar with cloud native container technologies and disaster recovery solutions.
  • Strong accountability, self-motivation, and teamwork spirit.

Responsibilities

  • Responsible for daily operations and troubleshooting of GPU/CPU infrastructure.
  • Manage and operate Kubernetes clusters and ML platforms.
  • Drive automation of operational workflows.

Skills

GPU/ML principles
Cloud platforms (AWS)
Docker/Kubernetes operations
Linux/Shell environments
Programming (Go/Python/Java)

Job description

Technology Engineering Group (TEG) is responsible for supporting the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.

What The Role Entails
  • Responsible for daily operations, hardware/software troubleshooting, and optimization of GPU/CPU computing infrastructure to enhance resource efficiency and service reliability.
  • Manage and operate Kubernetes clusters and ML platforms, including monitoring/alerting, version upgrades, disaster recovery optimization, and security drills to ensure system high availability and maintainability.
  • Drive automation of operational workflows covering resource management, change control, self-healing solutions, and user tools.
Who We Look For
  • Proficient in GPU/ML principles and cloud platforms (eg. AWS); Hands‑on experience in GPU hardware/drivers, CUDA, NCCL, and Mellanox network operations/optimization; Data center experience preferred.
  • Familiar with cloud native container technologies and disaster recovery solutions; Practical Docker/Kubernetes operations experience required.
  • Skilled in Linux/Shell environments; Proficient in ≥1 language (Go/Python/Java); Adept at leveraging automation/AI-driven methods to further enhance service stability and efficiency.
  • Strong accountability and self‑motivation; Excellent learning/communication skills with demonstrated logical analysis, abstraction capabilities, and teamwork spirit.
Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Native SRE: Kubernetes, GPUs & ML Platforms
Cloud Native SRE: Kubernetes, GPUs & ML Platforms

Tencent • Singapore

On-site
SGD 60,000 - 90,000
Cloud Engineer
Cloud Engineer

Lightspeed Studios • Singapore

On-site
SGD 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

Lightspeed Studios • Singapore

On-site
SGD 50,000 - 70,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Lightspeed Studios • Singapore

On-site
SGD 60,000 - 80,000
Big Data Development Engineer Intern
Big Data Development Engineer Intern

Tencent • Singapore

On-site
SGD 90,000 - 130,000
Cloud Site Relibility Engineer - DCS Singapore Regular
Cloud Site Relibility Engineer - DCS Singapore Regular

ByteDance • Singapore

On-site
SGD 75,000 - 100,000
AI Compute Intern
AI Compute Intern

Tencent • Singapore

On-site
SGD 60,000 - 90,000
Research Engineer
Research Engineer

Lightspeed Studios • Singapore

On-site
SGD 80,000 - 120,000
[Tencent Cloud 2026 Internship Program] Cloud Reliability Engineering Internship (Shenzhen)
[Tencent Cloud 2026 Internship Program] Cloud Reliability Engineering Internship (Shenzhen)

Tencent • Singapore

On-site
SGD 8,500 - 17,000
Internship allowance
Accommodation allowance
Return flight ticket allowance
+3
Tencent Cloud - Senior Solution Architect, Big Data (Singapore)
Tencent Cloud - Senior Solution Architect, Big Data (Singapore)

Tencent • Singapore

On-site
SGD 150,000 - 210,000