AI Compute Intern

Proxima Beta Pte. Limited

Singapore

On-site

SGD 20,000 - 33,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tencent Overseas IT is seeking a motivated Network Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI-accelerated HPC infrastructure. You will collaborate with internal business teams and vendor engineers to build, operate and optimize clusters, gaining hands-on experience in cutting‑edge AI infrastructure.

Currently pursuing or recently completed a Bachelor's or Master's in computer engineering, CS, or a related field; you will work with

Qualifications

  • Pursuing or recently completed Bachelor's or Master's in computer engineering, CS, or related field.
  • Basic understanding of server hardware and Linux environments.
  • Exposure to Bash or Python scripting.
  • Familiarity with configuration management, CI/CD tools, workload managers, and cluster software (Slurm, Kubernetes, BCM).
  • Ability to work independently and as part of a team.

Responsibilities

  • Support the deployment, configuration, and maintenance of high-end servers, storage servers, networking equipment, and software components in secure environments.
  • Assist with hardware diagnostics, system functionality checks, and firmware updates as required.
  • Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, HPC clusters, Kubernetes, Slurm, etc.).
  • Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.
  • Document incident details, resolutions, and lessons learned to improve future problem-solving.
  • Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team.
  • Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.

Job description

About the Hiring Team

Tencent Overseas IT has the mission to empower Tencent’s rapid global growth with future ready, global IT platforms, applications and services. We are chartered to lead the Overseas IT strategy, architecture, roadmap and execution. Satisfying our internal/external customers and becoming a world class global IT team are our top aspirations.

What the Role Entails

Role Summary

We are seeking a motivated Network Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations and maintenance of AI-accelerated high-performance computing (HPC) infrastructure. In this role, you will work closely with internal business team and vendor's engineering team to build and operate clusters. This is a hands‑on opportunity to gain direct exposure to cutting‑edge AI infrastructure.

Team Introduction

The AI Compute Centre sits within Tencent's Overseas IT department, acting as the bridge between internal AI infrastructure demand and the external resources that fulfil it. We play key roles in the full lifecycle of AI infrastructure clusters — from requirement gathering and capacity planning, through architectural design and development, to delivery, operations, and DevOps — across regions worldwide.

Key Responsibilities
  • Support the deployment, configuration, and maintenance of high-end servers, storage servers, networking equipment, and software components in secure environments.
  • Assist with hardware diagnostics, system functionality checks, and firmware updates as required.
  • Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, HPC clusters, Kubernetes, Slurm, etc.).
  • Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.
  • Document incident details, resolutions, and lessons learned to improve future problem‑solving.
  • Maintain clear, accurate, and up‑to‑date documentation to support knowledge sharing across the team.
  • Participate in team meetings and knowledge‑sharing sessions to foster collaboration and continuous learning.
Who We Look For
  • Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field.
  • Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system‑level security standards.
  • Exposure to scripting languages such as Bash or Python.
  • Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes, BCM), and observability tools (e.g., Prometheus, Grafana, ELK).
  • Assist senior engineers in checking cluster network connectivity, including InfiniBand or RoCE links. Run predefined tests, collect results, and flag link or performance anomalies.
  • Monitor dashboards for compute, storage, and management networks. Help record bandwidth, latency, and RDMA test results, and compare them with agreed baselines.
  • Support AI infrastructure incident investigations by collecting switch, NIC, and host logs, updating incident records, and tracking vendor follow‑up actions under guidance.
  • Ability to work both independently and as part of a team.
Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Who we are

Tencent is a world‑leading internet and technology company that develops innovative products and services to improve the quality of life for people around the world.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. AI infra engineer
Sr. AI infra engineer

Proxima Beta Pte. Limited • Singapore

On-site
SGD 180,000 - 260,000
AI Compute SRE Intern — HPC & Clusters
AI Compute SRE Intern — HPC & Clusters

Proxima Beta Pte. Limited • Singapore

On-site
SGD 20,000 - 33,000
AI IT Engineer Intern
AI IT Engineer Intern

Lightspeed Studios • Singapore

On-site
SGD 12,000 - 18,000
AI Compute Intern
AI Compute Intern

Tencent • Singapore

On-site
SGD 20,000 - 33,000
Exposure to AI infrastructure
Software Engineer II
Software Engineer II

Proxima Beta Pte. Limited • Singapore

On-site
SGD 120,000 - 180,000
Data Engineer Intern
Data Engineer Intern

Proxima Beta Pte. Limited • Singapore

On-site
SGD 90,000 - 150,000
Data Engineer Intern
Data Engineer Intern

Tencent • Singapore

On-site
SGD 110,000 - 180,000
Data Engineer Intern
Data Engineer Intern

Tencent International Service Pte. Ltd. • Singapore

On-site
SGD 90,000 - 180,000
Senior Data Center Operations Engineer
Senior Data Center Operations Engineer

WeChat International Pte. Ltd. • Singapore

On-site
SGD 90,000 - 170,000
Site Reliability Engineer Intern
Site Reliability Engineer Intern

Tencent • Singapore

On-site
SGD 120,000 - 180,000