L3 Escalation Support Engineer (Cloud Infrastructure / HPC)

DADACONSULTANTS PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DADACONSULTANTS PTE. LTD. is seeking an experienced Infrastructure Engineer to own the most complex customer escalations and platform incidents. You will bridge front-line support with Engineering/SRE to drive permanent improvements for a cloud platform serving AI and HPC workloads.

Requirements include 5+ years in cloud infrastructure, strong Linux/networking skills, and a proven incident-resolution track record. On-call rotation and cross-functional collaboration are essential.

Qualifications

  • 5+ years of hands-on experience in cloud infrastructure support or SRE-adjacent roles.
  • Strong Linux, networking, and incident-management skills.
  • Experience resolving complex production incidents and root-cause analysis is required.

Responsibilities

  • Own complex, escalated customer issues and platform incidents end-to-end.
  • Lead deep-dive troubleshooting across GPU compute, networking/SDN, storage, and billing systems.
  • Coordinate with SRE, Compute and R&D teams to implement permanent fixes.
  • Lead incident response and post-incident reviews to improve runbooks and monitoring.
  • Mentor junior engineers and participate in on-call escalation rotations.
  • Collaborate with engineering and product teams to feed learnings into the platform roadmap.

Skills

Linux
Cloud infrastructure
SRE practices
On-call readiness

Job description

About the Role

Our client is a world-leading technology company operating at the intersection of high-performance computing, Bitcoin mining, and AI cloud services, with a globally distributed infrastructure footprint spanning multiple continents and a multi-gigawatt energy portfolio. Headquartered in Singapore, they are scaling rapidly and investing heavily in next-generation datacenter and cloud capabilities. This role sits at the deepest technical tier of the operations centre, owning the most complex customer escalations and platform incidents as the critical bridge between front-line support and Engineering/SRE teams. It is an exceptional opportunity for a seasoned infrastructure engineer to drive permanent, systemic improvements rather than just resolving individual tickets. You will directly shape the reliability and quality of a cloud platform serving high-demand AI and HPC workloads.

Key Responsibilities
  • Own complex, escalated customer issues and platform incidents end-to-end, driving each case through to full resolution
  • Conduct deep-dive troubleshooting across GPU compute , networking/SDN, storage, drivers, control plane, and billing systems
  • Serve as the primary liaison to SRE, Compute, and R&D teams - leading root-cause analysis and ensuring permanent fixes are implemented
  • Lead and support incident response and post-incident reviews, translating findings into improved runbooks and monitoring enhancements
  • Identify recurring escalation patterns and convert them into lasting platform or process improvements
  • Mentor L1 and L2 support engineers, raising escalation quality and expanding the team knowledge base
  • Participate in an on-call escalation rotation, providing incident leadership during critical platform events
  • Collaborate cross-functionally with engineering and product teams to feed operational learnings back into the platform roadmap
Requirements
  • Minimum 5 years of hands-on experience in cloud infrastructure technical support, escalation engineering, or SRE-adjacent roles
  • Strong practical proficiency in Linux, networking, and cloud infrastructure; experience with GPU/CUDA/HPC environments (is a bonus)
  • Demonstrated track record of resolving complex production incidents and conducting thorough root-cause analysis
  • Excellent written English communication skills, with the ability to work effectively across technical and non-technical stakeholders
  • Comfortable with on-call responsibilities and experienced in taking ownership during high-pressure incident scenarios
  • Familiarity with SDN, distributed storage systems, or control plane architecture (is a bonus)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Escalation Engineer - Cloud Infra & HPC
Senior Escalation Engineer - Cloud Infra & HPC

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infrastructure Support Engineer
Senior AI Infrastructure Support Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Escalation / L3 Support Engineer
Escalation / L3 Support Engineer

Bitdeer Group • Singapore

On-site
SGD 110,000 - 170,000
L3 Infrastructure Engineer (Experience in Dell, HP, Lenovo are preferred; Break-fix technical role)
L3 Infrastructure Engineer (Experience in Dell, HP, Lenovo are preferred; Break-fix technical role)

RECRUIT EXPRESS PTE LTD • Singapore

On-site
SGD 100,000 - 150,000
Senior L3 Cloud Support Engineer - Escalations
Senior L3 Cloud Support Engineer - Escalations

Bitdeer Group • Singapore

On-site
SGD 110,000 - 170,000
L3 Infrastructure Engineer (Experience in Dell, HP, Lenovo are preferred; Break-fix technical role, up to $9000, Woodlands area) #IRT
L3 Infrastructure Engineer (Experience in Dell, HP, Lenovo are preferred; Break-fix technical role, up to $9000, Woodlands area) #IRT

RECRUIT EXPRESS PTE LTD • Singapore

On-site
SGD 78,000 - 100,000
AI Infrastructure Support Engineer
AI Infrastructure Support Engineer

NSCALE OPERATIONS APAC PTE. LTD. • Singapore

On-site
SGD 65,000 - 95,000
Senior AI Infrastructure Reliability Engineer
Senior AI Infrastructure Reliability Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Cloud Support Engineer
Cloud Support Engineer

AlphaSense Oy • Singapore

On-site
SGD 60,000 - 92,000
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 140,000 - 190,000