Staff System Engineer – AI Infrastructure

Cloudera

Singapore

On-site

SGD 180,000 - 260,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Generous PTO Policy
Flexible WFH Policy
Mental & Physical Wellness programs
Phone and Internet Reimbursement
Continued Career Development

Job summary

Cloudera in Singapore is seeking a Staff Systems Engineer – AI Infrastructure to design and implement scalable GPU-backed infrastructure for AI workloads across on‑premises and cloud environments. You will translate complex requirements into concrete engineering solutions and drive hands-on delivery across the stack.

This role emphasizes experimentation, benchmarking, and collaboration with product teams to improve performance, reliability, and scalability of AI platforms.

Qualifications

  • 8+ years in systems software, distributed infrastructure, platform engineering or related field.
  • Hands-on experience with production GPU-based infrastructure supporting AI workloads.
  • Ability to investigate ambiguous problems and validate solutions via experimentation and prototyping.

Responsibilities

  • Investigate AI workloads across heterogeneous GPU environments and validate concepts.
  • Develop infrastructure for production AI workloads including inference and serving.
  • Diagnose Linux, GPU runtimes, containers, networking, storage issues and optimize.
  • Investigate GPU resource efficiency and sharing/partitioning.
  • Build tooling for diagnostics, benchmarking, deployment validation.
  • Translate infrastructure findings into requirements and product improvements.

Skills

Distributed systems
Linux
Containers
GPU infrastructure
Performance engineering
Cloud integration

Tools

Kubernetes
NVIDIA MIG
NVIDIA NIM
Benchmarking tools

Job description

Business Area: Professional Services Seniority Level: Mid-Senior level

At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we’re the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises. As adoption of AI grows across our data and AI platform, customers are using Cloudera AI for increasingly diverse workloads across on-premises, private, sovereign, and public-cloud environments.

We are seeking a Staff Systems Engineer – AI Infrastructure to understand the infrastructure requirements and constraints behind these workloads and drive hands‑on engineering solutions that improve our product. This is a hands‑on system engineering role. You will investigate ambiguous technical scenarios, identify underlying systems constraints, and develop and validate solutions through experimentation, prototyping, benchmarking, and engineering work. Success means turning diverse workload and infrastructure scenarios into validated technical solutions and, where appropriate, scalable product capabilities and improvements.

As a Staff System Engineer, you will:

  • Work across AI Infrastructure & Systems Engineering: Investigate and design solutions for AI workloads across heterogeneous GPU environments, on‑premises datacenters, private infrastructure, and public clouds; reproduce complex scenarios and validate solutions through hands‑on experimentation and proof‑of‑concepts.
  • Work across AI Workload & Inference Engineering: Develop and optimize infrastructure solutions for production AI workloads, including inference and serving, considering workload characteristics, performance, GPU capacity, resource utilization, and deployment constraints.
  • Systems & Performance Engineering: Diagnose issues across Linux, GPU runtimes and drivers, containers, networking, storage, and hardware; identify root causes and validate solutions to performance, reliability, and scalability challenges.
  • GPU Resource Efficiency: Investigate approaches for efficiently allocating and utilizing GPU resources across AI workloads, including workload‑aware sharing and partitioning where appropriate.
  • Infrastructure Tooling & Validation: Build diagnostic, benchmarking, deployment, and validation tooling to reproduce complex infrastructure scenarios and evaluate product performance across different environments.
  • Product & Engineering Collaboration: Translate infrastructure findings into technical requirements, product improvements, performance optimizations, and reusable platform capabilities in partnership with product and engineering teams.

We’re excited about you if you have:

  • 8+ years of experience in systems software, distributed infrastructure, platform engineering, performance engineering, or a related field, with a track record of independently solving complex, ambiguous engineering problems.
  • Hands‑on experience with production GPU‑based infrastructure supporting AI workloads, with a strong understanding of the infrastructure characteristics and constraints that affect them.
  • Ability to take ambiguous problems, develop hypotheses, investigate root causes, and build or validate solutions through experimentation, debugging, prototyping, and measurement without requiring step‑by‑step direction.
  • Strong understanding of distributed systems, Linux, containers, and production infrastructure, with the ability to reason across multiple layers of the technology stack.
  • Demonstrated ability to diagnose and optimize bottlenecks involving GPU utilization, compute, memory, networking, I/O, or workload/runtime behavior.
  • Hands‑on experience designing or operating production infrastructure in on‑premises, private‑cloud, and/or public‑cloud environments, with an understanding of the practical constraints of heterogeneous environments.

You may also have:

  • Experience with modern AI inference and serving technologies such as NVIDIA NIM, vLLM, SGLang, Triton, or equivalent.
  • Experience with Kubernetes or distributed AI workload orchestration.
  • Experience with GPU resource management, sharing, or partitioning, including technologies such as NVIDIA MIG.
  • Experience diagnosing high‑performance GPU networking or distributed communication issues.
  • Experience building infrastructure diagnostics, benchmarks, or proof‑of‑concept systems to evaluate new architectures or technologies.

What you can expect from us:

  • Generous PTO Policy
  • Support work life balance with Unplugged Days
  • Flexible WFH Policy
  • Mental & Physical Wellness programs
  • Phone and Internet Reimbursement program
  • Access to Continued Career Development
  • Comprehensive Benefits and Competitive Packages
  • Paid Volunteer Time
  • Employee Resource Groups

EEO/VEVRAA #LI-RC1

If you have any questions about our privacy practices, please contact us at privacy@cloudera.com.

EEO/VEVRAA If you need assistance with applying for a position, please email our office at talentacquisition@cloudera.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Infrastructure Systems Engineer
Staff AI Infrastructure Systems Engineer

Cloudera • Singapore

On-site
SGD 180,000 - 260,000
Generous PTO Policy
Flexible WFH Policy
Mental & Physical Wellness programs
+2
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SWAPETECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Cloudera • Singapore

On-site
SGD 180,000 - 240,000
Generous PTO
WFH Policy
Wellness programs
+1
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 190,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Enterprise Data & AI Platform Engineer
Enterprise Data & AI Platform Engineer

OVERSEA-CHINESE BANKING CORPORATION LIMITED • Singapore

On-site
SGD 120,000 - 180,000
Cloudera Certifications
AWS Certifications
AB03 - AI Infrastructure Engineer
AB03 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000