Member of Technical Staff, Kernel Engineering

INFERACT SINGAPORE PTE. LTD.

Singapore

In loco

SGD 167.000 - 335.000

Tempo pieno

9 giorni fa
Generatore di candidature

Ricevi una risposta da questo datore di lavoro — un curriculum e una lettera di presentazione personalizzati, che corrispondono esattamente a ciò che sta cercando.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Equity
Medical, dental, vision

Descrizione del lavoro

Inferact Singapore PTE. LTD. seeks a performance engineer to squeeze every FLOP from modern accelerators and write kernels and low-level optimizations for vLLM. You will ensure compatibility across NVIDIA GPUs and emerging silicon, collaborating with hardware teams to push peak throughput.

The role requires deep CUDA experience, strong C++/Python skills, and expertise in profiling tools. Fresh grads welcome; equity is offered along with medical, dental, and vision benefits.

Competenze

  • Bachelor's degree or equivalent in CS/Engineering or similar.
  • Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
  • Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
  • Proficiency in C++ and Python with ability to write high-performance code.
  • Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.
  • Obsession with benchmarks and squeezing every percentage point of speedup.

Mansioni

  • Write kernels and low-level optimizations for vLLM to maximize performance across accelerator types.
  • Collaborate with hardware vendors to extract maximum throughput on new chips.
  • Develop portable, high-performance C++/Python code and profiling workflows.

Conoscenze

CUDA kernels
GPU architecture
C++
Python
Nsight/rocprof
benchmarking
performance optimization

Formazione

Bachelor's degree in CS/Engineering

Strumenti

CuTeDSL
Triton
TileLang
Pallas

Descrizione del lavoro

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware.

Skills and Qualifications
Minimum qualifications:
  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
  • Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
  • Proficiency in C++ and Python with demonstrated ability to write high-performance code.
  • Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.
  • Obsession with benchmarks and squeezing every percentage point of speedup.
Preferred qualifications:
  • Experience with ML-specific kernel optimization (FlashAttention, fused kernels).
  • Knowledge of quantization techniques (INT8, FP8, mixed-precision).
  • Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).
  • Experience with compiler technologies (LLVM, MLIR, XLA).
Bonus points if you have:
  • Kernel-related contributions to vLLM or other inference engine projects.
  • Contributions to open-source GPU, ML systems, or compiler optimization projects
  • Written deep technical blogs on GPU optimization.
Logistics

Fresh graduates are welcome to apply. Minimum years of work experience: 0.

Compensation: Monthly salary of S$15,000 to S$30,000, depending on background, skills, and experience, plus equity.

Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Member of Technical Staff, AMD GPU Performance Engineering
Member of Technical Staff, AMD GPU Performance Engineering

Inferact • Singapore

In loco
SGD 200.000 - 400.000
Medical coverage
Dental coverage
Vision coverage
+1
Member of Technical Staff, AMD GPU Performance Engineering
Member of Technical Staff, AMD GPU Performance Engineering

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Equity
Medical insurance
Dental
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, TPU Performance Engineering

Inferact • Singapore

In loco
SGD 200.000 - 400.000
Medical, dental, and vision coverage
Equity options
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, TPU Performance Engineering

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Medical coverage
Dental coverage
Vision coverage
+1
Member of Technical Staff, Cloud Orchestration
Member of Technical Staff, Cloud Orchestration

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Medical coverage
Dental coverage
Vision coverage
+1
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

Inferact • Singapore

In loco
SGD 200.000 - 400.000
Medical coverage
Dental coverage
Vision coverage
+1
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Equity
Medical, dental, and vision coverage
Kernel Performance Engineer – GPU AI Inference
Kernel Performance Engineer – GPU AI Inference

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Equity
Medical, dental, vision
Staff Engineer, Performance & Scale (Distributed Inference)
Staff Engineer, Performance & Scale (Distributed Inference)

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Equity
Medical coverage
Dental coverage
+2
AMD GPU Inference Performance Engineer
AMD GPU Inference Performance Engineer

INFERACT SINGAPORE PTE. LTD. • Singapore

In loco
SGD 167.000 - 335.000
Equity
Medical insurance
Dental