GPGPU Software Architect/ Principal Engineer

XPENG

Santa Clara (CA)

On-site

USD 241,800 - 409,200

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
Vision insurance
401(k)

Job summary

A leading smart technology company is seeking a GPGPU Software Architect in Santa Clara, CA. The role focuses on developing a comprehensive software stack compatible with CUDA and collaborating with cross-functional teams in AI and GPU technologies. Candidates should have over ten years of experience in systems software with proficiency in CUDA. Benefits include competitive salary, medical insurance, and 401(k).

Qualifications

  • 10+ years in systems software, with 5+ years in designing CUDA Compute stacks.
  • Experience leading end-to-end development of GPU Runtime or AI acceleration library.
  • Comprehensive mastery of PTX/SASS, CUDA Driver API, and cuBLAS/cuDNN internals.

Responsibilities

  • Develop and refine a roadmap for a software stack compatible with CUDA.
  • Build an observability platform for real-time metrics.
  • Collaborate across teams to develop Device Plugins and GPU Operators.

Skills

Software Technical Strategy
CUDA expertise
AI framework collaboration
Benchmarking and observability

Tools

CUDA
PCIe
LLVM

Job description

GPGPU Software Architect/ Principal Engineer

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.

Our pioneering first-generation NPU, utilizing DSA architecture, has successfully entered mass production. We're currently validating the architecture of our second generation and are making the strategic decision to transition towards General Purpose GPU (GPGPU) architecture. We're completely overhauling our software stack and embracing the CUDA ecosystem. Our goal is to achieve over 90% compatibility with cuBLAS/cuDNN on Linux across PCIe and CXL connections, all while delivering at least 1.3 times the performance of existing solutions on Transformer and Stable-Diffusion workloads.

Job Responsibilities
  • Software Technical Strategy: Develop and refine a comprehensive 3-year roadmap for a software stack compatible with CUDA, encompassing Runtime, Driver, Compiler, Profiler, Debugger, and AI acceleration libraries
  • Define binding specifications that link our upcoming GPU ISA to CUDA APIs, ensuring forward compatibility with CUDA 12.x features
  • Evaluate and integrate the latest technological advancements: CUDA Graph, Transformer Engine, virtual memory management, CUDA dynamic CUTLASS 3.x, TMA, Blackwell FP4, among others
  • Define the task launch protocol, including Queue, Stream, Event, and Graph, as well as the memory model
  • Design a dual-mode (JIT & offline) compiler supporting LTO, PGO, Auto-Tuning, and efficient PTX→ISA microcode caching
  • Develop GPU virtualization schemes (MIG) that work across processes and containers
Performance & Observability
  • Build an observability platform: Nsys-compatible traces, real-time Metric-QPS dashboards, and an AI Advisor for identifying bottlenecks automatically
  • Manage internal AI benchmarks as the single source of truth. Benchmark includes MLPerf Inference, Stable Diffusion XL, and 70B LLM
Cross-functional Collaboration
  • Co-design ISA compatible with CUDA Compute Capability 12.x with our hardware architecture team
  • Collaborate with AI framework teams (PyTorch, TensorFlow, JAX, ONNX Runtime) to build fully reusable kernel libraries
  • Partner with Cloud and Kubernetes teams to co-develop Device Plugins, GPU Operators, and RDMA Network Policies
  • 10 years+ in systems software, with at least 5 years in designing CUDA Compute stacks
  • Led end-to-end development of a GPU Runtime or AI acceleration library generation
  • Comprehensive mastery of PTX/SASS, CUDA Driver API, and cuBLAS/cuDNN internals; experience with LLVM NVPTX backend
  • Profound understanding of GPU micro-architecture, including SM architecture, Warp Scheduler, Shared-Memory conflicts, and Tensor Core pipelines
  • Proficiency with PCIe/CXL/RDMA topologies, NUMA settings, and GPU Direct RDMA/Storage
Compensation

The base salary range for this full-time position is $241,800 - $409,200 in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.

Equity, Benefits & Equal Opportunity

We are an Equal Opportunity Employer. It is our policy to provide equal employment opportunities to all qualified persons without regard to race, age, color, sex, sexual orientation, religion, national origin, disability, veteran status or marital status or any other prescribed category set forth in federal or state regulations.

Location

Santa Clara, CA

Experience & Seniority
  • Seniority level: Director
  • Employment type: Full-time
  • Job function: Industries — Motor Vehicle Manufacturing
Benefits
  • Medical insurance
  • Vision insurance
  • 401(k)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Stock options
Comprehensive benefits package
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity opportunities
Comprehensive benefits package
Principal Software Engineer
Principal Software Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
System Software Engineer, Performance – CUDA Driver
System Software Engineer, Performance – CUDA Driver

Nvidia • Jasper (AL)

On-site
USD 124,000 - 196,000
Equity
Benefits
Senior Staff Engineer - Employee Productivity
Senior Staff Engineer - Employee Productivity

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Senior GPU PCIe and Boot Architect
Senior GPU PCIe and Boot Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Principal Software Engineer
Principal Software Engineer

Talentify • Redmond (WA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

NVIDIA • Santa Clara (CA)

On-site
USD 248,000 - 391,000
Equity
Benefits
Senior Director, NCP and ISV Business Development
Senior Director, NCP and ISV Business Development

NVIDIA • California (MO)

On-site
USD 304,000 - 460,000