Senior GPU Performance Software Engineer

Intel Corporation

Hillsboro (OR)

Hybrid

USD 195,000 - 276,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health benefits
Retirement plan

Job summary

Intel Corporation in Hillsboro, Oregon is seeking a software engineer to contribute to oneDNN low-level optimization and GPU kernel development. The role focuses on optimizing math primitives and parallel algorithms for AI workloads on Intel hardware.

You will design JIT/codegen infrastructure, implement fusion and memory-traffic optimizations, and collaborate with hardware and compiler teams to shape future accelerator capabilities.

Qualifications

  • BSc, MSc, or PhD in CS/CE/Math/Physics or related field.
  • 5+ years of professional software development with modern C++.
  • 2+ years of GPU/kernel optimization experience (CUDA/OpenCL/SYCL/HIP).
  • Strong knowledge of computer architecture and SIMD.

Responsibilities

  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads.
  • Design scalable JIT and codegen infrastructure for GPU kernel generation.
  • Implement fusion and memory-traffic optimizations to maximize hardware utilization.
  • Optimize mixed-precision and quantized execution paths (BF16, FP16, INT8, FP8, FP4).
  • Profile and eliminate performance bottlenecks across GPU primitives and runtime paths.
  • Co-design GPU primitives and kernel architectures for next-generation Intel GPUs.
  • Partner with hardware and compiler teams to shape accelerator capabilities and software stacks.
  • Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads.

Skills

C++ expert
GPU kernel optimization
Computer architecture
Parallel programming

Education

BSc/MSc/PhD in CS/CE/Math/Physics

Tools

SYCL/DPC++/OpenCL/CUDA/HIP

Job description

About the Role

The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of oneDNN, a complex, cross-platform, open-source performance library that serves as the foundation for deep learning applications (github.com/uxlfoundation/oneDNN). Please Note: This is a low-level software engineering and hardware-acceleration role. It does not involve building, training, or tuning machine learning models. Instead, you will focus on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks (such as OpenVINO, TensorFlow, PyTorch, and ONNX Runtime) on Intel hardware.



Key Responsibilities

Kernel Development and Architecture


  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads

  • Design scalable JIT and codegen infrastructure for GPU kernel generation


Low-Level Optimization


  • Implement fusion and memory-traffic optimizations to maximize hardware utilization

  • Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)


Performance Modeling and Profiling


  • Build analytical and empirical performance models for kernel dispatch and tuning

  • Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths


Hardware and Software Co-Design


  • Co-design GPU primitives and kernel architectures for next-generation Intel GPUs

  • Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks


Infrastructure and Validation


  • Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads



Why Join Us


  • Massive Scale Work on a global, high-impact open-source library that scales AI performance across millions of devices worldwide

  • Cutting-Edge Hardware Get early access to and influence the software stack for Intel's roadmap of next-generation discrete GPUs

  • Expert Collaboration Work alongside industry-leading experts in GPU compilers, hardware architecture, and performance libraries



Total Rewards

Enjoy a competitive package including stock programs, quarterly bonuses, robust healthcare, and highly flexible hybrid/remote working options.



What We're Looking For


  • A strong ownership mindset - you take initiative on complex, ambiguous technical problems and drive them to resolution

  • A collaborative approach - you work effectively across hardware, compiler, and framework teams to align on shared technical goals

  • A performance-driven curiosity - you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems



Qualifications

Education


  • BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field


Core Language


  • 5+ years of professional software development experience with expert-level modern C++


Performance Optimizations


  • 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs


Hardware Architecture


  • Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization)


Preferred Qualifications


  • Math Libraries: Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)

  • Low-Level Tuning: Hands-on experience with GPU assembly-level tuning or compiler optimization

  • Parallel APIs: Familiarity with parallel programming APIs such as OpenMP or oneTBB

  • AI Workload Context: Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks



Job Type

Experienced Hire



Shift

Shift 1 (United States of America)



Primary Location

US, Oregon, Hillsboro



Additional Locations

US, California, Santa Clara



Posting Statement

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.



Work Model for this Role

This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site.



Benefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.



Annual Salary Range for jobs which could be performed in the US

$195,200.00-275,580.00 USD



Salary Range Details

The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Performance Software Engineer
Senior GPU Performance Software Engineer

Talanto • Hillsboro (OR), Northern (KY)

Hybrid
USD 195,000 - 276,000
Hybrid work model
Stock bonuses
Health insurance
+2
AI Performance Library Architect
AI Performance Library Architect

Intel • Hillsboro (OR)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Systems Software Engineer - Neuromorphic Computing
AI Systems Software Engineer - Neuromorphic Computing

Intel Corporation • Hillsboro (OR)

Hybrid
USD 129,000 - 211,000
GPU Performance Engineer
GPU Performance Engineer

Jobtailor • Folsom (CA)

On-site
USD 141,000 - 270,000
Competitive pay
Stock bonuses
Health benefits
+2
AI Performance Library Architect
AI Performance Library Architect

Intel Corporation • Hillsboro (OR)

Hybrid
USD 170,000 - 316,000
Competitive pay
Stock bonuses
Health and retirement benefits
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Sr. Director, Software Engineering
Sr. Director, Software Engineering

Intel Corporation • Northern (KY)

Hybrid
USD 272,000 - 383,000
Stock bonuses
Health benefits
Vacation
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Hillsboro (OR)

Hybrid
USD 195,200 - 361,200
Hybrid work model
Competitive compensation