Senior AI Inference Systems Engineer

NVIDIA

Toronto

On-site

CAD 170,000 - 275,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking highly skilled software engineers to build AI inference systems that scale across multi-GPU, multi-node, and multi-cloud environments. You will optimize kernels, contribute to vLLM, and push the boundaries of accelerated computing for AI.

You’ll collaborate across inference, compiler, scheduling, and performance teams, conducting research and publishing results while delivering high-performance software for NVIDIA’s AI ecosystem.

Qualifications

  • Bachelor's degree in CS/CE/SE with 7+ years of experience, or Master’s with 5+ years, or PhD with top-tier ML Systems publications
  • Strong programming in Python and C/C++, Go or Rust a plus, solid CS fundamentals
  • Performance engineering for ML frameworks (e.g., PyTorch) and model serving systems (vLLM, SGLang)
  • CUDA GPU programming expertise and profiling tools (Nsight)
  • Experience with containers/orchestration (Docker, Kubernetes, Slurm) and Linux features

Responsibilities

  • Contribute features to vLLM and optimize inference framework with advanced parallelism and GPU features
  • Develop, optimize, and benchmark GPU kernels; build high-level DSLs and compiler infrastructure for peak hardware utilization
  • Define and build inference benchmarking methodologies; contribute to MLPerf Inference submissions
  • Architect scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds
  • Conduct and publish original research integrating ideas into NVIDIA software products

Skills

Python
C/C++
Go
Rust
Algorithms & Data Structures
Operating Systems
Computer Architecture
Parallel Programming
Distributed Systems
Deep Learning Theories

Education

Bachelor's degree in CS/CE/SE
Master's degree in CS/CE/SE
PhD in ML Systems / GPU Architecture / HPC

Tools

CUDA
Nsight Systems/Compute
Docker
Kubernetes
Slurm
Linux Namespaces

Job description

NVIDIA is seeking highly skilled software engineers to build AI inference systems that scale across multi-GPU, multi-node, and multi-cloud environments. You will optimize kernels, contribute to vLLM, and push the boundaries of accelerated computing for AI.

You’ll collaborate across inference, compiler, scheduling, and performance teams, conducting research and publishing results while delivering high-performance software for NVIDIA’s AI ecosystem.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

NVIDIA Corporation • Toronto

Hybrid
CAD 170,000 - 275,000
Equity
Benefits
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA Corporation • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA Gruppe • Toronto

Hybrid
CAD 170,000 - 275,000
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

NVIDIA Corporation • Toronto

Hybrid
CAD 135,000 - 220,000
Equity
Benefits
Hybrid work
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

NVIDIA Gruppe • Toronto

Hybrid
CAD 135,000 - 220,000
Equity
Benefits
Hybrid work model
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

NVIDIA • Toronto

On-site
CAD 135,000 - 220,000
Equity
Benefits
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

United States Digital Space LLC • Toronto

Hybrid
CAD 135,000 - 220,000
Equity
Benefits
Senior ML Infra Engineer: Scale GPU Training & Reliability
Senior ML Infra Engineer: Scale GPU Training & Reliability

Veeda AI • Toronto

On-site
CAD 130,000 - 185,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Toronto

On-site
CAD 150,000 - 210,000