Senior Software Engineer, AI Inference Systems

NVIDIA

Italia

In loco

EUR 120.000 - 180.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Descrizione del lavoro

NVIDIA is seeking highly skilled software engineers to build AI inference systems that scale large models efficiently across multi-GPU, multi-node, and multi-cloud environments. You will architect and implement high-performance inference stacks, optimize GPU kernels and compilers, and collaborate across inference, compiler, scheduling, and performance teams.

You’ll drive industry benchmarks, contribute to vLLM, SGLang, and related tooling, and publish research that advances ML systems.

Competenze

  • Bachelor's degree in CS/CE/SE or equivalent with 7+ years of experience; or Master's with 5+ years; or PhD with top-tier publications.
  • Strong programming skills in Python and C/C++, Go or Rust a plus; solid CS fundamentals.
  • Knowledgeable about performance engineering in ML frameworks (PyTorch) and inference engines (vLLM, SGLang).
  • Familiar with GPU programming: CUDA, memory hierarchy, streams, NCCL; profiling/debug tools Nsight Systems/Compute.
  • Experience with containers and orchestration (Docker, Kubernetes, Slurm); Linux namespaces and cgroups.

Mansioni

  • Contribute features to vLLM and optimize inference framework with modern GPU features and parallelism.
  • Develop, optimize, and benchmark GPU kernels using fusion, autotuning, and memory/layout optimization.
  • Define and build inference benchmarking methodologies; contribute to MLPerf Inference submissions.
  • Architect scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.
  • Publish original research and translate ideas into NVIDIA software products.

Conoscenze

Python
C/C++
Go or Rust
Algorithms & data structures
Operating systems
Parallel programming
Distributed systems
Deep learning theories
Performance engineering
Python ML frameworks

Formazione

Bachelor's degree in CS/CE/SE
Master's degree in CS/CE/SE
PhD in ML Systems or related fields

Strumenti

CUDA
Nsight Systems/Compute
Docker
Kubernetes
Slurm

Descrizione del lavoro

We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You’ll collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated computing for AI.

What You’ll Be Doing
  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation.
  • Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization; build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization.
  • Define and build inference benchmarking methodologies and tools; contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite.
  • Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.
  • Conduct and publish original research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products.
What We Need To See
  • Bachelor’s degree (or equivalent expeience) in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE) with 7+ years of experience; alternatively, Master’s degree in CS/CE/SE with 5+ years of experience; or PhD degree with the thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.
  • Strong programming skills in Python and C/C++; experience with Go or Rust is a plus; solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, distributed systems, deep learning theories.
  • Knowledgeable and passionate about performance engineering in ML frameworks ("e.g., PyTorch") and inference engines ("e.g., vLLM" and "e.g., SGLang").
  • Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools ("e.g., Nsight Systems/Compute").
  • Experience with containers and orchestration (Docker, Kubernetes, Slurm); familiarity with Linux namespaces and cgroups.
  • Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.
Ways to stand out from the crowd
  • Experience building and optimizing LLM inference engines ("e.g., vLLM, SGLang").
  • Hands-on work with ML compilers and DSLs ("e.g., Triton", "e.g., TorchDynamo/Inductor", "e.g., MLIR/LLVM", "e.g., XLA"), GPU libraries ("e.g., CUTLASS") and features ("e.g., CUDA Graph", "e.g., Tensor Cores").
  • Experience contributing to containerization/virtualization technologies such as containerd/CRI-O/CRIU.
  • Experience with cloud platforms (AWS/GCP/Azure), infrastructure as code, CI/CD, and production observability.
  • Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.

At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential. Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior AI Compute Engineer
Senior AI Compute Engineer

NVIDIA Gruppe • Italia

In loco
EUR 66.000 - 114.000
Infiniband Network Engineer
Infiniband Network Engineer

NVIDIA ITALY S.R.L. • Italia

In loco
EUR 50.000 - 114.000
Senior Sales Account Manager, Smart Spaces and Local Government, Southern EMEA
Senior Sales Account Manager, Smart Spaces and Local Government, Southern EMEA

NVIDIA Corporation • Roma

Ibrido
EUR 114.000 - 259.000
Senior Platform Engineer
Senior Platform Engineer

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Hermes Corporate • Italia

In loco
EUR 70.000 - 100.000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

NVIDIA Gruppe • Italia

In loco
EUR 66.000 - 114.000
Senior AI Cloud Engineer
Senior AI Cloud Engineer

PLP Group • Milano

In loco
EUR 55.000 - 75.000
Learning Friday
Smart Working
Equity
AI Infrastructure Architect
AI Infrastructure Architect

Accenture Italia • Milano

In loco
EUR 36.000 - 66.000
Senior ML Ops Engineer — Remote, Scalable GPU Inference
Senior ML Ops Engineer — Remote, Scalable GPU Inference

Pragmatike • Italia

In loco
EUR 90.000 - 120.000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Skillvue • Milano

Remoto
EUR 70.000 - 90.000
Competitive compensation
Flexible work
Budget for conferences and training
+1