Senior AI Inference Architect: Ultra-Scale MoE on GPU Clusters

NVIDIA

Poland

On-site

PLN 293,000 - 507,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA in Poland is seeking a Senior Solutions Architect with deep experience in large-scale production AI inference. You will lead technical direction for scalable, high-performance inference across multi-node GPU deployments and collaborate with EMEA AI Natives, infrastructure providers, and enterprises.

You will optimize inference pipelines, shaping performance, efficiency, and reliability across quantization, MoE workloads, and advanced architectures while engaging ML engineers, researchers,

Qualifications

  • MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience.
  • 5+ years of experience in Neural Networks inference optimization.
  • Solid understanding of transformers inference optimization: quantization, disaggregated inference, speculative decoding, continuous batching, KV cache optimization.
  • Practical experience in MoE inference at scale: expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale.
  • Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level.

Responsibilities

  • Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters.
  • Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs.
  • Improve inference efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments.
  • Collaborate with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) to accelerate customer success.
  • Animate the AI inference developer’s community across EMEA through technical workshops, hackathons, and reference architectures.

Skills

NN inference optimization
MoE inference at scale
Transformers inference optimization
Engage with ML engineers/researchers

Education

MS or PhD in Computer Science/Engineering or equivalent

Tools

NVIDIA Dynamo
NIXL
TensorRT-LLM

Job description

NVIDIA in Poland is seeking a Senior Solutions Architect with deep experience in large-scale production AI inference. You will lead technical direction for scalable, high-performance inference across multi-node GPU deployments and collaborate with EMEA AI Natives, infrastructure providers, and enterprises.

You will optimize inference pipelines, shaping performance, efficiency, and reliability across quantization, MoE workloads, and advanced architectures while engaging ML engineers, researchers,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Architect – Large-Scale Neural Nets
Senior AI Inference Architect – Large-Scale Neural Nets

NVIDIA • Poland

On-site
PLN 375,000 - 650,000
Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior Solutions Architect – Large Scale Neural Networks Inference
Senior Solutions Architect – Large Scale Neural Networks Inference

NVIDIA • Poland

On-site
PLN 375,000 - 650,000
Senior Multimodal AI Solutions Architect
Senior Multimodal AI Solutions Architect

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior Solutions Architect – Large Scale AI Training
Senior Solutions Architect – Large Scale AI Training

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

NVIDIA • Polska

On-site
PLN 293,000 - 507,000
Senior AI Infra & Data Center Architect
Senior AI Infra & Data Center Architect

NVIDIA • Poland

On-site
PLN 293,000 - 650,000
NVIDIA AI Solutions Architect — Multimodal Systems Lead
NVIDIA AI Solutions Architect — Multimodal Systems Lead

Ernst & Young Advisory Services Sdn Bhd • Warszawa

On-site
PLN 100,000 - 130,000
Continuous learning
Transformative leadership
Diverse and inclusive culture
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Polska

On-site
PLN 293,000 - 507,000
Senior AI Solutions Architect - Multimodal & Real-Time
Senior AI Solutions Architect - Multimodal & Real-Time

EY • Wrocław

Hybrid
PLN 80,000 - 120,000
Continuous learning
Flexible work environment
Diverse and inclusive culture
+1