Senior Solutions Architect – Large Scale Neural Networks Inference

NVIDIA

Poland

On-site

PLN 375,000 - 650,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

NVIDIA seeks a Senior Solutions Architect with deep expertise in large-scale neural network inference to lead collaboration with frontier AI labs and enterprises deploying AI at scale across EMEA. You will define the technical direction for AI inference, identify bottlenecks, and drive scalable, high-impact solutions.

By aligning key NVIDIA and customer teams, you will influence strategic decisions and guide deployment across Poland and beyond, shaping the next generation of inference stack,

Qualifications

  • MS or PhD in Computer Science, Engineering, or equivalent.
  • 8+ years in AI/ML infrastructure with production deployments.
  • Deep understanding of transformer inference acceleration.
  • Knowledge of GPU memory hierarchies and low-latency networking.
  • Proven ability to lead technical initiatives.
  • Excellent communication with researchers, engineers, and executives.

Responsibilities

  • Lead the inference strategy for an EMEA customer portfolio from POC to production-scale deployments.
  • Identify latency, efficiency, cost per token, and memory bottlenecks in deployments.
  • Architect and optimize high-performance inference pipelines using Dynamo, TensorRT-LLM, vLLM, and more.
  • Translate insights into product feedback for roadmap of NVIDIA stack.

Skills

MS/PhD in CS/Engineering
8+ years AI/ML infra
Transformer inference acceleration
GPU memory hierarchies & low-latency
lead technical initiatives
strong cross-group communication

Education

MS or PhD in Computer Science, Engineering, or equivalent

Tools

TensorRT-LLM
NVIDIA Dynamo
Kubernetes

Job description

We are seeking a Senior Solutions Architect with deep expertise in large-scale neural network inference and a proven ability to lead technical collaboration with frontier AI labs and enterprises deploying AI at scale across EMEA. In this role, you will define the technical direction for AI inference across EMEA by identifying critical bottlenecks and driving the development of scalable, high-impact solutions. By aligning key team members within NVIDIA and customer organizations, you will influence strategic technology decisions to develop the deployment of next-generation AI inference at scale.

What You Will Be Doing
  • Lead the inference strategy for a portfolio of EMEA AI Natives customers, guiding engagements from initial proof of concept to production-scale deployments.
  • Identify inference challenges across customer deployments including latency, efficiency, cost per token, memory utilization, and low-latency networking.
  • Architect and optimize high-performance inference pipelines using NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other inference backends, improving GPU utilization and AI cluster efficiency.
  • Translate customer insights and deployment patterns into actionable product feedback that develops the roadmap for NVIDIA stack such as Dynamo, TensorRT-LLM, and NIM.
What We Need To See
  • MS or PhD in Computer Science, Engineering, or equivalent experience in the field.
  • 8+ years in AI/ML infrastructure, with deep expertise in LLM/VLM inference optimization and production deployment at scale.
  • Deep understanding of transformer inference acceleration: quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models.
  • Understanding of GPU memory hierarchies and low-latency networking along with their influence on inference performance.
  • Proven track record to lead technical initiatives.
  • Excellent communication skills, effective with research scientists, infrastructure engineers, and executive team members.
Ways To Stand Out From The Crowd
  • Experience with NVIDIA's inference stack, including TensorRT-LLM, Triton Inference Server, NIM, and NVIDIA Dynamo.
  • Experience with GPU orchestration on Kubernetes.
  • You have operated inference at scale inside a frontier AI lab or hyperscale's inference team.
  • Contributions to open-source inference projects such as vLLM, SGLang, KServe, or NVIDIA Dynamo.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer www.nvidiabenefits.com/

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5. , , JR2024432

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior Solutions Architect – Large Scale AI Training
Senior Solutions Architect – Large Scale AI Training

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Polska

On-site
PLN 293,000 - 507,000
Senior AI Inference Architect – Large-Scale Neural Nets
Senior AI Inference Architect – Large-Scale Neural Nets

NVIDIA • Poland

On-site
PLN 375,000 - 650,000
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Poland

On-site
PLN 221,000 - 384,000
Senior Solutions Architect, Customer Success and Partnership
Senior Solutions Architect, Customer Success and Partnership

NVIDIA • Poland

On-site
PLN 293,000 - 650,000
Senior MLOps Engineer - DSX Enablement
Senior MLOps Engineer - DSX Enablement

NVIDIA • Poland

On-site
PLN 293,000 - 650,000
Competitive salary
Generous benefits
Manager, Deep Learning Algorithms
Manager, Deep Learning Algorithms

NVIDIA • Poland

On-site
PLN 345,000 - 702,000
Senior Solutions Architect - Multimodal AI
Senior Solutions Architect - Multimodal AI

NVIDIA • Poland

On-site
PLN 293,000 - 507,000
Senior Deep Learning Engineer, Accuracy Evaluation
Senior Deep Learning Engineer, Accuracy Evaluation

NVIDIA • Warszawa

On-site
PLN 390,000 - 650,000