LLM Inference & Systems Optimization Engineer

Snowflake

Bellevue (KY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Snowflake is seeking talented systems developers and researchers to join the AI Research team and advance LLM inference systems and optimization. You will work across distributed serving, runtime systems, GPU kernels, and model-system co-design to push latency, throughput, scalability, and cost improvements.

This role emphasizes AI-native engineering and collaboration with model researchers, infrastructure, and product teams to deploy research into production.

Qualifications

  • 5+ years in LLM inference systems, distributed AI, GPU systems, or HPC.
  • Strong understanding of modern LLM inference architectures and tradeoffs.
  • Hands-on experience with LLM inference and serving frameworks like vLLM, SGLang, TensorRT-LLM.
  • Experience designing, extending, or optimizing inference runtimes (scheduling, batching, KV-cache).
  • Strong understanding of GPU architectures and CUDA/Triton.
  • Experience with performance libraries such as CUTLASS, cuBLAS, cuDNN.
  • Experience profiling with Nsight tools.
  • Independent problem solving and cross-team collaboration.
  • AI-native engineering approaches to accelerate development and deployment.

Responsibilities

  • Design and develop high-performance LLM inference systems across distributed serving and GPU kernels.
  • Develop techniques to improve latency, generation speed, throughput, memory efficiency, scalability, and cost.
  • Explore adaptive parallelism, parallel decoding, batching, and KV-cache optimization.
  • Create adaptive intelligent inference systems for new model architectures and hardware.
  • Apply AI-driven approaches to profiling, configuration search, and performance tuning.
  • Identify bottlenecks and prototype solutions from research to production.
  • Design distributed inference strategies across GPUs and nodes.
  • Multi-model serving, dynamic resource management, and workload-aware scheduling.

Skills

LLM inference systems
Distributed AI systems
GPU systems
High-performance computing
Independent problem solver
Communication skills

Education

Bachelor’s in Computer Science or Electrical Engineering
Master’s degree or PhD preferred

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
CUTLASS
cuBLAS
cuDNN
Nsight Systems
Nsight Compute

Job description

Snowflake is seeking talented systems developers and researchers to join the AI Research team and advance LLM inference systems and optimization. You will work across distributed serving, runtime systems, GPU kernels, and model-system co-design to push latency, throughput, scalability, and cost improvements.

This role emphasizes AI-native engineering and collaboration with model researchers, infrastructure, and product teams to deploy research into production.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior AI Systems Research Engineer - LLM Inference
Senior AI Systems Research Engineer - LLM Inference

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
AI Systems Research and Development Engineer – LLM Inference Systems & Optimization
AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

Snowflake • Bellevue (KY)

On-site
USD 180,000 - 240,000
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000