Senior LLM Inference Systems Engineer

Snowflake

Menlo Park (CA)

On-site

USD 236,000 - 310,000

Full time

23 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
Paid holidays

Job summary

Snowflake is seeking talented systems developers and researchers to advance LLM inference systems and optimization. You will design high-performance, intelligent inference stacks spanning distributed serving, runtime systems, and GPU kernels, collaborating with model researchers and product teams.

The role emphasizes AI-native engineering to accelerate both model performance and inference development, with opportunities to publish innovations and contribute to production-scale AI.

Qualifications

  • Bachelors in CS/EE; advanced degree preferred.
  • 5+ years in LLM inference or distributed AI.
  • Experience with LLM architectures and serving tradeoffs.
  • Hands-on with LLM inference frameworks (e.g., vLLM, TensorRT-LLM).
  • Familiarity with GPU programming (CUDA, Triton).

Responsibilities

  • Design high-performance LLM inference systems across devices.
  • Develop techniques to reduce latency and improve throughput.
  • Explore adaptive parallelism, decoding, disaggregation, batching, KV-cache.
  • Build adaptive inference systems for new models and hardware.
  • Apply AI-native engineering to profiling, configuration search, and optimization.
  • Identify bottlenecks and drive solutions from research to production.
  • Design distributed inference across GPUs and nodes.
  • Develop multi-model serving, dynamic resource management, model loading.
  • Profile GPUs and operators for attention, MoE, etc.
  • Collaborate with researchers, infra, product; publish innovations.
  • Open-source contributions and blogs.

Skills

LLM inference
Distributed AI
GPU systems
Performance optimization
System scalability
AI-native engineering

Education

Bachelor's degree in CS/EE
Master's or PhD preferred

Tools

CUDA
Triton
TensorRT-LLM
vLLM
SGLang
cuDNN
cuBLAS
CUTLASS
Nsight

Job description

Snowflake is seeking talented systems developers and researchers to advance LLM inference systems and optimization. You will design high-performance, intelligent inference stacks spanning distributed serving, runtime systems, and GPU kernels, collaborating with model researchers and product teams.

The role emphasizes AI-native engineering to accelerate both model performance and inference development, with opportunities to publish innovations and contribute to production-scale AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
Senior AI Systems Research Engineer - LLM Inference
Senior AI Systems Research Engineer - LLM Inference

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000