Senior AI Systems Research Engineer - LLM Inference

Snowflake

Bellevue (WA)

On-site

USD 236,000 - 310,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Bonus and equity plan

Job summary

Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product teams to push the frontier of AI systems.

The role emphasizes experimentation, production deployment, and sharing innovations through blogs and conferences, with a base salary range disclosed for U.S.

Qualifications

  • Bachelor’s degree in CS, EE, or related field; advanced degrees preferred.
  • 5+ years in LLM inference systems, distributed AI, GPU systems, or HPC.
  • Strong understanding of modern LLM inference architectures and tradeoffs.
  • Hands-on with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM.
  • Experience designing or optimizing inference runtimes (scheduling, batching, KV-cache).
  • Strong GPU expertise (CUDA, Triton) and performance libraries (CUTLASS, cuBLAS, cuDNN).
  • Experience profiling with Nsight or equivalent, independent problem solving, cross-team collaboration.

Responsibilities

  • Design and develop high-performance LLM inference systems across distributed serving and GPUs.
  • Develop techniques to improve latency, throughput, memory efficiency, and cost.
  • Explore adaptive parallelism, speculative decoding, disaggregation, and scheduling.
  • Create adaptive inference systems for new models, hardware, and workloads.
  • Apply AI-native approaches to profiling, configuration search, and performance tuning.
  • Identify high-impact performance problems and drive solutions from research to production.
  • Design distributed inference strategies across GPUs and nodes, including various parallelisms.
  • Develop efficient multi-model serving, dynamic resource management, and model swapping.
  • Analyze and optimize GPU kernels for attention, MoE, and other components.
  • Explore model-system co-design for more efficient inference and performance benchmarking.
  • Profile end-to-end workloads to identify bottlenecks across compute, memory, and networking.
  • Collaborate with researchers, infra, and product teams to deploy innovations.
  • Open-source and publish innovations via blogs and top ML/system conferences.

Skills

LLM inference systems
distributed AI systems
GPU systems
high-performance computing

Education

Bachelor's degree
Master's degree
PhD

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product teams to push the frontier of AI systems.

The role emphasizes experimentation, production deployment, and sharing innovations through blogs and conferences, with a base salary range disclosed for U.S.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Staff Research Scientist, Agentic AI & LLMs
Staff Research Scientist, Agentic AI & LLMs

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 340,000
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Senior AI Researcher: Build Production-Grade LLM Systems
Senior AI Researcher: Build Production-Grade LLM Systems

Innovaccer • San Francisco (CA)

On-site
USD 190,000 - 275,000
PTO 20 days
Parental leave
Rewards & Recognition
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000