Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake

Menlo Park (CA)

On-site

USD 236,000 - 310,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
Paid holidays

Job summary

Snowflake is seeking talented systems developers and researchers to advance LLM inference systems and optimization. You will design high-performance, intelligent inference stacks spanning distributed serving, runtime systems, and GPU kernels, collaborating with model researchers and product teams.

The role emphasizes AI-native engineering to accelerate both model performance and inference development, with opportunities to publish innovations and contribute to production-scale AI.

Qualifications

  • Bachelors in CS/EE; advanced degree preferred.
  • 5+ years in LLM inference or distributed AI.
  • Experience with LLM architectures and serving tradeoffs.
  • Hands-on with LLM inference frameworks (e.g., vLLM, TensorRT-LLM).
  • Familiarity with GPU programming (CUDA, Triton).

Responsibilities

  • Design high-performance LLM inference systems across devices.
  • Develop techniques to reduce latency and improve throughput.
  • Explore adaptive parallelism, decoding, disaggregation, batching, KV-cache.
  • Build adaptive inference systems for new models and hardware.
  • Apply AI-native engineering to profiling, configuration search, and optimization.
  • Identify bottlenecks and drive solutions from research to production.
  • Design distributed inference across GPUs and nodes.
  • Develop multi-model serving, dynamic resource management, model loading.
  • Profile GPUs and operators for attention, MoE, etc.
  • Collaborate with researchers, infra, product; publish innovations.
  • Open-source contributions and blogs.

Skills

LLM inference
Distributed AI
GPU systems
Performance optimization
System scalability
AI-native engineering

Education

Bachelor's degree in CS/EE
Master's or PhD preferred

Tools

CUDA
Triton
TensorRT-LLM
vLLM
SGLang
cuDNN
cuBLAS
CUTLASS
Nsight

Job description

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.

Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.

Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.

Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development.

This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI.

Responsibilities
  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
  • Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.
  • Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.
  • Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.
  • Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
  • Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
  • Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.
  • Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.
  • Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.
  • Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.
  • Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.
Requirements
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.
  • 5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
  • Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.
  • Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.
  • Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.
  • Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.
  • Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.
  • Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.
  • Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.
  • Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.
  • Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.
  • Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

  • The estimated base salary range for this role is $236,000 - $309,750.
  • Additionally, this role is eligible to participate in Snowflake’s bonus and equity plan.

The successful candidate’s starting salary will be determined based on permissible, non-discriminatory factors such as skills, experience, and geographic location. This role is also eligible for a competitive benefits package that includes: medical, dental, vision, life, and disability insurance; 401(k) retirement plan; flexible spending & health savings account; at least 12 paid holidays; paid time off; parental leave; employee assistance program; and other company benefits.

To comply with pay transparency requirements and other statutes, you can notify us if you believe that a job posting is not compliant by completing this form.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Intermediate/ Senior Software Engineer - Cortex LLM Training Platform
Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

Snowflake • Bellevue (WA)

On-site
USD 200,000 - 288,000
Medical, dental, vision
401(k) retirement plan
Paid holidays and time off
Staff Research Scientist, AI Agents & LLMs
Staff Research Scientist, AI Agents & LLMs

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 340,000
Staff/Principal AI Software Engineer - Snowflake CoWork
Staff/Principal AI Software Engineer - Snowflake CoWork

Snowflake • Menlo Park (CA)

On-site
USD 264,000 - 380,000
Medical, Dental, Vision insurance
401(k) retirement plan
Equity plan
+1
Principal / Staff Applied Research Scientist
Principal / Staff Applied Research Scientist

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 340,000
Medical, dental, and vision insurance
401(k) retirement plan
Paid time off and parental leave
Staff Research Scientist, AI Agents & LLMs
Staff Research Scientist, AI Agents & LLMs

Snowflake • Bellevue (WA)

On-site
USD 130,000 - 200,000
Senior Software Engineer - Cloud Efficiency
Senior Software Engineer - Cloud Efficiency

Snowflake • Bellevue (WA)

On-site
USD 200,000 - 288,000
Medical, dental, vision
401(k) retirement plan
Paid holidays
Senior AI/ML Architect, Applied Field Engineering
Senior AI/ML Architect, Applied Field Engineering

Snowflake • Washington

On-site
USD 165,000 - 217,000
Parental leave
Paid time off
Paid holidays
+2
Senior Forward Deployed Engineer, Applied AI
Senior Forward Deployed Engineer, Applied AI

Snowflake • Menlo Park (CA)

On-site
USD 200,000 - 263,000