Senior LLM Inference Systems Engineer

Snowflake

Menlo Park (CA)

On-site

USD 236,000 - 310,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
Bonus & equity plan
401(k) retirement plan
12 paid holidays

Job summary

Snowflake is seeking talented systems developers and researchers to join the AI Research team in Menlo Park. You will advance the state of the art in LLM inference systems and optimization, spanning distributed serving, runtime systems, and GPU kernels.

The role focuses on building high-performance, adaptive inference platforms that scale with evolving models and hardware, while collaborating with model researchers and product teams to deploy innovations in production.

Qualifications

  • 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
  • Strong understanding of modern LLM inference architectures and performance tradeoffs for large-scale models.
  • Hands-on experience with LLM inference and serving frameworks such as vLLM, SGLang, or TensorRT-LLM.

Responsibilities

  • Design and develop high-performance LLM inference systems across distributed serving, runtime systems, GPU execution, and kernels.
  • Develop techniques to reduce latency and increase throughput and memory efficiency.
  • Explore adaptive parallelism, speculative decoding, disaggregated inference, scheduling and batching, and KV-cache optimization.

Skills

LLM inference systems
Distributed AI systems
GPU systems
High-performance computing

Education

Bachelor's degree in CS/EE or related field
Master's degree or PhD preferred

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Snowflake is seeking talented systems developers and researchers to join the AI Research team in Menlo Park. You will advance the state of the art in LLM inference systems and optimization, spanning distributed serving, runtime systems, and GPU kernels.

The role focuses on building high-performance, adaptive inference platforms that scale with evolving models and hardware, while collaborating with model researchers and product teams to deploy innovations in production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Inference Engineer — High-Performance Rust Systems
LLM Inference Engineer — High-Performance Rust Systems

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
LLM Systems Engineer & Research
LLM Systems Engineer & Research

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3