Senior AI Systems Engineer: LLM Inference & Optimization

Showcify

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Snowflake is seeking talented systems developers and researchers to advance LLM inference systems and optimization in the AI Research team. You will build high-performance, adaptive inference stacks spanning distributed serving, kernels, and runtime optimization to accelerate model deployment and experimentation.

You will collaborate with researchers and engineers to push latency, throughput, and cost boundaries while enabling rapid support for new models and hardware.

Qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
  • Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.

Responsibilities

  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
  • Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.
  • Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.
  • Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.
  • Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
  • Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
  • Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.
  • Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.
  • Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.
  • Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.
  • Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.

Job description

Snowflake is seeking talented systems developers and researchers to advance LLM inference systems and optimization in the AI Research team. You will build high-performance, adaptive inference stacks spanning distributed serving, kernels, and runtime optimization to accelerate model deployment and experimentation.

You will collaborate with researchers and engineers to push latency, throughput, and cost boundaries while enabling rapid support for new models and hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Senior/Staff System Research Engineer – LLM Inference Optimization
Senior/Staff System Research Engineer – LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
Senior/Staff System Research Engineer - LLM Inference Optimization
Senior/Staff System Research Engineer - LLM Inference Optimization

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Engineer - Production LLM Optimizer
Senior AI Inference Engineer - Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Senior AI Inference Engineer: High-Throughput LLMs
Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Systems Engineer & Research
LLM Systems Engineer & Research

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3