Cloud Inference Engineer

SupportFinity™

San Francisco (CA)

On-site

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A pioneering AI technology firm in San Francisco is seeking a founding member to optimize and serve models on Luminal Cloud. The role involves deploying models with advanced optimization techniques, conducting performance reviews, and enhancing scheduling processes. Ideal candidates are experienced in CUDA and GPU optimization, with hands-on knowledge of vLLM, SGLang, or TensorRT-LLM. A degree is not required, reflecting a modern approach to tech recruitment.

Qualifications

  • Proficiency in CUDA and GPU optimization techniques.
  • Experience with vLLM, SGLang, or TensorRT-LLM is preferred.
  • Understanding of KV caching and distributed compute is a plus.

Responsibilities

  • Deploy and tune models with optimizations like KV caching and batch processing.
  • Conduct model performance reviews to assess efficiency.
  • Improve processes for scheduling and autoscaling.

Skills

CUDA + GPU inference optimization
vLLM, SGLang, or TensorRT-LLM experience
KV caching
distributed compute
no degree required

Job description

Qualifications
  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc.
  • Distributed compute (with GPUs is a super plus)
  • No degree required
Company

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Role

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

Day To Day Responsibilities
  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Cloud Inference Engineer (Low-Latency AI Serving)
Founding Cloud Inference Engineer (Low-Latency AI Serving)

SupportFinity™ • San Francisco (CA)

On-site
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Compiler Engineer
Senior Compiler Engineer

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Compiler Engineer
Compiler Engineer

Slope • San Francisco (CA)

On-site
USD 180,000 - 300,000
Tech Lead Manager, Inference
Tech Lead Manager, Inference

Luma • Redwood City (CA)

On-site
USD 210,000 - 320,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits