Founding Cloud Inference Engineer (Low-Latency AI Serving)

SupportFinity™

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A pioneering AI technology firm in San Francisco is seeking a founding member to optimize and serve models on Luminal Cloud. The role involves deploying models with advanced optimization techniques, conducting performance reviews, and enhancing scheduling processes. Ideal candidates are experienced in CUDA and GPU optimization, with hands-on knowledge of vLLM, SGLang, or TensorRT-LLM. A degree is not required, reflecting a modern approach to tech recruitment.

Qualifications

  • Proficiency in CUDA and GPU optimization techniques.
  • Experience with vLLM, SGLang, or TensorRT-LLM is preferred.
  • Understanding of KV caching and distributed compute is a plus.

Responsibilities

  • Deploy and tune models with optimizations like KV caching and batch processing.
  • Conduct model performance reviews to assess efficiency.
  • Improve processes for scheduling and autoscaling.

Skills

CUDA + GPU inference optimization
vLLM, SGLang, or TensorRT-LLM experience
KV caching
distributed compute
no degree required

Job description

A pioneering AI technology firm in San Francisco is seeking a founding member to optimize and serve models on Luminal Cloud. The role involves deploying models with advanced optimization techniques, conducting performance reviews, and enhancing scheduling processes. Ideal candidates are experienced in CUDA and GPU optimization, with hands-on knowledge of vLLM, SGLang, or TensorRT-LLM. A degree is not required, reflecting a modern approach to tech recruitment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Inference Engineer
Cloud Inference Engineer

SupportFinity™ • San Francisco (CA)

On-site
Founding Compiler Engineer - AI Models, CUDA & GPU
Founding Compiler Engineer - AI Models, CUDA & GPU

Slope • San Francisco (CA)

On-site
USD 180,000 - 300,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Senior Cloud AI LLM Serving Engineer
Senior Cloud AI LLM Serving Engineer

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus program
Potential RSU grants
Comprehensive benefits package
Founding Compiler Engineer - CUDA & Rust for AI Production (SF)
Founding Compiler Engineer - CUDA & Rust for AI Production (SF)

Luminal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Compiler Engineer
Senior Compiler Engineer

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Tech Lead Manager, Inference
Tech Lead Manager, Inference

Luma • Redwood City (CA)

On-site
USD 210,000 - 320,000
Staff ML Engineer: Efficient ML & Low-Latency AI
Staff ML Engineer: Efficient ML & Low-Latency AI

Embedding VC • San Francisco (CA)

On-site
USD 100,000 - 150,000