Founding Cloud Inference Engineer (Low-Latency AI Serving)
SupportFinity™
San Francisco (CA)
On-site
USD 180,000 - 320,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A pioneering AI technology firm in San Francisco is seeking a founding member to optimize and serve models on Luminal Cloud. The role involves deploying models with advanced optimization techniques, conducting performance reviews, and enhancing scheduling processes. Ideal candidates are experienced in CUDA and GPU optimization, with hands-on knowledge of vLLM, SGLang, or TensorRT-LLM. A degree is not required, reflecting a modern approach to tech recruitment.
Qualifications
Proficiency in CUDA and GPU optimization techniques.
Experience with vLLM, SGLang, or TensorRT-LLM is preferred.
Understanding of KV caching and distributed compute is a plus.
Responsibilities
Deploy and tune models with optimizations like KV caching and batch processing.
Conduct model performance reviews to assess efficiency.
Improve processes for scheduling and autoscaling.
Skills
CUDA + GPU inference optimization
vLLM, SGLang, or TensorRT-LLM experience
KV caching
distributed compute
no degree required
Job description
A pioneering AI technology firm in San Francisco is seeking a founding member to optimize and serve models on Luminal Cloud. The role involves deploying models with advanced optimization techniques, conducting performance reviews, and enhancing scheduling processes. Ideal candidates are experienced in CUDA and GPU optimization, with hands-on knowledge of vLLM, SGLang, or TensorRT-LLM. A degree is not required, reflecting a modern approach to tech recruitment.