A leading AI company in San Mateo is seeking a candidate to optimize AI models for throughput, latency, and cost. The role involves working with advanced GPU performance techniques and a serving stack, focusing on delivering models at a faster and more cost-effective rate without compromising quality. Ideal candidates will have experience in GPU performance and be ready to join an on-site team. Strong familiarity with AI deployment systems is essential.
Qualifications
Experience in GPU performance and optimization.
Familiarity with AI model deployment and serving stack.
Strong background in parallel computing techniques.
Responsibilities
Optimize models for performance and cost.
Work with GPU performance and quantization.
Implement systems for observability and scaling.
Skills
CUDA/Triton kernels
Parallelism: FSDP/ZeRO
Quantization/PEFT
TensorRT-LLM/Triton Inference Server
Observability tools (Prom/Grafana/OpenTelemetry)
Job description
A leading AI company in San Mateo is seeking a candidate to optimize AI models for throughput, latency, and cost. The role involves working with advanced GPU performance techniques and a serving stack, focusing on delivering models at a faster and more cost-effective rate without compromising quality. Ideal candidates will have experience in GPU performance and be ready to join an on-site team. Strong familiarity with AI deployment systems is essential.