AI Engineer — Model Performance & Inference Optimizer
Pantera Capital
San Francisco (CA)
Hybrid
USD 120,000 - 180,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Competitive compensation and benefits
Supportive environment for personal growth
Dynamic and collaborative engineering team
Job summary
Pantera Capital is looking for a Model Performance Engineer in San Francisco, California to optimize model inference speed, cost, and reliability. You will build fine-tuning infrastructure that accelerates the AI team’s processes. The role covers optimizing serving frameworks and ensuring efficient GPU use. Ideal candidates should have deep experience with LLM serving frameworks, substantial Python skills, and production experience in model fine-tuning. Competitive compensation and a supportive work environment are offered.
Qualifications
You must have experience tuning LLM serving frameworks like vLLM or SGLang.
Hands-on experience in quantization techniques is essential.
Production fine-tuning experience in frameworks like LoRA is required.
Strong proficiency in Python for infrastructure and benchmarking is needed.
Ability to analyze benchmarking results for performance bottlenecks is crucial.
Responsibilities
Own inference performance optimizing models for speed and cost.
Develop fine-tuning pipelines to streamline deployment of AI models.
Skills
Deep experience with LLM serving frameworks
Hands-on quantization experience
Production fine-tuning experience
Strong Python
Comfort with GPU profiling and performance analysis
Job description
Pantera Capital is looking for a Model Performance Engineer in San Francisco, California to optimize model inference speed, cost, and reliability. You will build fine-tuning infrastructure that accelerates the AI team’s processes. The role covers optimizing serving frameworks and ensuring efficient GPU use. Ideal candidates should have deep experience with LLM serving frameworks, substantial Python skills, and production experience in model fine-tuning. Competitive compensation and a supportive work environment are offered.