AI Engineer — Model Performance & Inference Optimizer

Pantera Capital

San Francisco (CA)

Hybrid

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation and benefits
Supportive environment for personal growth
Dynamic and collaborative engineering team

Job summary

Pantera Capital is looking for a Model Performance Engineer in San Francisco, California to optimize model inference speed, cost, and reliability. You will build fine-tuning infrastructure that accelerates the AI team’s processes. The role covers optimizing serving frameworks and ensuring efficient GPU use. Ideal candidates should have deep experience with LLM serving frameworks, substantial Python skills, and production experience in model fine-tuning. Competitive compensation and a supportive work environment are offered.

Qualifications

  • You must have experience tuning LLM serving frameworks like vLLM or SGLang.
  • Hands-on experience in quantization techniques is essential.
  • Production fine-tuning experience in frameworks like LoRA is required.
  • Strong proficiency in Python for infrastructure and benchmarking is needed.
  • Ability to analyze benchmarking results for performance bottlenecks is crucial.

Responsibilities

  • Own inference performance optimizing models for speed and cost.
  • Develop fine-tuning pipelines to streamline deployment of AI models.

Skills

Deep experience with LLM serving frameworks
Hands-on quantization experience
Production fine-tuning experience
Strong Python
Comfort with GPU profiling and performance analysis

Job description

Pantera Capital is looking for a Model Performance Engineer in San Francisco, California to optimize model inference speed, cost, and reliability. You will build fine-tuning infrastructure that accelerates the AI team’s processes. The role covers optimizing serving frameworks and ensuring efficient GPU use. Ideal candidates should have deep experience with LLM serving frameworks, substantial Python skills, and production experience in model fine-tuning. Competitive compensation and a supportive work environment are offered.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Model Performance Engineer: AI Inference & Fine-Tuning
Model Performance Engineer: AI Inference & Fine-Tuning

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive compensation and benefits
Supportive environment for innovation and growth
Senior AI Systems Performance Engineer: Drive SOTA Inference
Senior AI Systems Performance Engineer: Drive SOTA Inference

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Lead AI Inference Performance Architect
Lead AI Inference Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
AI Infrastructure Performance Engineer
AI Infrastructure Performance Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance