LLM Inference Engineer — High-Performance AI Serving

F5

San Jose (CA)

On-site

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 is seeking an AI Inference Engineer to bridge high-performance model development with optimized deployment. You will optimize Large Language Models for inference across GPU-rich data centers and edge devices, focusing on throughput, latency, and accuracy.

You will work on hardware acceleration, scalable infrastructure, and performance monitoring to ensure enterprise-grade reliability and efficient AI capabilities.

Qualifications

  • Experience building high-performance AI workflows and inference systems.
  • Strong hands-on with inference tooling and model optimization.
  • Experience with Docker, Kubernetes and cloud platforms (AWS, GCP, Azure).

Responsibilities

  • Build and maintain high-performance AI inference engines at scale.
  • Profile and optimize models for NVIDIA GPUs, Apple Silicon, and AI accelerators.
  • Design auto-scaling architectures for online and batch inference using Kubernetes.
  • Establish observability for TTFT, tokens/second, and memory bandwidth against SLAs.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes
AWS
GCP
Azure

Job description

F5 is seeking an AI Inference Engineer to bridge high-performance model development with optimized deployment. You will optimize Large Language Models for inference across GPU-rich data centers and edge devices, focusing on throughput, latency, and accuracy.

You will work on hardware acceleration, scalable infrastructure, and performance monitoring to ensure enterprise-grade reliability and efficient AI capabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
AI Inference Engineer
AI Inference Engineer

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000