High-Performance AI Inference Engineer

F5 Networks, Inc. 

San Jose (CA)

Hybrid

USD 177,000 - 265,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLM inference from data centers to edge devices. You will fine-tune software stacks, hardware backends, and orchestration to maximize throughput while preserving model accuracy.

You will work with vLLM, TensorRT, Llama.cpp, and Ollama, plus Kubernetes for real-time and batch inference, ensuring reliable, scalable AI endpoints across environments.

Qualifications

  • Proficiency in Python, C++, Rust, or Golang for high-performance AI workflows.

Responsibilities

  • High-Performance AI Serving: Build and maintain inference engines with vLLM, TGI, and NVIDIA Triton for scalable performance.

Skills

Python
C++
Rust
Golang
vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes
AWS
GCP
Azure
CUDA
Triton
MLOps
SRE

Tools

TensorRT
Llama.cpp
Ollama

Job description

F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLM inference from data centers to edge devices. You will fine-tune software stacks, hardware backends, and orchestration to maximize throughput while preserving model accuracy.

You will work with vLLM, TensorRT, Llama.cpp, and Ollama, plus Kubernetes for real-time and batch inference, ensuring reliable, scalable AI endpoints across environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Intel • Hillsboro (OR)

Hybrid
USD 195,200 - 361,200
Hybrid work model
Competitive compensation
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Benefits package
Equity
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 140,000 - 230,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1