Low-Latency AI Inference Engineer

Relha LLC

San Jose (CA)

Hybrid

USD 177,000 - 265,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments. You will optimize Large Language Models for inference across GPUs, edge devices, and data centers, focusing on throughput, latency, and accuracy.

Responsibilities include building inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, plus deploying scalable, low-latency solutions using Kubernetes and cloud platforms.

Qualifications

  • Proficiency in Python, C++, Rust, or Golang for high-performance AI workflows.
  • Hands-on experience with AI inference tools: vLLM, TensorRT, Llama.cpp, Ollama.
  • Infrastructure expertise with Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure).
  • Strong understanding of GPU/AI hardware, profiling, and optimization for accelerators like NVIDIA GPUs and TPUs.

Responsibilities

  • Build and maintain high-performance inference engines for scalable AI serving.
  • Optimize deployment for low latency across GPU-rich data centers and edge devices.
  • Profile and optimize models on NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and TPUs/LPUs.
  • Design auto-scaling architectures using Kubernetes for real-time and batch inference.
  • Establish observability for TTFT, tokens per second, and memory bandwidth against SLAs.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes
AWS
GCP
Azure
CUDA
Triton

Job description

F5 is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments. You will optimize Large Language Models for inference across GPUs, edge devices, and data centers, focusing on throughput, latency, and accuracy.

Responsibilities include building inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, plus deploying scalable, low-latency solutions using Kubernetes and cloud platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
AI Inference Engineer
AI Inference Engineer

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Low-Latency Inference Systems Engineer (Multi-GPU)
Low-Latency Inference Systems Engineer (Multi-GPU)

Digital Waffle • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior AI Inference Library Engineer - GPU-Optimized
Senior AI Inference Library Engineer - GPU-Optimized

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1