Get more replies from employers
Send a job-specific resume in minutes.
F5 is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments. You will optimize Large Language Models for inference across GPUs, edge devices, and data centers, focusing on throughput, latency, and accuracy.
Responsibilities include building inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, plus deploying scalable, low-latency solutions using Kubernetes and cloud platforms.
F5 is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments. You will optimize Large Language Models for inference across GPUs, edge devices, and data centers, focusing on throughput, latency, and accuracy.
Responsibilities include building inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, plus deploying scalable, low-latency solutions using Kubernetes and cloud platforms.