Get more replies from employers
Send a job-specific resume in minutes.
F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLM inference from data centers to edge devices. You will fine-tune software stacks, hardware backends, and orchestration to maximize throughput while preserving model accuracy.
You will work with vLLM, TensorRT, Llama.cpp, and Ollama, plus Kubernetes for real-time and batch inference, ensuring reliable, scalable AI endpoints across environments.
F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLM inference from data centers to edge devices. You will fine-tune software stacks, hardware backends, and orchestration to maximize throughput while preserving model accuracy.
You will work with vLLM, TensorRT, Llama.cpp, and Ollama, plus Kubernetes for real-time and batch inference, ensuring reliable, scalable AI endpoints across environments.