Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
F5 Networks, Inc. is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments.
The role focuses on optimizing Large Language Models for inference across data centers and edge devices, prioritizing throughput, low latency, and accuracy. You will build scalable inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, and optimize hardware usage on NVIDIA GPUs, Apple Silicon, TPUs, and other accelerators.
F5 Networks, Inc. is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments.
The role focuses on optimizing Large Language Models for inference across data centers and edge devices, prioritizing throughput, low latency, and accuracy. You will build scalable inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, and optimize hardware usage on NVIDIA GPUs, Apple Silicon, TPUs, and other accelerators.