Get more replies from employers
Send a job-specific resume in minutes.
F5 is seeking an AI Inference Engineer to bridge high‑performance model development with optimized deployment environments. The role focuses on optimizing Large Language Models for inference across GPU‑rich data centers to edge devices, emphasizing throughput, latency, and accuracy.
Responsibilities include building inference engines with vLLM, TensorRT, and Triton, profiling models for CUDA/TensorRT, CoreML, and AI accelerators, and designing scalable pipelines with Kubernetes.
F5 is seeking an AI Inference Engineer to bridge high‑performance model development with optimized deployment environments. The role focuses on optimizing Large Language Models for inference across GPU‑rich data centers to edge devices, emphasizing throughput, latency, and accuracy.
Responsibilities include building inference engines with vLLM, TensorRT, and Triton, profiling models for CUDA/TensorRT, CoreML, and AI accelerators, and designing scalable pipelines with Kubernetes.