Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
F5 Networks, Inc. is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments.
The role focuses on optimizing Large Language Models for inference across data centers and edge devices, prioritizing throughput, low latency, and accuracy. You will build scalable inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, and optimize hardware usage on NVIDIA GPUs, Apple Silicon, TPUs, and other accelerators.
At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.
Job Title: AI Inference Engineer
Role Objective The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy. This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.
Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale. Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.
Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs. Collaborate with hardware teams to maximize utilization and performance across various computational environments.
Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.
Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs). Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.
As an AI Inference Engineer at F5, success is measured by: Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems. Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement. Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals.
Equal Employment Opportunity It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting accommodations@f5.com.
Together, we're building a better digital world. Founded in 1996, F5 is a global leader in application delivery and security. Our premier platform helps customers secure and deliver every app, API, and piece of infrastructure across all environments. Backed by over three decades of expertise and 553 patents, our solutions protect against threats while ensuring fast, reliable digital experiences. With over 6,400 employees, we serve more than 23,000 customers in over 170 countries. To continue this work, we need people like you—the best minds in the industry. We're committed to a unique, human-first culture that encourages authenticity, prioritizes diversity and inclusion, and fosters the growth and success of our employees.