AI Inference Engineer

F5

San Jose (CA)

On-site

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 is seeking an AI Inference Engineer to bridge high-performance model development with optimized deployment. You will optimize Large Language Models for inference across GPU-rich data centers and edge devices, focusing on throughput, latency, and accuracy.

You will work on hardware acceleration, scalable infrastructure, and performance monitoring to ensure enterprise-grade reliability and efficient AI capabilities.

Qualifications

  • Experience building high-performance AI workflows and inference systems.
  • Strong hands-on with inference tooling and model optimization.
  • Experience with Docker, Kubernetes and cloud platforms (AWS, GCP, Azure).

Responsibilities

  • Build and maintain high-performance AI inference engines at scale.
  • Profile and optimize models for NVIDIA GPUs, Apple Silicon, and AI accelerators.
  • Design auto-scaling architectures for online and batch inference using Kubernetes.
  • Establish observability for TTFT, tokens/second, and memory bandwidth against SLAs.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes
AWS
GCP
Azure

Job description

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

Job Description

The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy.

This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.

Key Responsibilities
High-Performance AI Serving
  • Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.
  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.
Hardware Acceleration and Optimization
  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.
  • Collaborate with hardware teams to maximize utilization and performance across various computational environments.
Inference Orchestration and Scalability
  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.
  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.
Performance Monitoring and Observability
  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).
  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.
Technical Requirements
Required Skills:
  • Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
AI Engineer
AI Engineer

Worky • San Jose (CA)

Hybrid
USD 172,000 - 257,000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000