High-Performance AI Inference Engineer

F5 Networks, Inc. 

San Jose (CA)

Hybrid

USD 177,000 - 265,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLMs for low-latency, scalable inference across data centers and edge devices. You will build inference engines with vLLM, TensorRT, Triton, and orchestrate deployments with Kubernetes to handle real-time and batch workloads.

This role emphasizes hardware acceleration, MLOps practices, and observability metrics like TTFT and tokens per second, with a base pay range noted in the job description.

Qualifications

  • Programming Languages: Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Inference Tools: vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Infrastructure: Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
  • Hardware Optimization: profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs.
  • Experience deploying LLMs with techniques like Speculative Decoding or PagedAttention.
  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).
  • MLOps or SRE background focused on reliable high-performance AI endpoints.

Responsibilities

  • High-Performance AI Serving: Build and maintain robust inference engines using vLLM, TGI, and NVIDIA Triton.
  • Handle deployment optimizations to deliver low-latency AI serving solutions.
  • Hardware Acceleration and Optimization: Profile and optimize models for GPUs, CoreML, and AI accelerators.
  • Inference Orchestration and Scalability: Design auto-scaling architectures with Kubernetes for real-time and batch pipelines.
  • Performance Monitoring and Observability: Establish observability for TTFT, tokens/sec, and memory bandwidth against SLAs.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes
AWS
GCP
Azure

Job description

F5 Networks, Inc. is seeking an AI Inference Engineer to optimize LLMs for low-latency, scalable inference across data centers and edge devices. You will build inference engines with vLLM, TensorRT, Triton, and orchestrate deployments with Kubernetes to handle real-time and batch workloads.

This role emphasizes hardware acceleration, MLOps practices, and observability metrics like TTFT and tokens per second, with a base pay range noted in the job description.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
AI Inference Engineer
AI Inference Engineer

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
AI Inference Engineer
AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000