AI Inference Engineer — High-Performance, Low-Latency ML

F5

Dublin

On-site

EUR 90,000 - 150,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

F5 is seeking an AI Inference Engineer to bridge high‑performance model development with optimized deployment environments. The role focuses on optimizing Large Language Models for inference across GPU‑rich data centers to edge devices, emphasizing throughput, latency, and accuracy.

Responsibilities include building inference engines with vLLM, TensorRT, and Triton, profiling models for CUDA/TensorRT, CoreML, and AI accelerators, and designing scalable pipelines with Kubernetes.

Qualifications

  • Proficiency in Python, C++, Rust or Golang for high‑performance AI workflows.
  • Hands‑on experience with inference tools like vLLM, TensorRT, Llama.cpp, Ollama.
  • Familiar with infrastructure: Docker, Kubernetes, cloud platforms (AWS, GCP, Azure).
  • Knowledge of GPU/TPU hardware optimization and profiling techniques.

Responsibilities

  • Build and maintain high‑performance AI serving engines for diverse deployments.
  • Profile and optimize models for GPUs, Apple Silicon, and AI accelerators.
  • Design auto‑scaling inference pipelines using Kubernetes and efficient routing.
  • Establish observability for TTFT, tokens/second, and memory bandwidth, with load testing.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama
Docker
Kubernetes

Job description

F5 is seeking an AI Inference Engineer to bridge high‑performance model development with optimized deployment environments. The role focuses on optimizing Large Language Models for inference across GPU‑rich data centers to edge devices, emphasizing throughput, latency, and accuracy.

Responsibilities include building inference engines with vLLM, TensorRT, and Triton, profiling models for CUDA/TensorRT, CoreML, and AI accelerators, and designing scalable pipelines with Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

F5 • Dublin

On-site
EUR 90,000 - 150,000
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
Senior AI Inference Engineer - High-Throughput LLM Serving
Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Senior Site Reliability Engineer, AI Inference
Senior Site Reliability Engineer, AI Inference

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Senior Engineer - AI, Inference
Senior Engineer - AI, Inference

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Inference Performance Engineer
Inference Performance Engineer

adaption • Dublin

Hybrid
EUR 120,000 - 180,000
Flexible work
Adaption Passport travel stipend
Lunch stipend
+1
Inference Performance Engineer | Optimize AI Serving
Inference Performance Engineer | Optimize AI Serving

adaption • Dublin

Hybrid
EUR 120,000 - 180,000
Flexible work
Adaption Passport travel stipend
Lunch stipend
+1
AI Inference Performance Engineer
AI Inference Performance Engineer

Qualcomm • Cork

Hybrid
EUR 90,000 - 130,000
Salary review and performance bonus
Relocation support
Education Assistance
+2
Cloud AI Inference Performance Engineer
Cloud AI Inference Performance Engineer

Qualcomm • Cork

On-site
EUR 90,000 - 140,000
Salary, stock and performance bonus
Relocation support
Education Assistance
+4
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility