Hybrid/Remote Senior AI Inference Engineer

F5 Networks, Inc. 

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

F5 Networks, Inc. is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments.

The role focuses on optimizing Large Language Models for inference across data centers and edge devices, prioritizing throughput, low latency, and accuracy. You will build scalable inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, and optimize hardware usage on NVIDIA GPUs, Apple Silicon, TPUs, and other accelerators.

Qualifications

  • Programming in Python, C++, Rust or Go for high-performance AI workflows.
  • Experience with vLLM, TensorRT, Llama.cpp or Ollama for inference development.
  • Familiarity with Docker, Kubernetes and cloud platforms (AWS, GCP, Azure).
  • Understanding hardware optimization for GPUs/TPUs in AI workloads.

Responsibilities

  • Build and maintain high-performance AI serving engines for low latency inference.
  • Profile and optimize models on NVIDIA GPUs, Apple Silicon and other accelerators.
  • Design auto-scaling architectures using Kubernetes for online and batch inference.
  • Establish observability to monitor TTFT, tokens/second, and memory bandwidth against SLAs.

Skills

Python
C++
Rust
Golang

Tools

vLLM
TensorRT
Llama.cpp
Ollama

Job description

F5 Networks, Inc. is seeking an AI Inference Engineer to bridge high-performance model development and optimized deployment environments.

The role focuses on optimizing Large Language Models for inference across data centers and edge devices, prioritizing throughput, low latency, and accuracy. You will build scalable inference engines with vLLM, TensorRT, Llama.cpp, and Ollama, and optimize hardware usage on NVIDIA GPUs, Apple Silicon, TPUs, and other accelerators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer — High-Performance, Low-Latency ML
AI Inference Engineer — High-Performance, Low-Latency ML

F5 • Dublin

On-site
EUR 90,000 - 150,000
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
AI Inference Engineer
AI Inference Engineer

F5 • Dublin

On-site
EUR 90,000 - 150,000
Senior Site Reliability Engineer, AI Inference
Senior Site Reliability Engineer, AI Inference

F5 Networks, Inc.  • Dublin

Hybrid
EUR 120,000 - 180,000
Senior ML Systems Engineer (Inference) – Hybrid + Tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
Staff AI Model Optimization Architect — Scalable Inference
Staff AI Model Optimization Architect — Scalable Inference

Qualcomm • Cork

On-site
EUR 150,000 - 190,000
Salary and stock bonus
Relocation assistance
Education assistance
+4