AI Inference Platform Engineer

XpertDirect

Berlin

Vor Ort

EUR 70.000 - 110.000

Vollzeit

Vor 8 Tagen
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

XpertDirect in Berlin specializes in AI infrastructure, delivering scalable model serving on GPU-enabled platforms. The AI Inference Platform Engineer will build, optimise and operate the production inference stack across Kubernetes, vLLM, NVIDIA Triton, and cloud GPU environments.

You will focus on improving latency, throughput, and GPU utilisation while designing autoscaling, observability, and reliable deployment workflows.

Qualifikationen

  • 4+ years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering or similar roles.
  • Strong understanding of Linux and distributed production systems.

Aufgaben

  • Build and operate Kubernetes infrastructure for production AI inference
  • Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton
  • Improve GPU utilisation, throughput, and inference latency
  • Design autoscaling strategies for dynamic AI workloads
  • Build platform tooling and automation in Python
  • Provision and manage infrastructure using Terraform
  • Develop observability across models, GPUs, Kubernetes, and serving infrastructure
  • Profile and troubleshoot inference performance bottlenecks
  • Improve batching, concurrency, caching, and resource allocation strategies
  • Build reliable deployment workflows for new models and model versions
  • Partner with ML Engineers to move models efficiently into production

Kenntnisse

4+ years in AI Infrastructure, ML Infr
Linux & distributed systems

Tools

CUDA
NVIDIA GPU Operator
PyTorch
KServe
Ray Serve
LLM inference optimisation
Multi-GPU inference
AWS / GCP GPU infrastructure

Jobbeschreibung

AI Infrastructure | Model Serving | GPU Computing | Inference Engineering | ML Platforms

Our client, a growing AI Infrastructure company based in Berlin, is looking for an AI Inference Platform Engineer to build and optimise the platform used to serve production AI models across GPU-enabled infrastructure.

You'll work at the intersection of AI Infrastructure, Distributed Systems, and Platform Engineering, focusing on inference performance, GPU utilisation, autoscaling, latency, and reliability.

What You’ll Work On
  • Build and operate Kubernetes infrastructure for production AI inference
  • Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton
  • Improve GPU utilisation, throughput, and inference latency
  • Design autoscaling strategies for dynamic AI workloads
  • Build platform tooling and automation in Python
  • Provision and manage infrastructure using Terraform
  • Develop observability across models, GPUs, Kubernetes, and serving infrastructure
  • Profile and troubleshoot inference performance bottlenecks
  • Improve batching, concurrency, caching, and resource allocation strategies
  • Build reliable deployment workflows for new models and model versions
  • Partner with ML Engineers to move models efficiently into production
Core Skills
  • 4+ years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering, or similar roles
  • Strong understanding of Linux and distributed production systems
Nice to Have
  • CUDA
  • NVIDIA GPU Operator
  • PyTorch
  • KServe
  • Ray Serve
  • LLM inference optimisation
  • Multi-GPU inference
  • AWS / GCP GPU infrastructure
  • Experience operating high-throughput or latency-sensitive inference services
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Competitive compensation
Learning opportunities
Ownership of projects
+4
ML Deployment Engineer
ML Deployment Engineer

XpertDirect • München

Vor Ort
EUR 90.000 - 130.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 180.000
Infrastructure Operations Engineer
Infrastructure Operations Engineer

lightningai • Deutschland

Hybrid
EUR 138.000 - 172.000
Discretionary bonus
Equity
401(k) matching
+1
AI Infrastructure Architect (All Genders)
AI Infrastructure Architect (All Genders)

Accenture • Kronberg im Taunus

Vor Ort
EUR 120.000 - 160.000
AI Infrastructure Engineer
AI Infrastructure Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

Meyandy LLC • Berlin

Hybrid
EUR 110.000 - 150.000
Competitive compensation
Career growth and learning oppor tunun
Flexible working arrangements
Compute Solution Architect
Compute Solution Architect

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA • Deutschland

Vor Ort
EUR 293.000 - 507.000