Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.
XpertDirect in Berlin specializes in AI infrastructure, delivering scalable model serving on GPU-enabled platforms. The AI Inference Platform Engineer will build, optimise and operate the production inference stack across Kubernetes, vLLM, NVIDIA Triton, and cloud GPU environments.
You will focus on improving latency, throughput, and GPU utilisation while designing autoscaling, observability, and reliable deployment workflows.
AI Infrastructure | Model Serving | GPU Computing | Inference Engineering | ML Platforms
Our client, a growing AI Infrastructure company based in Berlin, is looking for an AI Inference Platform Engineer to build and optimise the platform used to serve production AI models across GPU-enabled infrastructure.
You'll work at the intersection of AI Infrastructure, Distributed Systems, and Platform Engineering, focusing on inference performance, GPU utilisation, autoscaling, latency, and reliability.