Lead AI Application Engineer – Infrastructure, LLMOps

Jobtailor

Deutschland

Vor Ort

EUR 120.000 - 180.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is looking for an experienced Platform Engineer to build and run a shared AI platform across cloud and on‑prem environments. You will design multi‑tenant infra, ensure high availability, and optimize costs while enabling scalable ML workflows.

You will implement LLMOps, curate AI services, expose as‑a‑service capabilities, and guide internal squads with standardized APIs. Expertise in Kubernetes, cloud platforms, and AI tooling is essential to deliver fast inference and robust data

Qualifikationen

  • 8+ years in Platform Engineering, DevOps or SRE.
  • 2+ years focused on AI/ML infrastructure or platforms.
  • Experience building Internal Developer Platforms (IDP) is a plus.
  • Hybrid cloud experience across AWS/Azure/GCP and on‑prem.

Aufgaben

  • Build and run the Shared AI Platform across cloud and on‑prem environments.
  • Architect and maintain a multi‑tenant AI Platform that supports the full ML lifecycle across cloud and on‑premises environments
  • Ensure high availability, low latency, and cost‑efficiency for all shared AI resources
  • Implement LLMOps/MLOps best practices, including automated deployment pipelines for models
  • Curate the AI Services Catalogue
  • Develop and expose “as‑a‑service” capabilities: Inference‑as‑a‑Service, Embeddings‑as‑a‑Service, and RAG‑as‑a‑Service
  • Standardize how squads interact with LLMs, providing unified APIs and abstraction layers to prevent vendor lock‑in
  • Manage AI Data Infrastructure
  • Own the deployment and scaling of Vector Databases (e.g., Pinecone, Milvus, Weaviate) and Feature Stores (e.g., Feast, Tecton, Hopsworks)
  • Optimize data retrieval patterns to support real‑time AI applications and agentic workflows
  • Oversee Model Hosting environments, utilizing Kubernetes (K8s) and GPU orchestration to manage compute resources efficiently
  • Enable Developer Self‑Service
  • Build and maintain a Self‑Service Portal or CLI that allows product squads to provision AI environments, models, and data stores independently
  • Reduce “Time‑to‑Inference” for new features by providing pre‑configured templates and blueprints
  • Conduct internal workshops and provide documentation to empower squads to use the platform effectively

Kenntnisse

Kubernetes
Docker
Terraform
Pulumi
OpenShift
NVIDIA Triton
vLLM
TGI
Vector Databases
Python
Go
Rust
AWS
Azure
GCP
On-Premises
Platform Engineering

Tools

Pinecone
Milvus
Weaviate
Feast
Tecton
Hopsworks
NVIDIA Triton
vLLM

Jobbeschreibung

Responsibilities
  • Build & Run the Shared AI Platform
  • Architect and maintain a multi‑tenant AI Platform that supports the full ML lifecycle across cloud and on‑premises environments
  • Ensure high availability, low latency, and cost‑efficiency for all shared AI resources
  • Implement LLMOps/MLOps best practices, including automated deployment pipelines for models
  • Curate the AI Services Catalogue
  • Develop and expose “as‑a‑service” capabilities: Inference‑as‑a‑Service, Embeddings‑as‑a‑Service, and RAG‑as‑a‑Service
  • Standardize how squads interact with LLMs, providing unified APIs and abstraction layers to prevent vendor lock‑in
  • Manage AI Data Infrastructure
  • Own the deployment and scaling of Vector Databases (e.g., Pinecone, Milvus, Weaviate) and Feature Stores (e.g., Feast, Tecton, Hopsworks)
  • Optimize data retrieval patterns to support real‑time AI applications and agentic workflows
  • Oversee Model Hosting environments, utilizing Kubernetes (K8s) and GPU orchestration to manage compute resources efficiently
  • Enable Developer Self‑Service
  • Build and maintain a Self‑Service Portal or CLI that allows product squads to provision AI environments, models, and data stores independently
  • Reduce “Time‑to‑Inference” for new features by providing pre‑configured templates and blueprints
  • Conduct internal workshops and provide documentation to empower squads to use the platform effectively
Requirements
  • Must‑Have Technical Skills
  • Infrastructure: Deep experience with Kubernetes (K8s), Docker, and Terraform/Pulumi
  • Hybrid Cloud: Proven experience managing workloads across AWS/Azure/GCP and On‑Premises (NVIDIA AI Enterprise, OpenShift)
  • AI/ML Tooling: Hands‑on experience with vLLM, TGI (Text Generation Inference), or NVIDIA Triton for model serving
  • Databases: Expertise in Vector DBs and traditional SQL/NoSQL databases
  • Languages: High proficiency in Python and Go or Rust for platform tooling
  • Experience 8+ years in Platform Engineering, DevOps, or Site Reliability Engineering (SRE)
  • 2+ years specifically focused on building AI/ML infrastructure or platforms
  • Experience building Internal Developer Platforms (IDP) is a massive plus
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Platform Engineer
AI Platform Engineer

Jobtailor • München

Vor Ort
EUR 90.000 - 120.000
Senior AI Infrastructure Engineer, LLM/AI Platforms
Senior AI Infrastructure Engineer, LLM/AI Platforms

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Senior AI DevOps / LLMOps
Senior AI DevOps / LLMOps

United States Digital Space LLC • Baden-Baden

Vor Ort
EUR 80.000 - 110.000
Senior AI Engineer
Senior AI Engineer

DRIMCO GmbH • München

Hybrid
EUR 90.000 - 140.000
Remote work not specified
Security and Loss Prevention Cluster Manager, Worldwide Operations Security
Security and Loss Prevention Cluster Manager, Worldwide Operations Security

Amazon VZ Berlin-Brandenburg GmbH • Schönefeld

Vor Ort
EUR 110.000 - 170.000
AI Infrastructure Architect (All Genders)
AI Infrastructure Architect (All Genders)

Accenture • Kronberg im Taunus

Vor Ort
EUR 120.000 - 160.000
Senior Generative AI Operations (GenAI Ops) Engineer
Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems • Deutschland

Hybrid
EUR 70.000 - 90.000
AI Engineer (all levels)
AI Engineer (all levels)

Secure Systems Engineering GmbH • Berlin

Hybrid
EUR 60.000 - 90.000
Flexible hybrid working
Comfortable travel policy
Continuous training programs
Staff AI Engineer - 2nd Horizon
Staff AI Engineer - 2nd Horizon

Grafana Labs • Deutschland

Remote
EUR 70.000 - 110.000
Lead Engineer - AI Studio
Lead Engineer - AI Studio

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000