Senior AI Infrastructure Engineer, LLM/AI Platforms

Jobtailor

Deutschland

Vor Ort

EUR 120.000 - 160.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is seeking a senior infrastructure/data engineer to scale LLM platforms in a production environment. You will design and maintain GPU clusters, optimize training and inference pipelines, and lead model lifecycle management with a focus on reproducibility.

You will collaborate with data scientists, product managers, and engineering teams to transform research prototypes into robust, scalable services, while championing best practices in MLOps and system reliability.

Qualifikationen

  • Bachelor’s degree in Computer Science, Data Engineering, or related STEM field; Master’s degree preferred.
  • 6+ years in Infrastructure/Data Engineering with focus on LLM platforms and pipelines.
  • Hands-on experience in LLM infra including cluster provisioning and inference pipelines.
  • Write clean, performant, well-tested code with focus on delivering results quickly.
  • Peer code reviews and resilient architecture design are essential.
  • Demonstrates technical leadership and mentorship capabilities.
  • Experience using AI technologies to improve workflows and business outcomes.

Aufgaben

  • Provision and configure large GPU clusters for LLM workloads.
  • Develop and optimize LLM serving infrastructure and inference frameworks.
  • Lead model lifecycle management including versioning and reproducibility.
  • Design evaluation frameworks for model performance and reliability.
  • Address GPU utilization bottlenecks with quantization, batching, and caching.
  • Architect data platforms and pipelines for LLMs and RAG systems.
  • Deliver production-ready code focusing on performance and testing.
  • Define best practices for MLOps/DataOps around LLMs with observability.
  • Document architectural designs and communicate decisions to stakeholders.
  • Collaborate with Data Scientists and Product Managers to scale prototypes.

Kenntnisse

LLM infra
Cluster provisioning
Data pipelines
Code quality
Technical leadership
Mentorship
AI-driven decisions

Ausbildung

Bachelor's in CS/Data Eng
Master's preferred

Tools

GPU clusters
LLM frameworks
Kubernetes
CI/CD

Jobbeschreibung

Responsibilities
  • Provision and configure large GPU clusters and compute resources for LLM training, finetuning, and inference workloads.
  • Develop and optimize LLM model‑serving infrastructure, including deployment and optimization of various inference frameworks.
  • Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments.
  • Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production‑ready standards.
  • Identify and address GPU utilization and GPU memory efficiency bottlenecks and apply techniques like quantization, batching, and caching.
  • Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval‑Augmented Generation (RAG), and AI Agentic Systems at scale.
  • Deliver production‑ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality.
  • Apply expertise in data modeling, normalization, and semantic cataloging for AI/ML workloads.
  • Define and enforce best practices for MLOps/DataOps surrounding LLMs, including monitoring, observability, and zero‑touch recovery mechanisms for AI services.
  • Document architectural designs thoroughly and communicate technical decisions clearly to stakeholders.
  • Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production‑grade services.
Requirements
  • Bachelor’s degree in Computer Science, Data Engineering, or a related STEM field; Master’s degree preferred.
  • 6+ years of experience in Infrastructure/Data Engineering, with at least 2 years focused on building and maintaining platforms/pipelines that support LLM‑based systems and applications.
  • Demonstrable hands‑on experience in LLM infrastructure engineering including cluster provisioning, optimizing training workloads, and maintaining inference pipelines.
  • Exceptional ability to write clean, elegant, performant, and well‑tested code, coupled with a strong focus on action and delivering results quickly.
  • Thorough understanding of engineering practices including effective peer code reviews and resilient architecture design.
  • Demonstrates technical leadership and mentorship capabilities.
  • Proven experience utilizing AI technologies to enhance decision‑making, streamline workflows and processes, improve efficiency and drive business outcomes.
Core Competencies

Expertise in provisioning and optimizing GPU clusters for LLM training and inference, with a strong focus on MLOps best practices and robust model lifecycle management. Proven ability to deliver high‑quality, production‑ready code while collaborating effectively with cross‑functional teams.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior AI & Software Engineer
Senior AI & Software Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Staff AI Engineer
Staff AI Engineer

Bluefish • Berlin

Hybrid
EUR 90.000 - 120.000
Senior AI Engineer
Senior AI Engineer

DRIMCO GmbH • München

Hybrid
EUR 90.000 - 140.000
Remote work not specified
Senior AI DevOps / LLMOps
Senior AI DevOps / LLMOps

United States Digital Space LLC • Baden-Baden

Vor Ort
EUR 80.000 - 110.000
AI Engineer (all levels)
AI Engineer (all levels)

Secure Systems Engineering GmbH • Berlin

Hybrid
EUR 60.000 - 90.000
Flexible hybrid working
Comfortable travel policy
Continuous training programs
AI Infrastructure Lead Architect (All Genders)
AI Infrastructure Lead Architect (All Genders)

Accenture DACH • Kronberg im Taunus

Vor Ort
EUR 120.000 - 180.000
AI Infrastructure Architect (All Genders)
AI Infrastructure Architect (All Genders)

Accenture • Kronberg im Taunus

Vor Ort
EUR 120.000 - 160.000
AI Infrastructure Lead Architect (All Genders)
AI Infrastructure Lead Architect (All Genders)

Accenture • Kronberg im Taunus

Vor Ort
EUR 120.000 - 180.000
AI Platform Engineer
AI Platform Engineer

Jobtailor • München

Vor Ort
EUR 90.000 - 120.000
Senior Generative AI Operations (GenAI Ops) Engineer
Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems • Deutschland

Hybrid
EUR 70.000 - 90.000