Senior AI DevOps / LLMOps

United States Digital Space LLC

Baden-Baden

Vor Ort

EUR 80.000 - 110.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

United States Digital Space LLC is seeking a Senior AI DevOps / LLMOps Specialist in Baden-Baden, Germany. This exciting opportunity involves automating AI build processes, designing CI/CD pipelines, and managing hybrid cloud environments with a focus on AI.

The ideal candidate will have over 10 years of experience in DevOps or Cloud Engineering and at least 2 years in MLOps, particularly in deploying large language models. Join a dynamic team that emphasizes growth and innovation!

Qualifikationen

  • 10+ years in DevOps, SRE, or Cloud Engineering.
  • 2+ years of hands-on experience in MLOps or LLMOps, moving LLMs to production.
  • Proven experience managing Hybrid Cloud environments.

Aufgaben

  • Automate Build-to-Production processes.
  • Design CI/CD pipelines for AI models.
  • Develop workflows for PromptOps and manage stateful AI workflows.
  • Provision high-performance compute environments using IaC.
  • Architect Progressive Delivery strategies for AI.
  • Establish observability into Inference Endpoints.

Kenntnisse

Orchestration: Advanced Kubernetes (K8s)
CI/CD & IaC: GitHub Actions/GitLab CI
AI Tooling: Weights & Biases, MLflow
Understanding of GPU virtualization
Familiarity with Open Policy Agent (OPA)

Tools

Terraform
Pulumi

Jobbeschreibung

At the company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOps specialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you.

Key Responsibilities
  • Automation of Build-to-Production.
  • Design and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code.
  • Develop specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code.
  • Automate the deployment of Agentic workflows, managing the complexities of stateful AI interactions and multi-agent handoffs.
  • AI Infrastructure as Code (IaC): Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, Pulumi, or Ansible.
  • AI Infrastructure as Code (IaC): Define and enforce Policy-as-Code for AI endpoints to ensure compliance with security, cost-usage limits, and data residency requirements.
  • AI Infrastructure as Code (IaC): Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless parity between On-Premises development and Cloud production.
  • Safe Experimentation & Controlled Releases: Architect Progressive Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in parallel with production to compare outputs).
  • Safe Experimentation & Controlled Releases: Build “Evaluation-in-the-Loop” gates within the pipeline to automatically test for bias, hallucination, and performance degradation before a release.
  • Safe Experimentation & Controlled Releases: Implement A/B testing frameworks specifically designed for LLM outputs and agentic behavior.
  • Monitoring & Observability: Establish deep observability into Inference Endpoints, tracking metrics like tokens-per-second, latency, and drift in model accuracy.
  • Monitoring & Observability: Integrate feedback loops that capture production “edge cases” to feed back into the training and fine-tuning pipelines.
Must-Have Technical Skills
  • Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or NVIDIA Triton.
  • CI/CD & IaC: Expertise in GitHub Actions/GitLab CI, and Terraform or Pulumi.
  • AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize Phoenix.
  • Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises hardware management.
  • Security: Familiarity with Open Policy Agent (OPA) and secret management (Vault).
Experience
  • 10+ years in DevOps, SRE, or Cloud Engineering.
  • 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs from notebook to production.
  • Proven experience managing Hybrid Cloud environments (e.g., AWS/Azure + Private Data Center).
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior AI Engineer
Senior AI Engineer

traide • Berlin

Vor Ort
EUR 70.000 - 110.000
Senior Applied AI Engineer (all genders)
Senior Applied AI Engineer (all genders)

Accenture DACH • Kronberg im Taunus

Vor Ort
EUR 90.000 - 130.000
Senior MLOps Engineer
Senior MLOps Engineer

Versatile People • Deutschland

Remote
EUR 90.000 - 150.000
Senior Artificial Intelligence Engineer
Senior Artificial Intelligence Engineer

IPI Technolab • Deutschland

Vor Ort
EUR 90.000 - 130.000
AI Large Language Model (LLM) Technology Architect (All Genders)
AI Large Language Model (LLM) Technology Architect (All Genders)

Accenture DACH • Würzburg

Vor Ort
EUR 110.000 - 150.000
DevOps Engineer (all genders)
DevOps Engineer (all genders)

Roland Berger GmbH • Bayern

Vor Ort
EUR 90.000 - 120.000
AI Devops Engineer - CST Time Zone
AI Devops Engineer - CST Time Zone

Ekshvaku Tech Innovation • Deutschland

Remote
EUR 110.000 - 140.000
Frontier Engineer (M/F/D)
Frontier Engineer (M/F/D)

Cognizant • Karlsruhe

Hybrid
EUR 90.000 - 130.000
Agentic AI Platform Engineer (with MLOps Expertise)
Agentic AI Platform Engineer (with MLOps Expertise)

Prisma Sync Tech • Deutschland

Remote
EUR 90.000 - 140.000
MLOps Engineer (m/f/d)
MLOps Engineer (m/f/d)

Advantest Corporation • Deutschland

Remote
EUR 90.000 - 120.000