Principal DevOps Engineer for Industrial AI Cloud (m/f/d)

SmartRecruiters, Inc.

España

A distancia

EUR 70.000 - 110.000

Jornada completa

hace 45 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

A complete application in a minute — tailored resume and cover letter, ready to send.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Coursera training
Language classes (English, Spanish, 2)
Life insurance
Social fund
Vacation 26+ days
Medical support
Salary during medical leave

Descripción de la vacante

T-Systems is seeking a Principal DevOps Engineer for Industrial AI Cloud to guide enterprise customers through onboarding, design of PoCs, and integration of AI workloads. You will deploy GPU clusters, optimize pipelines, and ensure scalable, secure operations across complex environments.

You will mentor teams, define architectures, and lead cross-functional collaborations to accelerate customer value in a fast-paced AI platform project.

Formación

  • 3+ years of experience designing and delivering systems based on IaaS, PaaS and SaaS.
  • Experience with GPU-based infrastructure.
  • Solid knowledge of Kubernetes container-based technologies.
  • Expert knowledge of Python/Bash scripting.
  • Expert knowledge of automation deployments (Ansible/Salt-Stack/Terraform/Helm).
  • Expert knowledge of CI/CD in a Kubernetes environment and repository management.
  • Expert in automation with Git (GitHub, GitLab) and CI/CD tools like GitHub Actions or GitLab CI/CD.
  • Strong experience with Linux OS.
  • Experience with monitoring and visualization tools (Grafana, Prometheus).
  • English at B2 level; German is a plus.

Responsabilidades

  • Consult customers on all technical aspects related to GPU infrastructure and platform usage.
  • Lead onboarding and training, mentoring customer specialists on usage of GPU clusters and AI environments.
  • Design and implement PoCs, including environment setup, data processing pipelines, and deployment workflows.
  • Conduct requirement engineering, translating business needs into technical specifications.
  • Assist customers with performance optimization, troubleshooting, and validation of delivered solutions.
  • Act as the key technical contact, coordinating cross-functional teams across infrastructure, networking, automation, security, and AI services.
  • Propose and develop automation concepts to improve services, processes, and operating models.
  • Ensure best practices in reliability, scalability, and security across the customer lifecycle.
  • Support monitoring, observability, and capacity planning for AI workloads.

Conocimientos

GPU infrastructure
Kubernetes
Python/Bash scripting
Automation tools
CI/CD
Git/GitHub/GitLab
Linux OS
Monitoring tools
English proficiency
Communication skills

Herramientas

Kubernetes
Ansible
Salt-Stack
Terraform
Helm
GitHub Actions
GitLab CI/CD
CI/CD in Kubernetes
Grafana
Prometheus

Descripción del empleo

Principal DevOps Engineer for Industrial AI Cloud (m/f/d)
  • Full-time

T‑Systems is part of the Deutsche Telekom Group, with around 30.000 employees worldwide. We create technology with purpose to generate a positive impact on society. We are looking for curious talent, eager to learn, take on challenges, and contribute ideas that transform our customers’ experience.

We trust people: we offer autonomy, continuous support, and a collaborative environment where you can grow without limits. We are one global team, guided by respect, integrity, and a passion for doing better every day.

NVIDIA and Deutsche Telekom are jointly developing industrial AI cloud for Europe. This AI factory in Germany will host 10,000 GPUs across NVIDIA DGX B200 systems and RTX Pro Servers. Deutsche Telekom provides secure, sovereign and fast infrastructure, including data centers, operations, security, and AI solutions.
As DevOps Engineer Principal you will guide enterprise customers through onboarding, training, and early adoption of the AI platform. Your responsibility includes understanding customer requirements, supporting solution design, executing Proofs of Concept (PoCs), and ensuring smooth integration of customer workloads (LLMs, GPU compute, AI pipelines). You act as a trusted technical advisor, helping customers efficiently use their GPU clusters and AI toolchains.
We are looking for a DevOps Engineer Principal position who will play a key role in helping our enterprise customers successfully adopt and scale our AI platform. In this position you will guide clients through onboarding, training, and early-stage implementation of cutting-edge AI solutions. You’ll work closely with them to understand their technical and business needs, support the design of tailored architectures, and lead Proofs of Concept that validate real-world value. You will ensure smooth integration of complex workloads — from LLM deployment and GPU compute optimization to building end-to-end AI/ML pipelines. As a trusted technical advisor, you will empower customers to use their GPU clusters and AI toolchains efficiently, troubleshoot challenges, and adopt best practices that accelerate their AI journey.

What will you do?

  • Consult customers on all technical aspects related to GPU infrastructure and platform usage.
  • Lead onboarding and training, mentoring customer specialists on optimal usage of their GPU clusters and AI environments.
  • Design and implement PoCs, including environment setup, data processing pipelines, and deployment workflows.
  • Conduct requirement engineering, translating business needs into technical specifications.
  • Assist customers with performance optimization, troubleshooting, fine-tuning, and validation of delivered solutions.
  • Act as the key technical point of contact, coordinating cross-functional teams across infrastructure, networking, automation, security, and AI services.
  • Propose and develop automation concepts to improve services, processes, and operating models.
  • Ensure best practices in reliability, scalability, and security are applied across the customer lifecycle.
  • Support monitoring, observability, and capacity planning for AI workloads and GPU utilization.
  • Have 3+ years of experience in the design and delivery of systems based on IaaS, PaaS and SaaS.
  • Have experience with GPU based infrastructure.
  • Possess solid knowledge of Kubernetes container-based technologies.
  • Have expert knowledge of scripting languages (Python/Bash).
  • Possess expert knowledge of Automation tools and automation deployments (Ansible/Salt-Stack/Terraform/Helm).
  • Have expert knowledge of CI/CD in a Kubernetes environment and repository management.
  • Are expert in automation with Git (GitHub, GitLab) and CI/CD tools like GitHub Actions or GitLab CI/CD.
  • Have strong experience with Linux OS.
  • Have experience with monitoring and visualization tools (Grafana, Prometheus,…)
  • Speak English at B2 level (German is an advantage).

Other skills:

  • Good communication skills, analytical thinking, team cooperation, presentation skills, negotiation skills.
  • Project Management- Basic.
  • Quality management- Intermediate.

What do we offer you?

Work environment & flexibility

  • International, dynamic and collaborative environment.
  • T-Social: social initiatives (sports, community, health, ...).
  • Customized training: access to Coursera to learn whatever you want, whenever you want.
  • Weekly language classes (English, Spanish & German).
  • International Mentoring Sessions & Experience Days.
  • Life and accident insurance.
  • Social fund.

Wellbeing & time off

  • 26+ working days of vacation per year.
  • Free access to specialist services (medical, legal, wellness).
  • 100% salary coverage during medical leave.

And many more advantages of being part of T-Systems!

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)

Iglubit • Sevilla

Híbrido
EUR 65.000 - 95.000
Health insurance
Meal vouchers
Telemedicine
+2
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)

SmartRecruiters, Inc. • España

A distancia
EUR 70.000 - 110.000
Life and accident insurance
Social fund
Weekly language classes
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)

SmartRecruiters, Inc. • Granada

Presencial
EUR 55.000 - 75.000
Coursera access
Language classes
Insurance (life/accident)
+2
AI SDLC Lead Engineer (m/f/d)
AI SDLC Lead Engineer (m/f/d)

T-Systems International • Granada

Presencial
EUR 90.000 - 120.000
Hybrid work
Health insurance
Meal vouchers
+4
Data Engineer PFS Admin & Tool (m/f/d)
Data Engineer PFS Admin & Tool (m/f/d)

SmartRecruiters, Inc. • Granada, Madrid

Presencial
EUR 45.000 - 65.000
Coursera training
Weekly language classes (EN/DE)
Life and accident insurance
+2
Cloud Architect (m/f/d)
Cloud Architect (m/f/d)

T-Systems Iberia • Granada

Presencial
EUR 60.000 - 85.000
Flexible schedule
Continuous training
Hybrid work model
+3
AI SWE / Code Quality Validation Engineer (m/f/d)
AI SWE / Code Quality Validation Engineer (m/f/d)

T-Systems Iberia • Valencia

Híbrido
EUR 70.000 - 90.000
Hybrid work model
Health insurance
Meal vouchers
+2
Data Engineer (m/f/d)
Data Engineer (m/f/d)

T-Systems Iberia • Valencia

Presencial
EUR 30.000 - 45.000
Hybrid work model
Coursera training access
Weekly language courses (English & Ger
+5
AI SDLC Lead Engineer (m/f/d)
AI SDLC Lead Engineer (m/f/d)

T-Systems Iberia • Comunidad Valenciana

Presencial
EUR 90.000 - 130.000
Language classes (English & German)
Coursera training access
Life and accident insurance
+2
DevOps Engineer – Experienced -- IMS & Voice AI Support (m/f/d)
DevOps Engineer – Experienced -- IMS & Voice AI Support (m/f/d)

T-Systems Iberia • Granada

Presencial
EUR 42.000 - 60.000
Hybrid work model (remote/on-site)
Flexible working hours
Coursera access for training
+3