T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

Magentateam

Warszawa

On-site

PLN 240,000 - 360,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
Sport and life insurance
Preferential terms
Contract-based cooperation
Referral bonus

Job summary

Magentateam in Warsaw is seeking an experienced DevOps/SRE professional to design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes with bare-metal GPU infrastructure.

You will manage NVIDIA GPU resources, automate model lifecycle, implement autoscaling, observability, and secure networking while collaborating with AI and Platform teams to deliver reliable, scalable AI services.

Qualifications

  • 5+ years of experience in DevOps/SRE/Platform/Infrastructure operations.
  • 2+ years supporting MLOps, AI infra, or LLM platforms.
  • Strong Kubernetes and OpenShift administration in production.
  • Experience deploying and operating vLLM-based inference platforms.
  • Solid understanding of LLM serving concepts and inference optimization.
  • Hands-on with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
  • Strong Python scripting and automation skills.
  • Experience with GitLab CI, Jenkins, ArgoCD, and IaC practices.

Responsibilities

  • Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes with bare-metal GPU infra.
  • Manage NVIDIA GPU resources including allocation and partitioning for multiple models/tenants.
  • Automate model lifecycle: onboarding, versioning, deployment, hot-swapping, rollback.
  • Implement HPA based on workload, queue depth, and GPU usage.
  • Build observability: metrics, logging, tracing; expose token/usage metrics.
  • Develop usage tracking for API keys, users, teams for quota management.

Skills

Kubernetes
OpenShift
Python
Automation
Observability
Linux
CI/CD
Security
Networking

Tools

OpenShift
Kubernetes
NVIDIA GPUs
Prometheus
Grafana
ELK Stack
OpenTelemetry
CI/CD Tools

Job description

  • Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
  • Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
  • Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
  • Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
  • Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
  • Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
  • Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
  • Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
  • Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
  • Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
  • Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
  • Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
  • Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
  • Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
WHAT SKILLS WILL BE APPRECIATED?

5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.

At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.

Strong experience with Kubernetes and OpenShift administration in production environments.

Proven experience deploying and operating vLLM-based inference platforms in production.

Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.

Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.

Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.

Strong Python programming skills with experience developing automation and operational tooling.

Experience with Bash scripting and Linux systems administration.

Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.

Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.

OUR OFFER FOR YOU

Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.

You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.

No dress code - you can just be yourself here

Medical, sport and life insurance packages at preferential terms

Access to our products and services at preferential terms

Employment contract-based cooperation

Know Talent - receive training or financial bonus for recommending new employees

WHAT WILL YOUR RECRUITMENT PROCESS BE LIKE?

A fair approach to all people who want to join T Hub means that:

  • The recruitment process is transparent;
  • Our recruitment decision is based solely on an assessment of your skills (your race, skin color, sexual orientation, gender identity, origin, disability, political view, appearance, or religion will not have any influence on he outcome of the process):
  • Regardless of the outcome of the process, you will get detailed feedback.
Let’s meet to better understand each other's expectations.

Our screening meeting will last about 20-30 minutes. We will ask you about our areas of interest and will be happy to field any questions you may have.

Hiring Manager will receive a recommendation for your application.

There’s nothing for you to do at this stage — we’ll take care of everything. During this time we process all the information we've gathered during the screening process and dig deeper into your CV.

Technical meetings with the Hiring Manager and/or team members.

We will invite you to one or two project meetings to confirm that we have a ‘perfect match'. The meetings may last between 30 minutes and 90 minutes, during which time we will talk to you about mutual expectations and the vision for our collaboration. You will also most likely meet your future supervisor and your teammate during this time.

At this stage, we will have decided that you are the person we want to develop our projects with. We will then come back to you with an offer of collaboration, hoping for your "yes". If for some reason, we are unable to offer you a position in our team, you will certainly receive feedback from us explaining our decision.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

T-Hub - AI Agents & LLM Engineer
T-Hub - AI Agents & LLM Engineer

Magentateam • Warszawa

On-site
PLN 180,000 - 360,000
Medical insurance
Life insurance
No dress code
+3
T Hub - IoT Service Operations Engineer (Mobile Core)
T Hub - IoT Service Operations Engineer (Mobile Core)

Magentateam • Warszawa

Hybrid
PLN 120,000 - 190,000
Private medical care
Life insurance
Hybrid work model
+3
T Hub - Data Architect – Microsoft Fabric & Data Intelligence
T Hub - Data Architect – Microsoft Fabric & Data Intelligence

Magentateam • Warszawa

Hybrid
PLN 220,000 - 320,000
Medical package
Life insurance
Training platforms - Percipio, Courser
+3
T Hub - AI Strategy Expert
T Hub - AI Strategy Expert

Magentateam • Warszawa

On-site
PLN 150,000 - 210,000
International exposure
Flexible workplace
No dress code
+3
T Hub - Senior Fullstack Developer - Backend-Heavy (Node.js/& React)
T Hub - Senior Fullstack Developer - Backend-Heavy (Node.js/& React)

Magentateam • Warszawa

On-site
PLN 180,000 - 240,000
Medical insurance
Sport insurance
Life insurance
+1
T Hub - IoT Test Automation Engineer
T Hub - IoT Test Automation Engineer

Magentateam • Warszawa

Hybrid
PLN 120,000 - 180,000
Medical insurance
Gift terms for employee referrals
Wellness program
T Hub - Lead SAP Integration & Replikation Expert with German
T Hub - Lead SAP Integration & Replikation Expert with German

Magentateam • Warszawa

On-site
PLN 180,000 - 260,000
Medical insurance
Sport and life insurance
Training/bonus for referrals
+1
T Hub - SAP SuccessFactors Employee Central Payroll & Time Consultant with German
T Hub - SAP SuccessFactors Employee Central Payroll & Time Consultant with German

Magentateam • Warszawa

On-site
PLN 180,000 - 240,000
Medical insurance
Preferential terms on products and ser
Wrocław T Hub - Research Area (R&D) Lead – Network Capabilities
Wrocław T Hub - Research Area (R&D) Lead – Network Capabilities

Magentateam • Poland

Remote
PLN 250,000 - 420,000
Insurance packages
Access to company products
Contract-based cooperation
+1
T Hub - Developer SAP ABAP HCM
T Hub - Developer SAP ABAP HCM

Magentateam • Poland

Remote
PLN 180,000 - 240,000
Medical, sport and life insurance
Preferential terms for products and/or
Employment contract-based cooperation
+1