A complete application in a minute — tailored resume and cover letter, ready to send.
Magentateam in Warsaw is seeking an experienced DevOps/SRE professional to design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes with bare-metal GPU infrastructure.
You will manage NVIDIA GPU resources, automate model lifecycle, implement autoscaling, observability, and secure networking while collaborating with AI and Platform teams to deliver reliable, scalable AI services.
5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
Strong experience with Kubernetes and OpenShift administration in production environments.
Proven experience deploying and operating vLLM-based inference platforms in production.
Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
Strong Python programming skills with experience developing automation and operational tooling.
Experience with Bash scripting and Linux systems administration.
Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.
Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
No dress code - you can just be yourself here
Medical, sport and life insurance packages at preferential terms
Access to our products and services at preferential terms
Employment contract-based cooperation
Know Talent - receive training or financial bonus for recommending new employees
A fair approach to all people who want to join T Hub means that:
Our screening meeting will last about 20-30 minutes. We will ask you about our areas of interest and will be happy to field any questions you may have.
There’s nothing for you to do at this stage — we’ll take care of everything. During this time we process all the information we've gathered during the screening process and dig deeper into your CV.
We will invite you to one or two project meetings to confirm that we have a ‘perfect match'. The meetings may last between 30 minutes and 90 minutes, during which time we will talk to you about mutual expectations and the vision for our collaboration. You will also most likely meet your future supervisor and your teammate during this time.
At this stage, we will have decided that you are the person we want to develop our projects with. We will then come back to you with an offer of collaboration, hoping for your "yes". If for some reason, we are unable to offer you a position in our team, you will certainly receive feedback from us explaining our decision.