AI Infrastructure Engineer

Vcluster

Deutschland

Remote

EUR 90.000 - 130.000

Vollzeit

vor 12 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive Salary
Premium Insurance
Flexible Working Schedule
Workplace Flexibility

Zusammenfassung

vCluster Labs is hiring an AI Infrastructure Engineer to drive end-to-end GPU-enabled Kubernetes deployments for neocloud and AI Factory customers. You will optimize bare metal GPU nodes, configure CNI, GPU Operator, and storage backends, and participate in pre-sales to reach a meaningful proof of value.

You will build playbooks, document architectures, and transfer knowledge to customer teams so they can operate and scale the platform.

Qualifikationen

  • 5+ years deploying and operating Kubernetes in production on bare metal or complex environments.
  • Hands-on experience with NVIDIA GPU Operators, CUDA tooling, and GPU node config.
  • Strong networking, storage, and distributed systems experience (Ceph, Rook, Longhorn).

Aufgaben

  • Lead end-to-end technical deployments for GPU neocloud and AI Factory customers from bare metal to validated vCluster environments.
  • Configure and troubleshoot bare metal GPU nodes, CNI, storage backends, and RDMA/InfiniBand.
  • Deploy and validate Kubernetes and vCluster for GPU-powered managed clusters.
  • Work with customer teams to build self-sufficiency and scalable playbooks.
  • Document reusable deployment architectures to accelerate future customer engagements.
  • Collaborate with Engineering and Product to feed field-driven infrastructure improvements and roadmap input.

Kenntnisse

Kubernetes
GPU Mastery
Networking
Storage
Automation
Kubernetes Operators
Scripting

Tools

NVIDIA GPU Operators
CUDA tooling

Jobbeschreibung

Join us to Build the Future of AI Infrastructure

We are a VC-backed tech startup in hyper-growth, redefining what's possible at the intersection of AI and cloud-native infrastructure. Our global remote-first culture is built for the best talent, wherever you are.

We're building the infrastructure that powers the world's largest AI Clouds. Filter by department or location to find your place in that work below.

Employment Type

Full time

Location Type

Remote

Department

Customer Engineering

Compensation

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.

Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.

Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.

Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.

Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.

Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.

Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

This role could be a fit for you if you bring:

Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.

GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.

Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.

Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.

Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.

Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Automation Skills: Experience writing automation scripts with Bash, Python, or Go.

Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.

AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.

Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

About vCluster Labs

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

We offer the following benefits:

Competitive Salary: We offer a competitive compensation package, including equity.

Premium Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.

Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Decisive Hiring
Values Driven Culture

Own it. Ship it. Wow them. Our five values aren't a list on our website; they're how decisions get made every day.

As we scale to meet the demands of the world’s largest AI Clouds and enterprises, your career trajectory scales with us.

Impact from Day 1

Your work directly influences our product roadmap and the global cloud-native ecosystem.

We use a data-driven approach to compensation to ensure our team is paid fairly and equitably.

Flexible Working Schedule

You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

Platinum-Level Insurance

Health, dental, vision, and life Insurance including plans for you and eligible dependents (benefits vary depending on country).

Workplace Flexibility

Work isn’t a place, it’s what you do. As a fully remote, globally distributed team, we give you the flexibility to work from where you are, without the need to relocate or commute. We’re designed to support great work, wherever it happens.

What to expect in our hiring process

1

Application Review

We review your resume, track record, and prior work examples to assess your skills and experience against the specific requirements of the role. This step ensures we respect your time by only moving forward if there is a strong potential match.

2

Talent Screening

You will have an initial conversation with a Talent Partner. This discussion focuses on your background, your motivations for seeking a new role, and an introduction to vCluster’s culture and mission. It is our chance to get to know you beyond the resume.

3

Department Screening

You will meet with the Hiring Manager for a 45-minute deep dive. This conversation focuses on the specific criteria for the role and gives you a chance to learn more about the team’s goals. This step determines if you move forward to the challenge and panel stages.

4

Our panel consists of a series of 1:1 conversations with future teammates. To ensure fairness and reduce bias, these interviews are structured with standardized questions, where each interviewer focuses on a specific area of expertise related to the role.

5

The Challenge

We believe in seeing your skills in action. Depending on the role, you will complete a work sample, presentation or technical assessment. This ensures our evaluation is objective and based on practical ability.

6

Executive Interview & Offer

The final step is a conversation with a member of our executive team. Following a successful debrief, our Talent Partner will extend an offer and welcome you to the team!

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Customer Success Engineer
Customer Success Engineer

vCluster Labs • Berlin

Remote
EUR 70.000 - 110.000
Competitive Salary
Premium Insurance
Flexible Working Schedule
+1
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Deutschland

Vor Ort
USD 150.000 - 200.000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Strategic Post-Sales Engineer for AI Infra
Strategic Post-Sales Engineer for AI Infra

vCluster • Deutschland

Vor Ort
USD 150.000 - 175.000
Competitive Salary + Equity
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior Staff Engineer - Kubernetes & Go (Remote)
Senior Staff Engineer - Kubernetes & Go (Remote)

vCluster • Deutschland

Vor Ort
USD 75.000 - 150.000
Flexible working schedule
Platinum-level health, dental, vision, life insurance
Equity options
Tech Lead, vNode Runtime & Kubernetes Isolation
Tech Lead, vNode Runtime & Kubernetes Isolation

vCluster • Deutschland

Hybrid
USD 138.000 - 180.000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Sr. Global Alliance Director, Channel and GSI
Sr. Global Alliance Director, Channel and GSI

Embedded Shishya • Deutschland

Vor Ort
EUR 194.000 - 241.000
Equity ownership
Platinum-level insurance
Flexible working schedule
+1
Senior Software Engineer | Kubernetes Automation
Senior Software Engineer | Kubernetes Automation

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Equity options
Remote-friendly environment
Learning budget
+4
Senior Software Engineer | Kubernetes Automation | EU
Senior Software Engineer | Kubernetes Automation | EU

Creandum • Deutschland

Remote
EUR 73.000 - 100.000
Equity options
Remote-first global environment
Learning budget
Engineering Tech Lead (vNode)
Engineering Tech Lead (vNode)

vCluster • Deutschland

Vor Ort
USD 138.000 - 180.000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior Software Engineer | Platform Engineering
Senior Software Engineer | Platform Engineering

Creandum • Deutschland

Remote
EUR 78.000 - 108.000
Remote-first environment
Equity options
Learning budget
+4