AI Infrastructure Engineer

vCluster Labs

United States

Remote

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
Workplace Flexibility

Job summary

vCluster Labs seeks an AI Infrastructure Specialist to drive GPU-focused Kubernetes deployments from bare metal to production-ready vCluster environments. You will work pre-sales as part of proof-of-value engagements and help customers scale installations across environments.

Ideal candidates have 5+ years operating Kubernetes in production, with hands-on NVIDIA GPU tooling and distributed storage experience. This role emphasizes fast-moving, collaborative problem solving in a remote-first setup.

Qualifications

  • 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.
  • Hands-on experience with NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.
  • Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity in layered environments.
  • Experience with persistent volumes and distributed storage backends (Ceph, Rook, Longhorn).
  • Ability to write playbooks and runbooks in real time in ambiguous, fast-moving environments.

Responsibilities

  • Lead end-to-end technical deployments for GPU neocloud and AI Factory customers, from bare metal to validated vCluster environments.
  • Configure and troubleshoot bare metal GPU node infrastructure, including CNI, GPU Operator, and storage backends.
  • Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.
  • Work with customer teams to build self-sufficiency and platform adoption.
  • Document reusable playbooks and deployment architectures for future customers.
  • Provide feedback to Engineering/Product on infrastructure challenges for the roadmap.
  • Partner with Sales in pre-sales to support value realization.

Skills

Kubernetes deployments
Bare metal GPU
Kubernetes validation
Customer enablement
Documentation
Pre-sales collaboration

Tools

NVIDIA GPU Operators
CUDA tooling
CNI
RDMA/InfiniBand
Ceph/Rook/Longhorn
Kubernetes Operators

Job description

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

  • Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.

  • Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.

  • Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.

  • Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.

  • Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer’s head start.

  • Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.

  • Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

This role could be a fit for you if you bring:

  • Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.

  • GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.

  • Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.

  • Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.

  • Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.

  • Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Bonus points for:
  • Automation Skills: Experience writing automation scripts with Bash, Python, or Go.

  • Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.

  • AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.

  • Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

About vCluster Labs

We are a venture-backed tech startup and the company pioneering Kubernetes virtualization for the AI era. We raised +$30M from top-tier VCs such as Khosla Ventures (first investor in OpenAI, GitLab, Stripe, Doordash) and are in a hyper-growth phase looking for motivated people to complement our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.

We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU.

We’re the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitHub stars, 40M+ virtual clusters created since 2021). Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI — a Kubernetes-native framework purpose-built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

  • Competitive Salary : We offer a competitive compensation package, including equity.

  • Platinum-Level Insurance : Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

  1. Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

  2. Own the Outcome: We understand that our responsibility doesn’t end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

  3. Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

  4. Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.

  5. Build Tomorrow’s Standards, Intentionally: We don’t just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission‑critical infrastructure platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

vCluster • New York (NY)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Vcluster • Northern (KY)

On-site
USD 140,000 - 165,000
Competitive Salary
Equity participation
Health, dental, vision, life Insurance
+2
Customer Success Engineer
Customer Success Engineer

vCluster Labs • Chicago (IL)

Remote
USD 90,000 - 140,000
Competitive Salary
Premium Insurance
Flexible Working Schedule
+1
Customer Success Engineer
Customer Success Engineer

vCluster Labs • San Francisco (CA)

Remote
USD 140,000 - 190,000
Competitive Salary
Premium Insurance
Flexible Working Schedule
+1
Technical Account Manager
Technical Account Manager

Vcluster • Northern (KY)

On-site
USD 150,000 - 175,000
Competitive Salary
Platinum‑Level Insurance
Flexible Working Schedule
+1
Social Media Manager
Social Media Manager

vCluster Labs • United States

Remote
USD 70,000 - 120,000
Equity
Health/Dental/Vision/Life
Flexible schedule
+1
Customer Success Engineer
Customer Success Engineer

Vcluster • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 175,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Enterprise Account Executive
Enterprise Account Executive

vCluster Labs • Chicago (IL)

Remote
USD 180,000 - 320,000
Competitive Salary
Premium Insurance
Flexible Working Schedule
+1
Social Media Manager vCluster Labs · USD 105k-125k/yr Workplace 6 hours ago
Social Media Manager vCluster Labs · USD 105k-125k/yr Workplace 6 hours ago

Content Creators • Northern (KY)

On-site
USD 80,000 - 110,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Sr. Global Alliance Director, NVIDIA and AI Clouds
Sr. Global Alliance Director, NVIDIA and AI Clouds

vCluster Labs • San Francisco (CA)

Remote
USD 170,000 - 250,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1