AI Infrastructure Engineer

Vcluster

Northern (KY)

Hybrid

USD 140,000 - 165,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Equity participation
Health, dental, vision, life Insurance
Flexible working schedule
Remote-first culture

Job summary

vCluster Labs is seeking an AI Infrastructure Specialist to lead deployments from bare metal GPUs to production-grade Kubernetes environments. This is a pre-sales / proof-of-value role integrated with the customer journey to scale playbooks for future customers.

The role requires 5+ years of production Kubernetes experience, hands-on GPU tooling, and strong networking/storage fundamentals. You will document architectures and work with Sales to drive customer value while maintaining a

Qualifications

  • 5+ years of production Kubernetes experience, ideally bare metal.
  • Experience deploying and operating Kubernetes in production.
  • Experience with NVIDIA GPU tooling and CUDA.
  • Strong networking knowledge: CNI, overlays, load balancing.
  • Experience with Ceph/Rook/Weka/Longhorn distributed storage.
  • Automation scripting with Bash, Python, or Go.
  • CKA certification or Kubernetes Operators experience.
  • AI/ML familiarity with inference serving and GPU scheduling.

Responsibilities

  • Lead end-to-end technical deployments for GPU neocloud and AI Factory customers.
  • Configure and troubleshoot bare metal GPU node infrastructure and related storage.
  • Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.
  • Collaborate with customer teams to build self-sufficiency and reusable playbooks.
  • Document deployment architectures to accelerate future customers.
  • Provide feedback to Engineering and Product to address recurring issues.
  • Support pre-sales activities with deep infrastructure work.

Skills

Kubernetes
GPU Operators
Networking
Storage
Automation
CUDA
Documentation
CKA
AI/ML

Tools

Ceph
Rook
Weka
Longhorn

Job description

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.

Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.

Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.

Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.

Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.

Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.

Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

This role could be a fit for you if you bring:

Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.

GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.

Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.

Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.

Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.

Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Automation Skills: Experience writing automation scripts with Bash, Python, or Go.

Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.

AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.

Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

About vCluster Labs

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

We offer the following benefits:

Competitive Salary: We offer a competitive compensation package, including equity.

Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.

Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Compensation Range:

  • United States: Estimated Base Salary $140K – $165K Offers Equity Offers Bonus
  • Australia: Estimated Base Salary A$150K – A$175K Offers Equity Offers Bonus
  • Germany: Estimated Base Salary €90K – €105K Offers Equity Offers Bonus
  • Singapore: Estimated Base Salary SGD155K – SGD175K Offers Equity Offers Bonus
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

vCluster • New York (NY)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Customer Success Engineer
Customer Success Engineer

Vcluster • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 175,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Technical Account Manager
Technical Account Manager

Vcluster • Northern (KY)

Hybrid
USD 150,000 - 175,000
Competitive Salary
Platinum‑Level Insurance
Flexible Working Schedule
+1
Social Media Manager vCluster Labs · USD 105k-125k/yr Workplace 6 hours ago
Social Media Manager vCluster Labs · USD 105k-125k/yr Workplace 6 hours ago

Content Creators • Northern (KY)

Hybrid
USD 80,000 - 110,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Corporate Counsel
Corporate Counsel

Vcluster • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Sr. Global Alliance Director, Channel and GSI
Sr. Global Alliance Director, Channel and GSI

vCluster • United States

On-site
USD 225,000 - 280,000
Competitive salary+
Equity participation
Flexible scheduling
+1
Product Marketing Manager
Product Marketing Manager

Vcluster • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 170,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Engineering Tech Lead (vMetal)
Engineering Tech Lead (vMetal)

Vcluster • Northern (KY)

Hybrid
USD 140,000 - 200,000
Equity
Health insurance
Flexible schedule
+1
Enterprise Account Executive
Enterprise Account Executive

vCluster • Town of Texas (WI)

On-site
USD 280,000 - 330,000
Competitive Salary
Platinum Insurance
Flexible Schedule
+1
Senior Product Manager (vMetal)
Senior Product Manager (vMetal)

vCluster • San Francisco (CA)

On-site
CAD 236,000 - 278,000
Competitive salary + equity
Platinum-level insurance
Flexible schedule
+1