AI Architect

Unify Technologies Ltd

Plano (TX)

On-site

USD 180,000 - 240,000

Full time

17 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Unify Technologies Ltd is seeking a senior platform architect to lead the GPU-as-a-Service infrastructure, owning orchestration, multi-tenancy, and scheduling across a GPU cluster for enterprise workloads.

You will drive design reviews, guide offshore delivery, and collaborate with the practice lead on customer engagements while setting scalable, repeatable deployment standards. This role centers on architecture and hands-on guidance with Kubernetes/OpenShift and GitOps tooling.

Qualifications

  • 7+ years platform engineering experience.
  • 4+ years Kubernetes in production; OpenShift valued.
  • Demonstrated GPU workload orchestration experience.
  • Multi-tenancy design experience: isolation, quotas, RBAC.
  • Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux.
  • Experience leading distributed or offshore teams via written standards.
  • Customer-facing credibility: capable of defending designs.

Responsibilities

  • Architect the GPUaaS control plane on Kubernetes/OpenShift.
  • Design multi-tenancy end-to-end with namespaces, RBAC, quotas.
  • Own GPU scheduling and allocation policy; topology-aware placement.
  • Define service catalog and GPU metering for chargeback/showback.
  • Build reusable platform blueprint: Terraform/Helm, GitOps, runbooks.
  • Lead an offshore delivery pod; set standards and gate deliverables.
  • Partner with practice lead on discovery, solution design, and escalation.

Skills

Kubernetes
OpenShift
GPU orchestration
IaC
GitOps
Offshore lead
Customer talk track

Tools

GPU Operator
Terraform
Helm
Argo CD
Flux
Kueue
Volcano

Job description

Duration: 12+ months

Practice: AI Infrastructure / GPU-as-a-Service

About the role

We are building a GPU-as-a-Service and AI factory practice from the ground up, delivering multi-tenant GPU platforms for enterprise and industrial customers. This is the senior technical seat on that platform. You will own the orchestration and multi-tenancy architecture that turns a GPU cluster into a consumable service, set the standards our global delivery team builds against, and serve as deputy to the practice lead in customer architecture engagements.

This is an architecture role. You will design, review, and defend — and lead an offshore engineering pod that executes.

What you’ll do
  • Architect the GPUaaS control plane on Kubernetes and OpenShift: NVIDIA GPU Operator, Network Operator, device plugin, MIG manager, node feature discovery.
  • Design multi-tenancy end to end — MIG partitioning strategy, time-slicing tiers, namespace and RBAC model, network policy, quotas, priority classes, and tenant onboarding.
  • Own GPU scheduling and allocation policy: gang scheduling (Kueue, Volcano), fair-share and preemption, topology-aware placement, and Slurm integration where customers run genuine batch HPC.
  • Define the service catalog — instance shapes, self-service request flow, and GPU metering for chargeback or showback from DCGM telemetry.
  • Build and own the reusable platform blueprint: reference architecture, Terraform and Helm modules, GitOps patterns, and runbooks that every engagement starts from.
  • Technically lead an offshore delivery pod — set standards, run design reviews, gate deliverables before they reach a customer.
  • Partner with the practice lead on customer discovery, solution design, and technical escalation; lead design sessions independently as the practice scales.
What you need
  • 7+ years platform engineering, with 4+ on Kubernetes in production; OpenShift experience valued.
  • Demonstrated GPU workload orchestration — GPU Operator, MIG, device plugin, GPU scheduling policy on real multi-node clusters.
  • Real multi-tenancy design experience: isolation, quota, RBAC, network segmentation, and the failure modes each produces.
  • Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux, Ansible.
  • Experience leading distributed or offshore engineering teams through written standards rather than direct supervision.
  • Customer-facing credibility — you can whiteboard a design for a CTO and defend it under challenge.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Platform Architect (GPUaaS)
Senior GPU Platform Architect (GPUaaS)

Unify Technologies Ltd • Plano (TX)

On-site
USD 180,000 - 240,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, AI Infra
Software Engineer, AI Infra

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000