Lead AI Infrastructure & Solutions Architect

Rafay

United States

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Rafay seeks a Principal Solutions Architect to enable enterprise customers deploying and scaling AI/ML workloads on Rafay’s GPU Platform-as-a-Service. The role combines hybrid leadership with hands-on architecture, building a team of Solutions Architects across platform engineering, MLOps, data science, and infrastructure.

You will design production-ready AI infrastructure on Kubernetes and GPU ecosystems, mentor engineers, and partner with sales and executives to drive scalable customer success.

Qualifications

  • 10+ years in infrastructure, platform, or solutions engineering roles.
  • 3+ years focused on AI/ML infrastructure or MLOps.
  • 2+ years direct people management, hiring and performance reviews.
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers.
  • Deep Kubernetes expertise including cluster lifecycle and RBAC.
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200).
  • Proficiency with distributed training concepts (NCCL, tensor parallelism).
  • Experience with LLM inference serving and optimization.
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics.
  • Strong scripting and automation skills (Python, Bash, Go).
  • Ability to explain complex concepts to diverse audiences and exec stakeholders.
  • Experience with AWS, Azure, or GCP platforms.
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry.
  • Understanding of GPU-based workloads and model serving.
  • Proven troubleshooting capabilities for infra issues.

Responsibilities

  • Recruit, hire, and onboard Solutions Architects as the team scales.
  • Directly manage a team of Solutions Architects with coaching and support.
  • Set goals; conduct 1:1s and performance reviews.
  • Own career development planning for direct reports and succession planning.
  • Foster an inclusive, high-performing team culture aligned with Rafay’s values.
  • Manage team capacity, prioritization, and staffing against demand.
  • Partner with sales, engineering, and executives on hiring plans.
  • Mentor and upskill both direct reports and junior staff.
  • Design AI/ML platform architectures for inference, training, data pipelines.
  • Develop reference architectures for GPU clusters and LLM serving infra.
  • Evaluate inference serving frameworks (vLLM, TGI, Triton).
  • Advise on GPU fabric topology for distributed training.
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF.
  • Translate infra requirements into actionable platform designs.
  • Deliver technical presentations, workshops, and PoCs.
  • Serve as trusted advisor on AI infra strategy, costs and scaling.
  • Partner with customer stakeholders to understand workload requirements.
  • Architect networking, identity, observability, and security integrations.
  • Monitor and troubleshoot production GPU workloads and cluster health.
  • Lead root cause analysis for complex customer issues.
  • Document reference architectures and best practices.

Skills

10+ years infra/platform/solutions eng
AI/ML infra or MLOps
People management
Leadership with hands-on approach
Kubernetes RBAC and lifecycle
NVIDIA GPU infra (H100/H200/B200)
Distributed training (NCCL)
LLM inference serving
GPU Operator/MIG/SR-IOV
Scripting: Python/Bash/Go
Communication of complex tech
Cloud platforms AWS/Azure/GCP
Monitoring: Prometheus/Grafana/OpenTel
GPU workloads/model serving
Troubleshooting infra
Customer-facing collaboration

Job description

Rafay seeks a Principal Solutions Architect to enable enterprise customers deploying and scaling AI/ML workloads on Rafay’s GPU Platform-as-a-Service. The role combines hybrid leadership with hands-on architecture, building a team of Solutions Architects across platform engineering, MLOps, data science, and infrastructure.

You will design production-ready AI infrastructure on Kubernetes and GPU ecosystems, mentor engineers, and partner with sales and executives to drive scalable customer success.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure & Solutions Architecture Lead
AI Infrastructure & Solutions Architecture Lead

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Senior AI/ML Solutions Architect – Kubernetes & GPU PaaS
Senior AI/ML Solutions Architect – Kubernetes & GPU PaaS

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Senior Solutions Engineer: AI/ML GPU-Kubernetes Architect
Senior Solutions Engineer: AI/ML GPU-Kubernetes Architect

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Principal Solutions Architect
Principal Solutions Architect

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Senior Solutions Engineer: AI/ML GPU on Kubernetes
Senior Solutions Engineer: AI/ML GPU on Kubernetes

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Principal Solutions Architect
Principal Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Solutions Architect
Solutions Architect

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Senior Platform Architect - Cloud & AI Infrastructure
Senior Platform Architect - Cloud & AI Infrastructure

Rafay Systems • Sunnyvale (CA)

On-site
USD 260,000 - 360,000
Remote GPU Platform Solutions Architect
Remote GPU Platform Solutions Architect

Rafay Systems • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Solutions Engineer
Senior Solutions Engineer

Rafay Systems • United States

On-site
USD 140,000 - 170,000