Principal Solutions Architect

Rafay

United States

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Rafay seeks a Principal Solutions Architect to enable enterprise customers deploying and scaling AI/ML workloads on Rafay’s GPU Platform-as-a-Service. The role combines hybrid leadership with hands-on architecture, building a team of Solutions Architects across platform engineering, MLOps, data science, and infrastructure.

You will design production-ready AI infrastructure on Kubernetes and GPU ecosystems, mentor engineers, and partner with sales and executives to drive scalable customer success.

Qualifications

  • 10+ years in infrastructure, platform, or solutions engineering roles.
  • 3+ years focused on AI/ML infrastructure or MLOps.
  • 2+ years direct people management, hiring and performance reviews.
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers.
  • Deep Kubernetes expertise including cluster lifecycle and RBAC.
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200).
  • Proficiency with distributed training concepts (NCCL, tensor parallelism).
  • Experience with LLM inference serving and optimization.
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics.
  • Strong scripting and automation skills (Python, Bash, Go).
  • Ability to explain complex concepts to diverse audiences and exec stakeholders.
  • Experience with AWS, Azure, or GCP platforms.
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry.
  • Understanding of GPU-based workloads and model serving.
  • Proven troubleshooting capabilities for infra issues.

Responsibilities

  • Recruit, hire, and onboard Solutions Architects as the team scales.
  • Directly manage a team of Solutions Architects with coaching and support.
  • Set goals; conduct 1:1s and performance reviews.
  • Own career development planning for direct reports and succession planning.
  • Foster an inclusive, high-performing team culture aligned with Rafay’s values.
  • Manage team capacity, prioritization, and staffing against demand.
  • Partner with sales, engineering, and executives on hiring plans.
  • Mentor and upskill both direct reports and junior staff.
  • Design AI/ML platform architectures for inference, training, data pipelines.
  • Develop reference architectures for GPU clusters and LLM serving infra.
  • Evaluate inference serving frameworks (vLLM, TGI, Triton).
  • Advise on GPU fabric topology for distributed training.
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF.
  • Translate infra requirements into actionable platform designs.
  • Deliver technical presentations, workshops, and PoCs.
  • Serve as trusted advisor on AI infra strategy, costs and scaling.
  • Partner with customer stakeholders to understand workload requirements.
  • Architect networking, identity, observability, and security integrations.
  • Monitor and troubleshoot production GPU workloads and cluster health.
  • Lead root cause analysis for complex customer issues.
  • Document reference architectures and best practices.

Skills

10+ years infra/platform/solutions eng
AI/ML infra or MLOps
People management
Leadership with hands-on approach
Kubernetes RBAC and lifecycle
NVIDIA GPU infra (H100/H200/B200)
Distributed training (NCCL)
LLM inference serving
GPU Operator/MIG/SR-IOV
Scripting: Python/Bash/Go
Communication of complex tech
Cloud platforms AWS/Azure/GCP
Monitoring: Prometheus/Grafana/OpenTel
GPU workloads/model serving
Troubleshooting infra
Customer-facing collaboration

Job description

Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.

Key Responsibilities
  • Recruit, hire, and onboard Solutions Architects as the team scales
  • Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support
  • Set individual and team goals; conduct regular 1:1s and performance reviews
  • Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning
  • Foster an inclusive, high-performing team culture aligned with Rafay's values
  • Manage team capacity, prioritization, and staffing against customer and project demand
  • Partner with sales, engineering, and executive leadership on hiring plans and team structure
  • Mentor and upskill both direct reports and junior team members across the broader organization
Technical & Customer-Facing Responsibilities
  • Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines
  • Develop reference architectures for GPU cluster deployment and LLM serving infrastructure
  • Evaluate inference serving frameworks including vLLM, TGI, and Triton
  • Advise on GPU fabric topology options for distributed training scenarios
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF
  • Translate infrastructure requirements into actionable platform designs
  • Deliver technical presentations, workshops, and proof-of-concept engagements
  • Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling
  • Partner with customer stakeholders to understand workload requirements
  • Architect networking, identity management, observability, and security integrations
  • Monitor and troubleshoot production environments for GPU utilization and cluster health
  • Lead root cause analysis for complex customer issues
  • Document reference architectures and implementation best practices
Required Qualifications
  • 10+ years in infrastructure, platform, or solutions engineering roles
  • 3+ years focused on AI/ML infrastructure or MLOps
  • 2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture
  • Deep Kubernetes expertise including cluster lifecycle and RBAC
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Proficiency with distributed training concepts (NCCL, tensor parallelism)
  • Experience with LLM inference serving and optimization
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics
  • Strong scripting and automation skills (Python, Bash, Go preferred)
  • Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders
  • Experience with AWS, Azure, or GCP platforms
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry
  • Understanding of GPU-based workloads and model serving
  • Proven troubleshooting capabilities for infrastructure issues
  • Excellent communication, coaching, and customer-facing skills
Preferred Qualifications
  • Experience building a Solutions Architecture or technical pre-sales team from the ground up
  • Formal people-management training or leadership certification
  • Enterprise customer support experience in cloud-native environments
  • Familiarity with PyTorch and TensorFlow frameworks
  • Experience with Run:AI and Slurm
  • GPU scheduling and autoscaling expertise
  • Multi-tenant Kubernetes environment knowledge
  • MLOps platform experience
  • Technical workshop leadership experience
  • Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect)
  • Understanding of multi-tenant GPU isolation technologies
Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Solutions Architect
Principal Solutions Architect

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Solutions Architect
Solutions Architect

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Senior Solutions Engineer
Senior Solutions Engineer

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Sr. Solutions Engineer
Sr. Solutions Engineer

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Technical Solutions Architect (GPU Platform)
Technical Solutions Architect (GPU Platform)

Rafay • United States

On-site
USD 180,000 - 240,000
Sr. Implementation Engineer (Kubernetes/AI)
Sr. Implementation Engineer (Kubernetes/AI)

Rafay • United States

On-site
USD 120,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
Sr. Principal Engineer — Platform
Sr. Principal Engineer — Platform

Rafay Systems • Sunnyvale (CA)

On-site
USD 260,000 - 360,000
Lead AI Infrastructure & Solutions Architect
Lead AI Infrastructure & Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Enterprise Account Executive (AI Infrastructure)
Enterprise Account Executive (AI Infrastructure)

Rafay • San Francisco (CA)

On-site
USD 180,000 - 280,000
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay Systems • Northern (KY)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive benefits
Stock options