Principal Solutions Architect

Rafay Systems

United States

Hybrid

USD 190,000 - 280,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options

Job summary

Rafay Systems seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on its GPU Platform-as-a-Service offering. This hybrid leadership role combines hands-on architecture with building and growing a team of Solutions Architects.

You will guide AI infrastructure strategy, collaborate with platform engineering and data science teams, and drive production-ready GPU/ML deployments on Kubernetes.

Qualifications

  • 8+ years in infrastructure, platform, or solutions engineering.
  • 3+ years AI/ML infrastructure or MLOps.
  • 2+ years direct people-management experience.
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers.
  • Deep Kubernetes expertise including cluster lifecycle and RBAC.
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200).
  • Proficiency with distributed training concepts (NCCL, tensor parallelism).
  • Experience with LLM inference serving and optimization.
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics.
  • Strong scripting and automation skills (Python, Bash, Go).
  • Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders.
  • Experience with AWS, Azure, or GCP platforms.
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry.
  • Understanding of GPU-based workloads and model serving.
  • Proven troubleshooting capabilities for infrastructure issues.
  • Excellent communication, coaching, and customer-facing skills.

Responsibilities

  • Recruit, hire, and onboard Solutions Architects as the team scales.
  • Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support.
  • Set individual and team goals; conduct regular 1:1s and performance reviews.
  • Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning.
  • Foster an inclusive, high-performing team culture aligned with Rafay's values.
  • Manage team capacity, prioritization, and staffing against customer and project demand.
  • Partner with sales, engineering, and executive leadership on hiring plans and team structure.
  • Mentor and upskill both direct reports and junior team members across the broader organization.
  • Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines.
  • Develop reference architectures for GPU cluster deployment and LLM serving infrastructure.
  • Evaluate inference serving frameworks including vLLM, TGI, and Triton.
  • Advise on GPU fabric topology options for distributed training scenarios.
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF.
  • Translate infrastructure requirements into actionable platform designs.
  • Deliver technical presentations, workshops, and proof-of-concept engagements.
  • Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling.
  • Partner with customer stakeholders to understand workload requirements.
  • Architect networking, identity management, observability, and security integrations.
  • Monitor and troubleshoot production environments for GPU utilization and cluster health.
  • Lead root cause analysis for complex customer issues.
  • Document reference architectures and implementation best practices.

Skills

Infra/Platform experience
AI/ML infra
People management
Kubernetes expertise
NVIDIA GPUs
Distributed training
LLM serving
Prometheus
Grafana
Cloud platforms
Scripting languages
Communication

Tools

NVIDIA GPU Operator
Slurm
PyTorch
TensorFlow
Run:AI

Job description

Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.

Key Responsibilities
  • Recruit, hire, and onboard Solutions Architects as the team scales
  • Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support
  • Set individual and team goals; conduct regular 1:1s and performance reviews
  • Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning
  • Foster an inclusive, high-performing team culture aligned with Rafay's values
  • Manage team capacity, prioritization, and staffing against customer and project demand
  • Partner with sales, engineering, and executive leadership on hiring plans and team structure
  • Mentor and upskill both direct reports and junior team members across the broader organization
Technical & Customer-Facing Responsibilities
  • Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines
  • Develop reference architectures for GPU cluster deployment and LLM serving infrastructure
  • Evaluate inference serving frameworks including vLLM, TGI, and Triton
  • Advise on GPU fabric topology options for distributed training scenarios
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF
  • Translate infrastructure requirements into actionable platform designs
  • Deliver technical presentations, workshops, and proof-of-concept engagements
  • Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling
  • Partner with customer stakeholders to understand workload requirements
  • Architect networking, identity management, observability, and security integrations
  • Monitor and troubleshoot production environments for GPU utilization and cluster health
  • Lead root cause analysis for complex customer issues
  • Document reference architectures and implementation best practices
Required Qualifications
  • 8+ years in infrastructure, platform, or solutions engineering roles
  • 3+ years focused on AI/ML infrastructure or MLOps
  • 2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture
  • Deep Kubernetes expertise including cluster lifecycle and RBAC
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Proficiency with distributed training concepts (NCCL, tensor parallelism)
  • Experience with LLM inference serving and optimization
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics
  • Strong scripting and automation skills (Python, Bash, Go preferred)
  • Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders
  • Experience with AWS, Azure, or GCP platforms
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry
  • Understanding of GPU-based workloads and model serving
  • Proven troubleshooting capabilities for infrastructure issues
  • Excellent communication, coaching, and customer-facing skills
Preferred Qualifications
  • Experience building a Solutions Architecture or technical pre-sales team from the ground up
  • Formal people-management training or leadership certification
  • Enterprise customer support experience in cloud-native environments
  • Familiarity with PyTorch and TensorFlow frameworks
  • Experience with Run:AI and Slurm
  • GPU scheduling and autoscaling expertise
  • Multi-tenant Kubernetes environment knowledge
  • MLOps platform experience
  • Technical workshop leadership experience
  • Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect)
  • Understanding of multi-tenant GPU isolation technologies
Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Solutions Architect
Principal Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Solutions Architect
Solutions Architect

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Sr. Solutions Engineer
Sr. Solutions Engineer

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Senior Solutions Engineer
Senior Solutions Engineer

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Sr. Implementation Engineer (Kubernetes/AI)
Sr. Implementation Engineer (Kubernetes/AI)

Rafay • United States

On-site
USD 120,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
Technical Solutions Architect (GPU Platform)
Technical Solutions Architect (GPU Platform)

Rafay • United States

On-site
USD 180,000 - 240,000
Lead AI Infrastructure & Solutions Architect
Lead AI Infrastructure & Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Enterprise Account Executive (AI Infrastructure)
Enterprise Account Executive (AI Infrastructure)

Rafay • San Francisco (CA)

On-site
USD 180,000 - 280,000
AI Infrastructure & Solutions Architecture Lead
AI Infrastructure & Solutions Architecture Lead

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay • San Francisco (CA)

On-site
USD 140,000 - 180,000
Stock options
Comprehensive benefits
Career mentorship