Solutions Architect

Rafay Systems

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Rafay Systems seeks a Solutions Architect to help customers deploy, operate, and scale AI/ML workloads on our GPU PaaS. You will collaborate with platform engineering, MLOps, data science, and infrastructure teams to design production-ready AI infrastructure on Kubernetes and GPU environments.

You will onboard customers, optimize workload performance, automate infrastructure, and act as a trusted technical advisor throughout the lifecycle.

Qualifications

  • 8+ years of experience in Solutions Architecture, DevOps, platform engineering or related fields.
  • Hands-on Kubernetes experience in production environments.
  • Proficiency with Python or Go for automation and tooling.
  • Experience with AWS, Azure, or GCP including networking and IAM.
  • Strong knowledge of Infrastructure as Code and automation tools (Terraform, Helm, GitOps, CI/CD).
  • Familiar with monitoring/observability (Prometheus, Grafana, OpenTelemetry).
  • Understanding AI/ML infrastructure incl. GPU workloads, model serving, pipelines.

Responsibilities

  • Partner with customer teams to translate AI/ML workload requirements into scalable platform architectures.
  • Design and deploy Kubernetes-based solutions for training, fine-tuning, and inference.
  • Onboard customers to the GPU PaaS platform across cloud and hybrid environments.
  • Configure networking, IAM, observability, and security integrations with enterprise systems.
  • Build and maintain automation assets (Terraform modules, Helm charts, GitOps, CI/CD).
  • Monitor production environments for GPU utilization, performance, health, and cost.
  • Support root cause analysis and remediation for customer issues.
  • Serve as day-to-day technical advisor and primary contact for assigned customers.
  • Document best practices and provide feedback to Product/Engineering.
  • Collaborate with internal teams to drive customer adoption and expansion.

Skills

Customer-facing communication
Troubleshooting
Cross-functional collaboration
Problem solving

Tools

Kubernetes
Python
Go
Terraform
Helm
GitOps
CI/CD
Prometheus
Grafana
OpenTelemetry
AWS
Azure
GCP
PyTorch
TensorFlow

Job description

We are seeking a Solutions Architect to help customers successfully deploy, operate, and scale AI/ML workloads on our GPU Platform-as-a-Service (PaaS) offering. In this customer-facing role, you will work closely with platform engineering, MLOps, data science, and infrastructure teams to design and implement production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.

You will help customers onboard to the platform, optimize workload performance, automate infrastructure, and ensure reliable operations while serving as a trusted technical advisor throughout the customer lifecycle.

Responsibilities
  • Partner with customer platform, MLOps, and data science teams to understand AI/ML workload requirements and translate them into scalable platform architectures.
  • Design and deploy Kubernetes-based solutions for model training, fine-tuning, and inference workloads.
  • Assist customers with onboarding and implementation of the GPU PaaS platform across cloud and hybrid environments.
  • Configure networking, identity management, observability, and security integrations with enterprise systems.
  • Build and maintain automation assets including Terraform modules, Helm charts, GitOps workflows, and CI/CD pipelines.
  • Monitor and troubleshoot production environments, including GPU utilization, workload performance, cluster health, and cost efficiency.
  • Support root cause analysis and remediation efforts for customer issues.
  • Serve as a technical advisor and day-to-day point of contact for assigned customers.
  • Document best practices and provide feedback to Product and Engineering teams to improve platform capabilities.
  • Collaborate with internal teams to ensure successful customer adoption and expansion.
Required Qualifications
  • 8+ years of experience in Solutions Architecture, DevOps, Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, or related fields.
  • Strong hands-on experience with Kubernetes in production environments.
  • Experience with at least one programming language such as Python or Go.
  • Experience with AWS, Azure, or GCP, including networking, IAM, and managed Kubernetes services.
  • Knowledge of Infrastructure as Code and automation tools such as Terraform, Helm, GitOps, and CI/CD platforms.
  • Familiarity with monitoring and observability technologies including Prometheus, Grafana, OpenTelemetry, or similar.
  • Understanding of AI/ML infrastructure concepts including GPU-based workloads, model serving, training pipelines, and resource optimization.
  • Strong troubleshooting, communication, and customer-facing skills.
Preferred Qualifications
  • Experience supporting enterprise customers in cloud-native environments.
  • Familiarity with AI/ML frameworks such as PyTorch and TensorFlow.
  • Experience with GPU scheduling, autoscaling, and workload optimization.
  • Understanding of multi-tenant Kubernetes environments and platform operations.
  • Experience working with MLOps or AI infrastructure platforms.
Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Solutions Architect
Principal Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Sr. Solutions Engineer
Sr. Solutions Engineer

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Senior Solutions Engineer
Senior Solutions Engineer

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Principal Solutions Architect
Principal Solutions Architect

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Technical Solutions Architect (GPU Platform)
Technical Solutions Architect (GPU Platform)

Rafay • United States

On-site
USD 180,000 - 240,000
Senior AI/ML Solutions Architect – Kubernetes & GPU PaaS
Senior AI/ML Solutions Architect – Kubernetes & GPU PaaS

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Sr. Implementation Engineer (Kubernetes/AI)
Sr. Implementation Engineer (Kubernetes/AI)

Rafay • United States

On-site
USD 120,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
Enterprise Account Executive (AI Infrastructure)
Enterprise Account Executive (AI Infrastructure)

Rafay • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Solutions Engineer: AI/ML GPU on Kubernetes
Senior Solutions Engineer: AI/ML GPU on Kubernetes

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Lead AI Infrastructure & Solutions Architect
Lead AI Infrastructure & Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000