Senior Solutions Engineer

Rafay Systems

United States

On-site

USD 140,000 - 170,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Rafay Systems seeks a Senior Solutions Engineer who partners with enterprise customers to architect, deploy, and optimize Rafay platform-based cloud-native infrastructure and Kubernetes ecosystems.

The role focuses on leading customer engagements, designing scalable solutions, and guiding teams through GPU-accelerated AI/ML workloads with performance and reliability in mind. Strong communication is essential.

Qualifications

  • 10+ years of experience in customer-facing technical roles, including implementation or consulting.
  • Hands-on AI/ML engineering experience with model inference and deployment workflows.
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200 preferred).
  • Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals.
  • Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton.
  • Strong Kubernetes and cloud-native environment expertise.
  • Excellent written and verbal communication abilities.
  • Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms.
  • Linux systems administration and distributed systems knowledge.
  • Proven ability to lead technical projects and manage multiple engagements.

Responsibilities

  • Lead customer implementation engagements.
  • Gather requirements, perform architecture reviews, and design solutions.
  • Design Rafay platform deployments across public, private, and hybrid clouds.
  • Troubleshoot complex infra, networking, Kubernetes, and virtualization issues.
  • Collaborate with CS, Product, and Engineering on technical challenges.
  • Manage customer issues within SLAs and set expectations.
  • Create technical docs, runbooks, and implementation guides.
  • Mentor junior engineers and drive process improvements.
  • Stay current on Rafay capabilities and industry trends.
  • Deploy GPU-accelerated Kubernetes clusters for AI/ML workloads with GPU Operator, MIG, and SR-IOV setup.
  • Validate LLM inference stacks (vLLM, TGI, Triton).
  • Troubleshoot GPU cluster health, utilization, and GPU fabric networking.
  • Configure GPU observability with DCGM and OpenTelemetry; use Grafana dashboards.
  • Advise on distributed training and inference topology for NVIDIA GPUs.

Skills

Customer-facing experience
Technical leadership
Kubernetes
Cloud-native architecture
AI/ML engineering
Linux administration
NVIDIA GPU infrastructure
LLM inference serving

Education

Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience

Tools

NVIDIA GPU Operator
MIG
SR-IOV
DCGM
Prometheus/Grafana

Job description

Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads.

Key Responsibilities

Core Implementation Responsibilities

  • Serve as primary technical lead for customer implementation engagements
  • Partner with customers on requirements gathering, architecture reviews, and solution design
  • Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments
  • Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues
  • Collaborate with Customer Success, Product, and Engineering teams on technical challenges
  • Manage customer issues within SLAs while maintaining expectations
  • Reproduce and analyze customer-reported issues; communicate findings to internal teams
  • Develop technical documentation, implementation guides, and runbooks
  • Mentor junior engineers and contribute to process improvements
  • Stay current on Rafay platform capabilities and emerging industry trends

GPU & AI/ML Infrastructure

  • Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup
  • Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations
  • Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2)
  • Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling
  • Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)
Minimum Qualifications
  • 10+ years of experience in customer-facing technical roles, including implementation or consulting
  • Hands-on AI/ML engineering experience with model inference and deployment workflows
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals
  • Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton
  • Strong Kubernetes and cloud-native environment expertise
  • Excellent written and verbal communication abilities
  • Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms
  • Linux systems administration and distributed systems knowledge
  • Proven ability to lead technical projects and manage multiple engagements
  • Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience
Preferred Qualifications
  • Experience with Run:AI and Slurm for GPU scheduling
  • GPU autoscaling and multi-tenant GPU isolation expertise
  • Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards
  • Relevant certifications (CKA, CKAD, or NVIDIA certifications)
Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Solutions Engineer
Sr. Solutions Engineer

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Principal Solutions Architect
Principal Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Sr. Implementation Engineer (Kubernetes/AI)
Sr. Implementation Engineer (Kubernetes/AI)

Rafay • United States

On-site
USD 120,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
Solutions Architect
Solutions Architect

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Principal Solutions Architect
Principal Solutions Architect

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Technical Solutions Architect (GPU Platform)
Technical Solutions Architect (GPU Platform)

Rafay • United States

On-site
USD 180,000 - 240,000
Senior Solutions Engineer: AI/ML GPU on Kubernetes
Senior Solutions Engineer: AI/ML GPU on Kubernetes

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay Systems • Northern (KY)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive benefits
Stock options
Senior Solutions Engineer: AI/ML GPU-Kubernetes Architect
Senior Solutions Engineer: AI/ML GPU-Kubernetes Architect

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay • San Francisco (CA)

On-site
USD 140,000 - 180,000
Stock options
Comprehensive benefits
Career mentorship