Platform Support Architect

Ddn

Sacramento (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DDN is expanding its AI platform offerings and seeks an AI Data Platform Solutions Architect to lead supportability and enablement for the AI stack on HyperPOD. You will act as a trusted technical advisor across Support, OEMs, and NVIDIA partners, combining solution architecture mindset with hands-on L3 support.

You will own triage across GPU, NVAIE services, vector databases, and storage, and develop reusable assets, runbooks, and diagnostics to accelerate problem resolution and customer

Qualifications

  • 5+ years in Linux-based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support).
  • Hands-on experience with containers and Kubernetes (Docker/containerd, Helm, Operators).
  • Experience operating GPU-accelerated workloads in production (NVIDIA GPUs, CUDA).
  • Familiarity with NVIDIA AI Enterprise components and toolchain (NIM, NeMo, Triton, TensorRT).
  • Experience with vector databases (Milvus, pgVector, OpenSearch vectors).
  • Strong understanding of RAG/GenAI workflows and how they interact with vector search and GPU inference.

Responsibilities

  • Act as primary NVIDIA AI Enterprise and vector DB expert for HyperPOD environments; guide diagnosis and design.
  • Own end-to-end triage across GPU, NVAIE, Milvus, Kubernetes, networking, and storage.
  • Diagnose performance bottlenecks in RAG and agentic AI workflows; optimize configurations.
  • Collect logs/telemetry; build reproducible defect reports for escalation to partners and engineering.
  • Author and maintain runbooks for HyperPOD covering NVAIE, Milvus, and related components.

Skills

Linux-based infrastructure
Kubernetes
GPU-accelerated workloads
NVIDIA AI Enterprise components
MLOps / GenAI pipelines
Observability tools
Communication skills

Tools

Milvus
NVIDIA GPU Operator
Triton Inference Server
Docker
Kubernetes (CSI, CNI)
Prometheus/Grafana
ELK/NetQ
NVIDIA DGX/HGX familiarity

Job description

DDN is expanding our Enterprise and Sovereign AI Solution offerings, for example Hyperpod - a turnkey NVIDIA AI Data Platform built on DDN Infinia storage, NVIDIA AI Enterprise (NVAIE), and Supermicro reference hardware, optimized for inference and RAG workloads. Our support organization is deep on storage (Infinia, EXAScaler); we are now hiring an AI platform specialist to lead supportability and enablement for the AI side of the stack – NVIDIA AI Enterprise services (NIMs, NeMo, Triton, GPU Operator, licensing), vector databases (initially Milvus), RAG/agentic workflows, and the high‑performance storage and networking fabric that underpins them.

You will be a trusted technical advisor within Support and across OEM and NVIDIA partner teams, combining the mindset of a solutions architect (architecture, reference patterns, PoCs, reusable assets) with that of a L3 support engineer. You’ll help DDN and our partners operate AI Data solutions as a cohesive AI platform, not just a collection of components.

Key Responsibilities
Platform support
  • Act as the primary NVIDIA AI Enterprise and vector database solutions expert for HyperPOD customer environments, bringing deep knowledge of NVAIE services (e.g., NIMs, NeMo, Triton, TensorRT/TensorRT‑LLM, GPU Operator, licensing/NLS) and vector databases (e.g., Milvus) to guide diagnosis, optimization, and solution design.

  • Own complex end‑to‑end triage across GPU, NVAIE services, vector DB, Kubernetes, Docker, high‑speed networking, and Infinia storage, distinguishing product defects from environmental and integration issues.

  • Diagnose and resolve performance bottlenecks in RAG and agentic AI workflows, from model selection and prompt/RAG configuration throughto vector search, GPU utilization, and data access patterns.

  • Collect and interpret logs and telemetry across Linux, containers, Kubernetes, GPU stack, vector DB, and storage/networking; build minimal repros and high‑quality defect reports for escalation to NVIDIA, vector‑DB vendors, OEMs, and internal engineering.

Runbooks, diagnostics, and supportability
  • Author and maintain support triage runbooks and checklists for HyperPOD covering NVAIE services, Milvus/vector DB, GPU stack, Docker, Kubernetes resources, and their interaction with Infinia and the network fabric.

  • Define and validate unified diagnostics bundles that capture the right logs/configs/metrics from all relevant layers (Infinia, GPUs, NVAIE, Milvus, Kubernetes, network) to enable fast problem isolation and high‑signal escalations.

  • Collaborate with observability and tools teams to shape Prometheus/Grafana/ELK/NetQ or equivalent dashboards that surface both platform health and RAG/service‑level metrics (e.g., TTFT, retrieval latency, error rates, throughput).

Enablement, PoCs, and reusable assets
  • Build hands‑on labs and PoCs that mirror customer RAG and agentic AI use cases on HyperPOD, validating supportability and capturing “known good” configurations and troubleshooting patterns.

  • Develop reusable technical assets – implementation guides, best‑practice playbooks, tuning checklists, example architectures – to accelerate time‑to‑value for customers, PS, and Support.

Design feedback, readiness, and cross‑functional leadership
  • Provide structured feedback from early field cases and PoCs into Product Management and Engineering on stack compatibility, upgrade order, rollback constraints, and observability needs for NVAIE, Milvus/cuVS, Infinia, and networking.

  • Collaborate closely with NVIDIA solutions architects, OEM architects, PS, and Support Innovation to align reference architectures and best practices with real‑world support experience.

Required Experience & Skills
Technical
  • 5+ years in Linux‑based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support) supporting production systems; 8+ years total technical experience preferred.

  • Strong hands‑on experience with containers and Kubernetes (Docker/containerd, Helm, Operators; debugging pods, DaemonSets, CSI, CNI, and ingress/load balancers).

  • Demonstrated experience operating GPU‑accelerated workloads in production:

    • NVIDIA GPUs, drivers, CUDA concepts, GPU utilization/perf triage

    • NVIDIA GPU Operator and Kubernetes‑based GPU lifecycle management

    • Familiarity with DGX / HGX or similar GPU cluster platforms.

  • Practical experience with AI storage and networking for HPC/AI clusters:

    • High‑performance storage systems (e.g., EXAScaler/Lustre, GPFS, Ceph, distributed object storage, enterprise NAS/SAN).

    • RDMA‑accelerated and/or high‑speed Ethernet/InfiniBand networking, including fabrics, switch topologies, and large‑scale deployments.

    • Hybrid cloud or cloud‑adjacent patterns (Kubernetes CSI, cloud‑native fabrics, data locality).

  • Experience with one or more vector databases (Milvus, Qdrant, Pinecone, pgVector, OpenSearch/Elasticsearch vectors, etc.), including schema design, ingestion, and operations.

  • Solid understanding of RAG and Generative AI workflows: embeddings, retrieval, reranking, prompt design, context management, and how these interplay with vector search and GPU inference at scale.

  • Familiarity with NVIDIA AI Enterprise components and toolchain, for example:

    • NVIDIA NIM inference microservices

    • NVIDIA NeMo framework / NeMo Retriever / NeMo Curator

    • Triton Inference Server, TensorRT / TensorRT‑LLM, CUDA libraries

    • NVIDIA blueprints for enterprise RAG and agentic AI.

  • Experience designing, operating, or supporting MLOps / GenAI pipelines: CI/CD for models, deployment strategies, canarying/rollback, GPU resource management, monitoring and alerting for AI services.

  • Strong diagnostic skills across Linux, containers, Kubernetes, GPUs, storage, and networking; able to quickly narrow fault domains and propose experiments or configuration changes.

Support, architecture, and stakeholder skills
  • Track record of building reusable technical assets (runbooks, KBs, implementation guides, benchmarks, PoC templates) that improve support readiness and partner/customer success.

  • Excellent communication skills, capable of clearly explaining complex AI platform topics to both engineers and executive stakeholders, internally and with partners.

Preferred Qualifications
  • Prior experience with scale‑out storage in GPU/AI environments.

  • Direct experience crafting and operating RDMA‑accelerated HPC/AI clusters at scale, including spine‑leaf or fat‑tree network designs and large switch/router deployments.

  • Hands‑on work with NVIDIA reference blueprints (Enterprise RAG, VSS, AIQ, industry‑specific blueprints) or similar enterprise AI architectures.

  • Familiarity with AI observability and responsible AI practices (guardrails, monitoring for drift/toxicity, basic understanding of regulatory considerations like GDPR/HIPAA in the context of AI systems).

  • Experience with observability stacks (Prometheus, Grafana, Loki/ELK, NetQ, etc.) tuned for AI workloads, including service‑level dashboards and SLOs.

What Success Looks Like in This Role

Within 6–12 months, a successful AI Data Platform Solutions Architect will have:

  • Become the go‑to internal expert for “how this AI and networking stack actually works in production” across Support, PS, Product, and NPI for HyperPOD.

  • Drive speed and quality of support at solution level; NVAIE, vector DB, and AI‑workflow issues through high‑quality diagnostics, architecture insight, and well‑defined “golden stack” patterns.

  • Established clear, repeatable triage and escalation patterns for AI‑side incidents that L1/L2 storage engineers can follow with confidence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Technical Product Manager, AI Data Platforms
Sr. Technical Product Manager, AI Data Platforms

DDN • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Senior AI Platform Support Architect
Senior AI Platform Support Architect

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
AI Platform Architect — HyperPOD & NVAIE Expert
AI Platform Architect — HyperPOD & NVAIE Expert

Ddn • Sacramento (CA)

On-site
USD 140,000 - 210,000
AI Platform Architect
AI Platform Architect

Cerebras • Austin (TX)

On-site
USD 180,000 - 280,000
Sr. Technical Product Manager, AI Data Platforms
Sr. Technical Product Manager, AI Data Platforms

Data Direct Networks • Santa Clara (CA)

Hybrid
USD 175,000 - 225,000
Vacation plans
Paid holidays
Bonus programs
+5
Solutions Engineer - AI Lab
Solutions Engineer - AI Lab

SHI International Corp. • Piscataway Township (NJ)

On-site
USD 100,000 - 250,000
Senior Solutions Architect, AI Infrastructure
Senior Solutions Architect, AI Infrastructure

Jobtailor • California (MO)

On-site
USD 120,000 - 160,000
AI Solutions Architect - Central Region
AI Solutions Architect - Central Region

World Wide Technology • New Home (MO)

On-site
USD 185,000 - 235,000
Health, Dental, and Vision Care
Competitive pay
401k Plan with Company Matching
+2
Rack-Scale AI Platform Architect
Rack-Scale AI Platform Architect

Cerebras • Austin (TX)

On-site
USD 180,000 - 280,000