Lead Solution Architect

Yotta

Hinoba-an

On-site

PHP 9,404,000 - 13,166,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Yotta Data Services is seeking a senior AI infrastructure architect to design, deploy, and operate large-scale GPU clusters for distributed training and inference. You will lead multi-node GPU systems, optimize OS/kernel configurations, and craft compute blueprints for diverse AI workloads.

The role requires deep hands-on experience with NVIDIA GPUs, InfiniBand/RDMA, and high-performance storage, with a strong focus on security, reliability, and observability across multi-tenant environments.

Qualifications

  • Deep hands-on experience with GPU systems (NVIDIA).
  • Strong background in InfiniBand/RDMA/RoCEv2 and high‑performance storage.
  • Experience designing large-scale AI infrastructure with 100+ GPUs is a plus.

Responsibilities

  • Design and deploy large-scale GPU clusters for distributed training and inference.
  • Architect multi-node GPU systems with NVLink/NVSwitch and PCIe Gen5.
  • Define OS/kernel/driver/runtime configs for AI workloads (CUDA/ROCm, NCCL, UCX, OFED).
  • Develop compute blueprints for training, fine-tuning, retrieval, and batch inference.
  • Ensure secure, scalable, high-availability AI compute clusters.
  • Implement monitoring and performance optimization (Prometheus, Grafana, DCGM).
  • Architect secure multi-tenant GPU environments with IAM/RBAC and encryption in transit.

Skills

GPU systems
InfiniBand/RDMA/RoCEv2
High-performance storage
Multi-node GPU architectures
AI frameworks familiarity
Kubernetes GPU environments
Terraform
Ansible

Education

B.E./B.Tech or equivalent

Tools

MAAS/iPXE/Metal³
NVMe/NPF fabrics

Job description

Yotta Data Services is India’s leading sovereign AI infrastructure, cloud platform and data centre services company, enabling enterprises, governments, startups, and digital platforms to build, deploy, and scale next-generation AI and digital workloads securely within India.

With hyperscale data centre campuses in Navi Mumbai and Greater Noida (Delhi NCR), advanced GPU-powered AI infrastructure, and a comprehensive ecosystem of cloud, AI, hosting, cybersecurity, and managed platform services, Yotta delivers high-performance, scalable, and compliant digital infrastructure built for the AI era.

Yotta is at the forefront of powering India’s sovereign AI and digital transformation journey through world-class infrastructure, deep technology partnerships, and fully India-hosted enterprise-grade platforms.

Job Scope

The role is responsiblefor designing, deploying, and operating large-scale AI infrastructure,including GPU compute, high-performance networking, storage, andorchestration platforms. The incumbent will lead the architecture of secure,scalable, and high-availability AI clusters to support distributed training,fine-tuning, and inference workloads, while ensuring optimal performance,reliability, and compliance

  • 10+ years in systemsengineering, network engineering, cloud infrastructure, or datacenter design .
Key Responsibilities
  • Design anddeploy large-scale GPU clusters (H100, H200, GB200 and GB300) fordistributed training and inference.
  • Architectmulti-node GPU systems using:
    • NVLink/NVSwitch
    • PCIe Gen5
  • Define OS,kernel, driver, and runtime configurations optimized for AI workloads(CUDA/ROCm, NCCL, UCX, OFED).
  • Develophigh-performance compute blueprints for diverse use cases: training,fine-tuning, retrieval, and batch inference.
High-Performance Networking
  • Architect AIfabric networks including:
    • InfiniBandHDR/NDR/XDR/SPX
    • RoCEv2 /RDMA
    • 100/200/400/800Gbps Ethernet fabrics
  • Designlow-latency, high-bandwidth topologies (fat-tree, dragonfly+,multi-plane architectures).
  • Plan and tuneinter-node communication for distributed AI training (NCCL, MPI, UCX).
  • Implementnetwork segmentation, isolation, and multi-tenant security for AIcompute clusters.
  • Architecthigh-throughput storage solutions for AI:
    • Parallelfile systems (Lustre, BeeGFS, IBM Spectrum Scale)
    • Cloud-nativehigh-performance storage (FSx for Lustre, Azure ANF, GCS Filestore HighScale)
    • NVMe,NVMe-over-Fabrics, object storage
  • Optimize datapipelines for large-scale dataset ingestion, feature extraction,checkpointing, and streaming.
Platform Integration & Orchestration
  • Integratesystems with Kubernetes GPU environments (EKS/AKS/GKE, K8s on-prem,Kueue, Volcano).
  • Designinfrastructure to support distributed training frameworks:
    • PyTorch DDP
    • DeepSpeed
    • Ray Train
    • JAX / TPUalternatives
  • Enable robustscheduling, multi-tenancy, and job orchestration.
Reliability, Monitoring & Performance Optimization
  • Implementmonitoring for GPU utilization, network telemetry, I/O performance, andcluster health (Prometheus, Grafana, DCGM, NetQ).
  • Conductperformance tuning across:
    • NIC/driverstack
    • GPU topology
    • Storagethroughput
    • Networkcongestion management (ECN, PFC, QoS)
  • Designsystems for high availability, resilience, and disaster recovery.
Security & Compliance (Infra-Level)
  • Implementhardware-level and network-level security controls—IAM, RBAC, ACLs,segmentation, encryption in transit.
  • Architectsecure multi-tenant GPU environments, including confidential computingwhere supported.
  • Ensure systemcompliance with SOC2, ISO 27001, or industry-specific securityframeworks
Must-have skill

Deep hands-on experience with:

  • GPU systems (NVIDIA)
  • InfiniBand / RDMA / RoCEv2
  • High-performance storage solutions

Strong networking background (L2/L3 switching,routing, QoS, congestion control, BGP/EVPN).

Familiarity with AI frameworks and distributed training, even if not a data scientist.

Expertise with infrastructure automation:

  • Terraform
  • Ansible
Good-to-Have Skills
  • Experience building clusters for AI training at>100 GPUs scale.
  • Familiarity with AI data engineering systems (Kafka,Spark, Ray Data).
  • Experience with bare-metal provisioning tools (MAAS,iPXE, Metal³).
  • Knowledge of GPU virtualization, MIG/partitioning, ormulti-tenant GPU scheduling.
Qualifications Criteria
  • B.E/B.Tech or equivalent degree

3 Round of Interview

Company Values

Customer Centricity

Agility

Integrity

Trust andTransparency

Happiness for all

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU System/Fabrics Architect
Senior GPU System/Fabrics Architect

NVIDIA Corporation • Hinoba-an

On-site
PHP 7,524,000 - 11,912,000
Hybrid work model
GPU Architect
GPU Architect

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 1,200,000 - 1,800,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 2,000,000 - 4,500,000
Hybrid work model
System Software Engineer - Performance Verification Infrastructure
System Software Engineer - Performance Verification Infrastructure

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 900,000 - 1,300,000
Senior DevOps Engineer (AI, GPU & Cloud Infrastructure)
Senior DevOps Engineer (AI, GPU & Cloud Infrastructure)

Wisewit Solutions Private Limited • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
GPU System Architect for Scalable AI Data Centers
GPU System Architect for Scalable AI Data Centers

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 1,200,000 - 1,800,000
Senior System Integration and Validation Engineer
Senior System Integration and Validation Engineer

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 7,524,000 - 11,285,000
Competitive benefits
Flexible time off
Continuous learning
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA Corporation • Hinoba-an

On-site
PHP 2,400,000 - 3,600,000
SOLUTION ARCHITECT - TECHNOLOGY
SOLUTION ARCHITECT - TECHNOLOGY

PeopleStrong • Hinoba-an

On-site
PHP 1,592,000 - 2,387,000
Senior Software Engineer — Infra Agent Systems Remote India Together AI India
Senior Software Engineer — Infra Agent Systems Remote India Together AI India

Neura Market • Hinoba-an

Remote
INR 3,000,000 - 5,400,000