Senior OpenStack & GPU Infrastructure Architect (Position 2) | OCI Threads | Remote

OCI Threads

United States

Remote

USD 180,000 - 240,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OCI Threads is seeking a Senior OpenStack & GPU Infrastructure Architect to lead architecture, design, and deployment of private cloud and AI infrastructure in a remote capacity. You will own high-level and low-level designs, drive IaC deployments, and validate multi-node OpenStack environments with NVIDIA GPU integration.

Responsibilities include capacity planning, configuring OpenStack core services, benchmarking NCCL, validating DCGM workflows, and producing comprehensive runbooks and

Qualifications

  • Bachelor’s degree in a related field.
  • 8+ years of professional IT infrastructure architecture and cloud engineering experience.
  • Hands-on experience designing and implementing enterprise OpenStack private cloud environments.
  • Deep technical knowledge of OpenStack compute, networking, and storage architecture.
  • Strong Linux administration and automation / IaC experience.
  • NVIDIA GPU infrastructure experience including DCGM, Fabric Manager, and NCCL benchmarking.
  • HA, disaster recovery, performance tuning, and complex troubleshooting expertise.
  • Ability to develop HLD/LLD documentation.

Responsibilities

  • Lead OpenStack architecture workshops, HLD, and LLD development.
  • Design compute, network, storage, GPU acceleration, and security infrastructures.
  • Perform capacity planning, sizing, and hardware resource forecasting.
  • Design and support hardened Linux-based OpenStack environments.
  • Lead IaC-based OpenStack deployment across control and compute nodes.
  • Configure and troubleshoot OpenStack core services (Nova, Neutron, Cinder, Glance, Placement, Horizon).
  • Design GPU workloads and container runtime integration.
  • Review firmware compatibility for enterprise hardware (BIOS/BMC/GPU/NIC).
  • Implement and validate NVIDIA driver stack, Fabric Manager, and DCGM.
  • Perform NCCL bandwidth tests and GPU burn-in across multi-node clusters.
  • Lead testing for availability, resiliency, failover, and disaster recovery.
  • Provide L3 OpenStack architecture and operational support.

Skills

OpenStack
Linux
Infrastructure as Code
NVIDIA GPU infrastructure
NCCL benchmarking
DCGM
High Availability
Disaster recovery

Education

Bachelor’s degree in Computer Science, Electrical Engineering, Information Technology, or a related field

Tools

OpenStack Nova
OpenStack Neutron
OpenStack Cinder
OpenStack Glance
OpenStack Horizon
Fabric Manager
NVIDIA DCGM
NVIDIA driver stack

Job description

Position Summary:

OCI Threads is seeking an experienced, highly technical, and accomplished Senior OpenStack & GPU Infrastructure Architect (Second Opening) to join our team in a remote capacity. In this specialized private cloud and artificial intelligence infrastructure role, you will spearhead designing, deploying, validating, and optimizing large-scale OpenStack and NVIDIA GPU computing environments. You will work closely with cross-functional software architects, data science squads, and enterprise stakeholders to architect robust compute, network, storage, and GPU acceleration platforms. Ideal candidates bring a strong academic background in computer science or electrical engineering, deep practical mastery of hardened Linux systems, NVIDIA DCGM/NCCL toolchains, and OpenStack core services, with immediate availability for remote deployment.

Detailed Job Description:

As a Senior OpenStack & GPU Infrastructure Architect at OCI Threads in a remote capacity, you will take full ownership of leading architecture workshops, developing High-Level (HLD) and Low-Level (LLD) designs, and driving infrastructure-as-code deployments. Your day-to-day responsibilities encompass capacity planning, configuring OpenStack core components (Nova, Neutron, Cinder, Glance, Placement, Horizon), integrating NVIDIA GPU/HGX hardware, and validating driver stacks, Fabric Manager, and DCGM monitoring. You will perform NCCL bandwidth testing, GPU burn-in, resiliency validations, and complex troubleshooting across multi-node private cloud clusters. Working in a fast-paced enterprise environment, you will prepare comprehensive operational runbooks, conduct technical knowledge transfer sessions, and provide tier-3 architectural support.

Key Responsibilities:
  • Lead OpenStack architecture workshops, High-Level Design (HLD), and Low-Level Design (LLD) development.
  • Design enterprise compute, network, storage, GPU acceleration, and security infrastructure architectures.
  • Perform comprehensive capacity planning, infrastructure sizing, and hardware resource forecasting.
  • Design and support hardened Linux-based OpenStack environments across control and compute planes.
  • Lead Infrastructure-as-Code (IaC) based OpenStack deployment across control and compute nodes.
  • Configure, administer, and troubleshoot OpenStack core services including Nova, Neutron, Cinder, Glance, Placement, and Horizon.
  • Design and enable high-performance GPU workloads and container runtime integration.
  • Review BIOS, BMC, GPU/HGX, NVMe, and NIC firmware compatibility across enterprise hardware.
  • Implement, configure, and validate the NVIDIA driver stack, Fabric Manager, and Data Center GPU Manager (DCGM).
  • Perform NCCL bandwidth testing, GPU burn-in, and performance validation across AI cluster nodes.
  • Lead functional, performance, resiliency, failover, and disaster recovery testing.
  • Provide L3 OpenStack architecture, platform troubleshooting, and operational support.
Required Qualifications & Skills:
  • Bachelor’s degree in Computer Science, Electrical Engineering, Information Technology, or a related field.
  • Minimum 8+ years of professional IT infrastructure architecture and cloud engineering experience.
  • Strong, proven hands‑on experience designing and implementing enterprise OpenStack private cloud environments.
  • Deep technical knowledge of OpenStack compute, networking (Neutron), and storage architecture.
  • Strong Linux infrastructure administration and automation / Infrastructure-as-Code (IaC) experience.
  • Hands‑on experience with NVIDIA GPU infrastructure, DCGM, Fabric Manager, and NCCL benchmarking.
  • Demonstrated expertise in High Availability (HA), disaster recovery, performance tuning, and complex troubleshooting.
  • Ability to independently develop comprehensive HLD, LLD, and implementation architecture documentation.
  • Professional availability for remote employment.
Nice-to-Have Skills:
  • Professional certifications such as OpenStack Administrator (COA), Red Hat Certified Architect (RHCA), Linux Foundation Certified Engineer, or NVIDIA Deep Learning / Accelerated Computing credentials.
  • Advanced familiarity with container orchestration platforms (Kubernetes) and cloud-native storage (Ceph).
  • Experience integrating InfiniBand and high-speed RoCE (RDMA over Converged Ethernet) network fabrics.
  • Prior working experience in artificial intelligence infrastructure startups, hyperscale cloud providers, or enterprise systems integrators.
  • Contributions to open-source OpenStack projects or Linux kernel development.
Application Information:
  • Salary/Rate: Competitive and commensurate with experience
  • Deadline: Open until filled
  • Notice Period: Immediate to short notice preferred
  • Contract Duration: Permanent / Full-Time (Remote)
Recruitment Pro Tip:

When applying for senior OpenStack and GPU architecture roles, ensure your resume explicitly highlights your experience with NVIDIA DCGM/NCCL validation, automated IaC deployments, and multi-node private cloud sizing.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior OpenStack & GPU Infrastructure Architect | OCI Threads | Remote
Senior OpenStack & GPU Infrastructure Architect | OCI Threads | Remote

OCI Threads • United States

Remote
USD 140,000 - 220,000
Remote Senior OpenStack & GPU Infra Architect
Remote Senior OpenStack & GPU Infra Architect

OCI Threads • United States

Remote
USD 180,000 - 240,000
Senior OpenStack & GPU Infra Architect - Remote
Senior OpenStack & GPU Infra Architect - Remote

OCI Threads • United States

Remote
USD 140,000 - 220,000
Senior HPC Architect, Automation and At-Scale Deployment
Senior HPC Architect, Automation and At-Scale Deployment

NVIDIA • New Mexico

On-site
USD 184,000 - 287,500
Senior HPC Architect, Automation and At-Scale Deployment
Senior HPC Architect, Automation and At-Scale Deployment

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Health and wellness program
Equity
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior HPC Architect, Automation and At-Scale Deployment
Senior HPC Architect, Automation and At-Scale Deployment

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Inclusive work environment
Comprehensive benefits
GPU Cluster Architect
GPU Cluster Architect

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3
Infrastructure Solutions Architect - OEM
Infrastructure Solutions Architect - OEM

NVIDIA • Virginia (MN)

On-site
USD 152,000 - 288,000
Equity
Benefits