Position Summary:
OCI Threads is seeking an experienced, highly technical, and accomplished Senior OpenStack & GPU Infrastructure Architect to join our team in a remote capacity. In this specialized private cloud and artificial intelligence infrastructure role, you will spearhead designing, deploying, validating, and optimizing large-scale OpenStack and NVIDIA GPU computing environments. You will work closely with cross-functional software architects, data science squads, and enterprise stakeholders to architect robust compute, network, storage, and GPU acceleration platforms. Ideal candidates bring a strong academic background in computer science or electrical engineering, deep practical mastery of hardened Linux systems, NVIDIA DCGM/NCCL toolchains, and OpenStack core services, with immediate availability for remote deployment.
Detailed Job Description:
As a Senior OpenStack & GPU Infrastructure Architect at OCI Threads in a remote capacity, you will take full ownership of leading architecture workshops, developing High-Level (HLD) and Low-Level (LLD) designs, and driving infrastructure-as-code deployments. Your day-to-day responsibilities encompass capacity planning, configuring OpenStack core components (Nova, Neutron, Cinder, Glance, Placement, Horizon), integrating NVIDIA GPU/HGX hardware, and validating driver stacks, Fabric Manager, and DCGM monitoring. You will perform NCCL bandwidth testing, GPU burn-in, resiliency validations, and complex troubleshooting across multi-node private cloud clusters. Working in a fast-paced enterprise environment, you will prepare comprehensive operational runbooks, conduct technical knowledge transfer sessions, and provide tier-3 architectural support.
Key Responsibilities:
- Lead OpenStack architecture workshops, High-Level Design (HLD), and Low-Level Design (LLD) development.
- Design enterprise compute, network, storage, GPU acceleration, and security infrastructure architectures.
- Perform comprehensive capacity planning, infrastructure sizing, and hardware resource forecasting.
- Design and support hardened Linux-based OpenStack environments across control and compute planes.
- Lead Infrastructure-as-Code (IaC) based OpenStack deployment across control and compute nodes.
- Configure, administer, and troubleshoot OpenStack core services including Nova, Neutron, Cinder, Glance, Placement, and Horizon.
- Design and enable high-performance GPU workloads and container runtime integration.
- Review BIOS, BMC, GPU/HGX, NVMe, and NIC firmware compatibility across enterprise hardware.
- Implement, configure, and validate the NVIDIA driver stack, Fabric Manager, and Data Center GPU Manager (DCGM).
- Perform NCCL bandwidth testing, GPU burn-in, and performance validation across AI cluster nodes.
- Lead functional, performance, resiliency, failover, and disaster recovery testing.
- Provide L3 OpenStack architecture, platform troubleshooting, and operational support.
Required Qualifications & Skills:
- Bachelor’s degree in Computer Science, Electrical Engineering, Information Technology, or a related field.
- Minimum 8+ years of professional IT infrastructure architecture and cloud engineering experience.
- Strong, proven hands-on experience designing and implementing enterprise OpenStack private cloud environments.
- Deep technical knowledge of OpenStack compute, networking (Neutron), and storage architecture.
- Strong Linux infrastructure administration and automation / Infrastructure-as-Code (IaC) experience.
- Hands-on experience with NVIDIA GPU infrastructure, DCGM, Fabric Manager, and NCCL benchmarking.
- Demonstrated expertise in High Availability (HA), disaster recovery, performance tuning, and complex troubleshooting.
- Ability to independently develop comprehensive HLD, LLD, and implementation architecture documentation.
- Professional availability for remote employment.
Nice-to-Have Skills:
- Professional certifications such as OpenStack Administrator (COA), Red Hat Certified Architect (RHCA), Linux Foundation Certified Engineer, or NVIDIA Deep Learning / Accelerated Computing credentials.
- Advanced familiarity with container orchestration platforms (Kubernetes) and cloud-native storage (Ceph).
- Experience integrating InfiniBand and high-speed RoCE (RDMA over Converged Ethernet) network fabrics.
- Prior working experience in artificial intelligence infrastructure startups, hyperscale cloud providers, or enterprise systems integrators.
- Contributions to open-source OpenStack projects or Linux kernel development.
Application Information:
- Recruiter: OCI Threads Recruitment Team
- Contact Name: Unspecified (OCI Threads Talent Acquisition)
- Email: careers@oci-threads.com
- Phone: Unspecified
- Application URL: Unspecified
- Salary/Rate: Competitive and commensurate with experience
- Deadline: Open until filled
- Notice Period: Immediate to short notice preferred
- Contract Duration: Permanent / Full-Time (Remote)
Recruitment Pro Tip:
When applying for senior OpenStack and GPU architecture roles, ensure your resume explicitly highlights your experience with NVIDIA DCGM/NCCL validation, automated IaC deployments, and multi-node private cloud sizing.