The Data Center Operations Technician II - GPU Specialist is responsible for supporting and maintaining highly available GPU-based compute environments, engineering labs, and data center infrastructure. This role partners closely with hardware, software, QA, and systems engineering teams to deploy, troubleshoot, and optimize next-generation computing platforms. The ideal candidate combines strong data center operations experience with a deep understanding of GPU technologies, server hardware, Linux/Windows administration, and large-scale test infrastructure.
Key Responsibilities
Compute Farm & Infrastructure Operations
- Manage and maintain a high-performance compute farm consisting of builders, packagers, testers, and supporting infrastructure.
- Monitor system health, availability, and performance to ensure operational excellence and SLA compliance.
- Lead system recovery efforts and incident response activities to minimize downtime and restore services quickly.
- Support deployment, configuration, and lifecycle management of GPU servers, workstations, and test systems.
- Perform rack, stack, cabling, hardware installation, and equipment decommissioning activities within the data center.
Engineering Support
- Collaborate closely with system architects, hardware engineers, software engineers, QA teams, and platform operations teams to develop, test, debug, and release next-generation products.
- Troubleshoot hardware, software, networking, and infrastructure issues impacting engineering and validation environments.
- Provide technical support for GPU systems, PCBs, servers, storage systems, and network-connected devices.
- Assist engineering teams with validation, benchmarking, and deployment activities for new technologies and platforms.
Process Improvement & Documentation
- Gather operational metrics and performance data to identify trends, risks, and improvement opportunities.
- Develop, maintain, and enhance Standard Operating Procedures (SOPs), runbooks, and technical documentation.
- Drive continuous improvement initiatives that increase availability, throughput, operational efficiency, and test accuracy.
- Participate in change management activities and ensure documentation is kept current.
Systems Administration & Automation
- Support and troubleshoot Linux, Windows, and macOS environments.
- Utilize scripting and automation tools to streamline operational tasks and improve scalability.
- Maintain accurate asset and infrastructure records using DCIM systems.
- Assist with infrastructure automation and configuration management initiatives.
Required Qualifications
- Associate's degree or Bachelor's degree in Engineering, Information Technology, Computer Science, or a related technical field; equivalent experience will be considered.
- 5+ years of experience supporting data center operations, engineering labs, high-performance computing environments, or related technical infrastructure.
- Experience working with GPU-based systems, PCBs, servers, and large-scale system deployments.
- Proficiency with DCIM platforms such as Nautobot or similar infrastructure management tools.
- Experience with scripting and automation technologies including Shell, Python, and Ansible.
- Working knowledge of networking fundamentals and protocols including: TCP/IP DNS NFS SSL/TLS
- Experience administering and troubleshooting: Linux Windows macOS
- Strong troubleshooting and problem-solving skills across hardware, operating systems, networking, and infrastructure.
- Excellent written and verbal communication skills with the ability to present technical concepts to non-technical audiences.
- Strong teamwork skills and the ability to work effectively in cross-functional engineering environments.
Preferred Qualifications
- Experience managing High Performance Computing (HPC) environments.
- Experience utilizing cluster management and workload scheduling platforms such as: Bright Cluster Manager (BCM) Slurm
- Industry certifications such as CCNA or equivalent networking certifications.
- Advanced Windows and Linux systems administration experience.
- Understanding of modern data center architecture, including: Compute infrastructure Storage platforms Networking systems
- Knowledge of data center facilities infrastructure with emphasis on liquid-cooled environments.
- Experience supporting AI, machine learning, or GPU-intensive workloads.
- Strong mechanical aptitude and comfort performing hands-on hardware installation, maintenance, and repair tasks.
Core Competencies
- Data Center Operations
- GPU Infrastructure Management
- Linux & Windows Administration
- Hardware Troubleshooting
- Network Fundamentals
- Automation & Scripting
- HPC Cluster Support
- Incident Response
- Documentation & SOP Development
- Cross-Functional Collaboration
- Continuous Improvement
- Customer & Engineering Partner Support
This position may require access to hardware, software, technology, or technical data subject to U.S. export control laws, including the Export Administration Regulations (EAR) and, where applicable, the International Traffic in Arms Regulations (ITAR). Any offer, assignment, or continued access to controlled items is contingent upon the company’s determination that the individual is legally authorized to access such items or that any required government authorization can be obtained.