Data Center Technician L2

Covestic Inc

Reno (NV)

On-site

USD 90,000 - 130,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Covestic Inc is seeking a Data Center Operations Technician II - GPU Specialist in Reno, NV to support and maintain GPU compute environments, data center infrastructure, and engineering labs.

The role collaborates with hardware, software, QA, and systems teams to deploy, troubleshoot, and optimize platforms, with emphasis on Linux/Windows admin, scripting, and automation. Hands-on hardware experience is essential.

Qualifications

  • Bachelor's or Associate degree in a technical field or equivalent experience.
  • 5+ years data center operations or HPC infrastructure experience.
  • GPU systems, PCBs, servers deployment experience.
  • Experience with DCIM platforms like Nautobot.
  • Scripting: Shell, Python, Ansible.
  • Networking basics: TCP/IP, DNS, NFS, SSL/TLS.
  • Cross-functional communication and strong problem-solving.

Responsibilities

  • Manage and maintain a high-performance compute farm consisting of builders, packagers, testers, and supporting infrastructure.
  • Monitor system health, availability, and performance to ensure operational excellence and SLA compliance.
  • Lead system recovery efforts and incident response activities to minimize downtime and restore services quickly.
  • Support deployment, configuration, and lifecycle management of GPU servers, workstations, and test systems.
  • Perform rack, stack, cabling, hardware installation, and equipment decommissioning activities within the data center.

Skills

Data Center Operations
GPU Infrastructure Management
Linux Administration
Windows Administration
Automation & Scripting
HPC Cluster Support
Incident Response
Documentation & SOP Development
Cross-Functional Collaboration
Continuous Improvement
Customer & Engineering Partner Support

Education

Associate's or Bachelor's degree in Engineering/IT/CS

Tools

Nautobot
Bright Cluster Manager
Slurm

Job description

The Data Center Operations Technician II - GPU Specialist is responsible for supporting and maintaining highly available GPU-based compute environments, engineering labs, and data center infrastructure. This role partners closely with hardware, software, QA, and systems engineering teams to deploy, troubleshoot, and optimize next-generation computing platforms. The ideal candidate combines strong data center operations experience with a deep understanding of GPU technologies, server hardware, Linux/Windows administration, and large-scale test infrastructure.

Key Responsibilities
Compute Farm & Infrastructure Operations
  • Manage and maintain a high-performance compute farm consisting of builders, packagers, testers, and supporting infrastructure.
  • Monitor system health, availability, and performance to ensure operational excellence and SLA compliance.
  • Lead system recovery efforts and incident response activities to minimize downtime and restore services quickly.
  • Support deployment, configuration, and lifecycle management of GPU servers, workstations, and test systems.
  • Perform rack, stack, cabling, hardware installation, and equipment decommissioning activities within the data center.
Engineering Support
  • Collaborate closely with system architects, hardware engineers, software engineers, QA teams, and platform operations teams to develop, test, debug, and release next-generation products.
  • Troubleshoot hardware, software, networking, and infrastructure issues impacting engineering and validation environments.
  • Provide technical support for GPU systems, PCBs, servers, storage systems, and network-connected devices.
  • Assist engineering teams with validation, benchmarking, and deployment activities for new technologies and platforms.
Process Improvement & Documentation
  • Gather operational metrics and performance data to identify trends, risks, and improvement opportunities.
  • Develop, maintain, and enhance Standard Operating Procedures (SOPs), runbooks, and technical documentation.
  • Drive continuous improvement initiatives that increase availability, throughput, operational efficiency, and test accuracy.
  • Participate in change management activities and ensure documentation is kept current.
Systems Administration & Automation
  • Support and troubleshoot Linux, Windows, and macOS environments.
  • Utilize scripting and automation tools to streamline operational tasks and improve scalability.
  • Maintain accurate asset and infrastructure records using DCIM systems.
  • Assist with infrastructure automation and configuration management initiatives.
Required Qualifications
  • Associate's degree or Bachelor's degree in Engineering, Information Technology, Computer Science, or a related technical field; equivalent experience will be considered.
  • 5+ years of experience supporting data center operations, engineering labs, high-performance computing environments, or related technical infrastructure.
  • Experience working with GPU-based systems, PCBs, servers, and large-scale system deployments.
  • Proficiency with DCIM platforms such as Nautobot or similar infrastructure management tools.
  • Experience with scripting and automation technologies including Shell, Python, and Ansible.
  • Working knowledge of networking fundamentals and protocols including: TCP/IP DNS NFS SSL/TLS
  • Experience administering and troubleshooting: Linux Windows macOS
  • Strong troubleshooting and problem-solving skills across hardware, operating systems, networking, and infrastructure.
  • Excellent written and verbal communication skills with the ability to present technical concepts to non-technical audiences.
  • Strong teamwork skills and the ability to work effectively in cross-functional engineering environments.
Preferred Qualifications
  • Experience managing High Performance Computing (HPC) environments.
  • Experience utilizing cluster management and workload scheduling platforms such as: Bright Cluster Manager (BCM) Slurm
  • Industry certifications such as CCNA or equivalent networking certifications.
  • Advanced Windows and Linux systems administration experience.
  • Understanding of modern data center architecture, including: Compute infrastructure Storage platforms Networking systems
  • Knowledge of data center facilities infrastructure with emphasis on liquid-cooled environments.
  • Experience supporting AI, machine learning, or GPU-intensive workloads.
  • Strong mechanical aptitude and comfort performing hands-on hardware installation, maintenance, and repair tasks.
Core Competencies
  • Data Center Operations
  • GPU Infrastructure Management
  • Linux & Windows Administration
  • Hardware Troubleshooting
  • Network Fundamentals
  • Automation & Scripting
  • HPC Cluster Support
  • Incident Response
  • Documentation & SOP Development
  • Cross-Functional Collaboration
  • Continuous Improvement
  • Customer & Engineering Partner Support

This position may require access to hardware, software, technology, or technical data subject to U.S. export control laws, including the Export Administration Regulations (EAR) and, where applicable, the International Traffic in Arms Regulations (ITAR). Any offer, assignment, or continued access to controlled items is contingent upon the company’s determination that the individual is legally authorized to access such items or that any required government authorization can be obtained.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Technician L2 - 3 Openings
Data Center Technician L2 - 3 Openings

Covestic Inc • Hubbard (TX)

On-site
USD 60,000 - 90,000
Data Center Technician L3
Data Center Technician L3

Covestic Inc • Town of Texas (WI), Northern (KY)

On-site
USD 90,000 - 130,000
Data Center Technician L2
Data Center Technician L2

Covestic Inc • Town of Texas (WI), Northern (KY)

On-site
USD 55,000 - 80,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Data Center Technician - AI Infrastructure
Data Center Technician - AI Infrastructure

Hamilton Barnes Associates Limited • South Carolina

On-site
USD 70,000 - 85,000
Funded training and certifications
Stable full-time schedule: Monday–Friday, 9-5
Healthcare benefits
Data Center Technician II Cleveland
Data Center Technician II Cleveland

Cirrascale Corporation • Cleveland (OH)

On-site
USD 50,000 - 70,000
Data Center Technician/ Field Engineer
Data Center Technician/ Field Engineer

STN Inc • Odessa (TX)

On-site
USD 65,000 - 90,000