Data Center Technician (GA, US)

Penguin Computing

Georgia

On-site

USD 60,000 - 75,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
401k plan
Paid Time Off
Life Insurance
Employee Assistance Plan

Job summary

Penguin Solutions, onsite in Columbus, Georgia, seeks an experienced Data Center Professional for a fast-paced AI/HPC environment. You will diagnose compute and network issues at scale, own end-to-end resolutions, and drive improvements across HPC compute and storage environments.

The role emphasizes independent analysis, problem-solving, and guidance across teams, with opportunities to automate tasks, contribute to SOPs, and support 24x7 operations. Onsite position with comprehensive benefits.

Qualifications

  • Independent analysis and problem-solving in complex data center/IT infrastructure environments.
  • Proven track record of owning solutions and driving improvements.
  • Ability to influence decisions across cross-functional teams.
  • Strong English communication to articulate guidance and risks.

Responsibilities

  • AI/HPC hardware management: evaluate, install, and optimize GPU servers and accelerators.
  • Diagnose and troubleshoot complex server and network issues.
  • Drive process improvements and asset management for high-value AI/HPC equipment.
  • Provide guidance to remediate HPC network infrastructure outages.
  • Own diagnosis, recommendations, and solutions via ticketing system.
  • Develop and optimize SOPs and operational strategies for efficiency.
  • Collaborate with engineering and infrastructure teams to maintain HPC health.
  • Automate routine diagnostics and streamline workflows.
  • Support 24x7 operations including on-call rotations.

Skills

Independent analysis
Problem-solving
Technical judgment
Cross-functional leadership
English communication

Job description

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.

Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.

Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.

At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

Job Overview

We are looking for an experienced onsite Data Center Professional to apply their specialized expertise and technical judgment in a fast-paced, complex AI/HPC environment. Moving beyond the execution of established procedures, this role requires independent analysis and problem-solving to address complex infrastructure issues. You will take ownership of solutions and recommendations, driving process improvements and operational efficiencies at a large-scale data center.
This position involves diagnosing compute issues at scale, evaluating options to determine the best course of action, and providing technical guidance to influence decisions across teams. You will contribute heavily to continuous improvement and operational strategy, ensuring the stability and optimization of cloud-scale compute and storage environments

This position will be onsite in Columbus, Georgia at the customer’s data center.

Responsibilities
  • AI/HPC-Specific Hardware Management: Apply specialized expertise to evaluate, install, and optimize high-density GPU servers, AI accelerators, and specialized components. Own the end-to-end resolution of hardware issues rather than just issue identification and escalation.
  • Complex Troubleshooting & Maintenance: Exercise technical judgment and decision-making to address complex server and network equipment issues. Perform independent analysis and run complex diagnostics to ensure high-performance computing cluster stability.
  • Process Improvement & Asset Management: Drive process improvements and operational efficiencies for tracking high-value AI/HPC assets (e.g., GPUs, high-end switches) and managing the RMA lifecycle.
  • Network Infrastructure: Provide technical guidance to remediate physical layer outages for HPC cluster networks (InfiniBand, optical cables), evaluating options to determine the best course of action.
  • Risk & Impact Assessment: Identify risks, assess impacts, and recommend corrective actions during Data Center power and cooling events, focusing on high-density HVAC and liquid cooling systems critical for AI/HPC workloads.
  • Solution Ownership: Work within the client ticketing system to analyze root causes, develop recommendations, and own solutions for Systems and Network hardware problems.
  • Operational Strategy & SOPs: Drive continuous improvement by developing, reviewing, and optimizing Standard Operating Procedures (SOPs) and operational strategies, shifting focus from solely day-to-day support to long-term efficiency.
  • Cross-functional Leadership: Provide technical guidance and influence decisions across engineering and infrastructure teams to maintain overall HPC cluster health and optimization.
  • Automation: Contribute to operational efficiencies by applying technical expertise to automate routine diagnostic tasks and streamline workflows.
  • Operations & Security: Ensure compliance with strict safety guidelines and physical security best practices, taking ownership of risk mitigation. Successfully cover 24x7 shift rotations and participate in a weekly on-call rotation to support continuous data center operations.
Qualifications
  • Demonstrated experience in independent analysis, problem-solving, and exercising technical judgment in a complex data center or IT infrastructure environment.
  • Proven track record of taking ownership of solutions, driving process improvements, and contributing to overall operational strategy.
  • In-depth, specialized knowledge of data center environments, servers, and network equipment (experience with AI/HPC environments is highly preferred).
  • Ability to provide technical guidance, evaluate complex options, and influence decisions across cross-functional teams.
  • Exceptional ability to assess impacts, identify risks, and recommend corrective actions proactively.
  • NCA-AIIO, CompTIA ServerPlus, CompTIA Network, or CCNP certification is a plus.
  • Excellent English communication skills to clearly articulate technical guidance, risks, and recommendations to team members and clients.
Location

Onsite in Columbus, Georgia

Travel

None

Compensation & Benefits

The base pay range that the Company reasonably expects to pay for this position in Columbus, Georgia is $60,000 - $75,000; the pay ultimately offered may vary based on business considerations, including job-related knowledge, skills, experience, and education. The position is bonus-eligible, and there are medical, dental, and vision benefits available. There is a 401k saving plan and other benefits, such as Paid Time Off, Life Insurance, and an Employee Assistance Plan.

Inclusion & Belonging Statement

We are committed to creating an inclusive environment that embraces differences and fosters belonging for all.

Equal Opportunity Statement

We are an Afffulative Action/Equal Opportunity Employer and strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, national origin, race, ethnicity, creed, gender, disability, veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Technician
Data Center Technician

Penguin Solutions • United States

On-site
USD 60,000 - 75,000
Medical, dental, vision benefits
401k plan
Paid Time Off
+2
Sr. Data Center Technician
Sr. Data Center Technician

sghcorp.com • Georgia

On-site
USD 76,000 - 94,000
Medical insurance
Dental insurance
Vision insurance
+5
Sr. Data Center Technician (GA, US)
Sr. Data Center Technician (GA, US)

Penguin Computing • Georgia

On-site
USD 76,000 - 94,000
Bonus eligibility
Medical, dental and vision benefits
401k plan
+3
Supervisor, Technical Operations
Supervisor, Technical Operations

Penguin Solutions • United States

On-site
USD 95,000 - 118,000
Medical benefits
Dental benefits
Vision benefits
+4
Supervisor, Technical Operations (GA, US)
Supervisor, Technical Operations (GA, US)

Penguin Computing • Georgia

On-site
USD 95,000 - 118,000
Medical, dental, vision benefits
401k saving plan
Paid Time Off
+2
Supervisor, Technical Operations
Supervisor, Technical Operations

sghcorp.com • Georgia

On-site
USD 95,000 - 118,000
Bonus eligible
Medical, dental, and vision benefits
401(k) plan
+2
Data Center Technician
Data Center Technician

sghcorp.com • Georgia

On-site
USD 65,000 - 110,000
Onsite Data Center Technician – AI/HPC Infrastructure
Onsite Data Center Technician – AI/HPC Infrastructure

sghcorp.com • Georgia

On-site
USD 65,000 - 110,000
AI/HPC Data Center Engineer (Onsite)
AI/HPC Data Center Engineer (Onsite)

Penguin Computing • Georgia

On-site
USD 60,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+4
Senior Onsite Data Center Engineer — AI/HPC Ops
Senior Onsite Data Center Engineer — AI/HPC Ops

sghcorp.com • Georgia

On-site
USD 76,000 - 94,000
Medical insurance
Dental insurance
Vision insurance
+5