Data Center Technician

Penguin Solutions

United States

On-site

USD 60,000 - 75,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision benefits
401k plan
Paid Time Off
Life Insurance
Employee Assistance Plan

Job summary

Penguin Solutions is seeking an on-site Data Center Professional for a large-scale AI/HPC environment in Columbus, Georgia. You will diagnose and resolve complex infrastructure issues, own solutions, and drive efficiency across compute and storage domains.

Role requires strong problem-solving, ownership mindset, and ability to guide cross-functional teams. On-site, 24x7 coverage and on-call rotations are expected, with comprehensive safety and security practices in place.

Qualifications

  • Experience in independent analysis and problem solving in a complex data center environment.
  • Ability to take ownership of solutions and drive process improvements in operations.
  • Deep knowledge of data center hardware and networking, esp. AI/HPC environments.
  • Excellent English communication to articulate risks and guidance.
  • Certifications in AI/Networking are a plus.

Responsibilities

  • Manage AI/HPC hardware and GPU servers end-to-end.
  • Troubleshoot complex server and network issues.
  • Drive SOPs and operational strategy improvements.
  • Provide guidance across engineering and infrastructure teams.
  • Automate routine diagnostics and workflows.
  • Support 24x7 shift rotations and on-call schedule.
  • Assess risks during data center power and cooling events.
  • Ensure compliance with safety and physical security standards.

Skills

Independent analysis
Problem solving
Technical guidance
English communication
Cross-functional collaboration

Education

NCA-AIIO
CompTIA ServerPlus
CompTIA Network+
CCNP

Job description

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide. Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale. Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end‑to‑end services, and industry‑leading partner solutions in a full‑stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision. At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

Job Overview

We are looking for an experienced onsite Data Center Professional to apply their specialized expertise and technical judgment in a fast‑paced, complex AI/HPC environment. Moving beyond the execution of established procedures, this role requires independent analysis and problem‑solving to address complex infrastructure issues. You will take ownership of solutions and recommendations, driving process improvements and operational efficiencies at a large‑scale data center. This position involves diagnosing compute issues at scale, evaluating options to determine the best course of action, and providing technical guidance to influence decisions across teams. You will contribute heavily to continuous improvement and operational strategy, ensuring the stability and optimization of cloud‑scale compute and storage environments. This position will be onsite in Columbus, Georgia at the customer’s data center.

Responsibilities
  • AI/HPC‑Specific Hardware Management: Apply specialized expertise to evaluate, install, and optimize high‑density GPU servers, AI accelerators, and specialized components. Own the end‑to‑end resolution of hardware issues rather than just issue identification and escalation.
  • Complex Troubleshooting & Maintenance: Exercise technical judgment and decision‑making to address complex server and network equipment issues. Perform independent analysis and run complex diagnostics to ensure high‑performance computing cluster stability.
  • Process Improvement & Asset Management: Drive process improvements and operational efficiencies for tracking high‑value AI/HPC assets (e.g., GPUs, high‑end switches) and managing the RMA lifecycle.
  • Network Infrastructure: Provide technical guidance to remediate physical layer outages for HPC cluster networks (InfiniBand, optical cables), evaluating options to determine the best course of action.
  • Risk & Impact Assessment: Identify risks, assess impacts, and recommend corrective actions during Data Center power and cooling events, focusing on high‑density HVAC and liquid cooling systems critical for AI/HPC workloads.
  • Solution Ownership: Work within the client ticketing system to analyze root causes, develop recommendations, and own solutions for Systems and Network hardware problems.
  • Operational Strategy & SOPs: Drive continuous improvement by developing, reviewing, and optimizing Standard Operating Procedures (SOPs) and operational strategies, shifting focus from solely day‑to‑day support to long‑term efficiency.
  • Cross‑functional Leadership: Provide technical guidance and influence decisions across engineering and infrastructure teams to maintain overall HPC cluster health and optimization.
  • Automation: Contribute to operational efficiencies by applying technical expertise to automate routine diagnostic tasks and streamline workflows.
  • Operations & Security: Ensure compliance with strict safety guidelines and physical security best practices, taking ownership of risk mitigation. Successfully cover 24x7 shift rotations and participate in a weekly on‑call rotation to support continuous data center operations.
Qualifications
  • Demonstrated experience in independent analysis, problem‑solving, and exercising technical judgment in a complex data center or IT infrastructure environment.
  • Proven track record of taking ownership of solutions, driving process improvements, and contributing to overall operational strategy.
  • In‑depth, specialized knowledge of data center environments, servers, and network equipment (experience with AI/HPC environments is highly preferred). Ability to provide technical guidance, evaluate complex options, and influence decisions across cross‑functional teams.
  • Exceptional ability to assess impacts, identify risks, and recommend corrective actions proactively.
  • NCA‑AIIO, CompTIA ServerPlus, CompTIA Network, or CCNP certification is a plus.
  • Excellent English communication skills to clearly articulate technical guidance, risks, and recommendations to team members and clients.
Location

Onsite in Columbus, Georgia

Travel

None

Compensation & Benefits

The base pay range that the Company reasonably expects to pay for this position in Columbus, Georgia is $60,000 - $75,000; the pay ultimately offered may vary based on business considerations, including job‑related knowledge, skills, experience, and education. The position is bonus‑eligible, and there are medical, dental, and vision benefits available. There is a 401k saving plan and other benefits, such as Paid Time Off, Life Insurance, and an Employee Assistance Plan.

Inclusion & Belonging Statement

We are committed to creating an inclusive environment that embraces differences and fosters belonging for all.

Equal Opportunity Statement

We are an Affi­rmative Action/Equal Opportunity Employer and strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, national origin, race, ethnicity, creed, gender, disability, veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Center Technician (GA, US)
Sr. Data Center Technician (GA, US)

Penguin Computing • Georgia

On-site
USD 76,000 - 94,000
Bonus eligibility
Medical, dental and vision benefits
401k plan
+3
Data Center Technician (GA, US)
Data Center Technician (GA, US)

Penguin Computing • Georgia

On-site
USD 60,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+4
Sr. Data Center Technician
Sr. Data Center Technician

sghcorp.com • Georgia

On-site
USD 76,000 - 94,000
Medical insurance
Dental insurance
Vision insurance
+5
Supervisor, Technical Operations (GA, US)
Supervisor, Technical Operations (GA, US)

Penguin Computing • Georgia

On-site
USD 95,000 - 118,000
Medical, dental, vision benefits
401k saving plan
Paid Time Off
+2
Supervisor, Technical Operations
Supervisor, Technical Operations

Penguin Solutions • United States

On-site
USD 95,000 - 118,000
Medical benefits
Dental benefits
Vision benefits
+4
Data Center Technician
Data Center Technician

sghcorp.com • Georgia

On-site
USD 65,000 - 110,000
Supervisor, Technical Operations
Supervisor, Technical Operations

sghcorp.com • Georgia

On-site
USD 95,000 - 118,000
Bonus eligible
Medical, dental, and vision benefits
401(k) plan
+2
Network Engineer
Network Engineer

Penguin Solutions • Durham (NC)

On-site
USD 94,000 - 117,000
Bonus eligible
Medical, dental, and vision benefits
401k saving plan
+3
Onsite Data Center Technician – AI/HPC Infrastructure
Onsite Data Center Technician – AI/HPC Infrastructure

sghcorp.com • Georgia

On-site
USD 65,000 - 110,000
AI/HPC Data Center Engineer (Onsite)
AI/HPC Data Center Engineer (Onsite)

Penguin Computing • Georgia

On-site
USD 60,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+4