Sr. Data Center Technician (GA, US)

Penguin Computing

Georgia

On-site

USD 76,000 - 94,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus eligibility
Medical, dental and vision benefits
401k plan
Paid Time Off
Life Insurance
Employee Assistance Program

Job summary

Penguin Solutions in Columbus, Georgia seeks an experienced Senior Onsite Data Center Professional to own complex AI/HPC infrastructure issues, guide cross-functional teams, and drive strategic operational improvements in a fast-paced environment.

You will work onsite at the customer's data center, provide technical leadership, diagnose hardware and network challenges, and shape incident response, security, and compliance practices to ensure maximum cluster uptime.

Qualifications

  • 5+ years of experience in a data center technician or similar complex IT infrastructure environment.
  • Ability to analyze independently, exercise technical judgment, and make strategic decisions for infrastructure issues.
  • Proven ownership of technical solutions, driving process improvements and long-term operational strategy.
  • Extensive expertise installing, monitoring, and maintaining high-density data center equipment in AI/HPC environments.

Responsibilities

  • Lead advanced hardware diagnostics and root-cause analysis for AI/HPC issues, ensuring cluster stability.
  • Serve as definitive technical authority and primary escalation point for complex hardware and network challenges.
  • Drive continuous process improvements and document SOPs for operational efficiency.
  • Provide mentorship and training on ticket resolution, interventions, and safety protocols.

Skills

Independent analysis
Technical judgment
Ownership of solutions
AI/HPC infrastructure
English communication
Leadership guidance

Education

Networking/AI HPC certifications

Tools

InfiniBand expertise
High-bandwidth networks

Job description

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.

Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.

Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.

At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

Job Overview

We are looking for an experienced Senior Onsite Data Center Professional to apply specialized technical expertise and advanced judgment in a fast-paced, complex AI/HPC environment. Rather than merely executing established procedures, this role demands independent analysis, complex problem-solving, and decisive action to address large-scale infrastructure challenges. You will take full ownership of technical solutions and strategic recommendations, serving as a primary authority for issue resolution rather than simply escalating problems. In this senior capacity, you will drive process optimizations, provide critical technical guidance to cross-functional teams, and actively shape operational strategies to support cloud-scale compute and storage environments. Adaptability, advanced technical judgment, and the ability to influence outcomes are central to your success in this role.

This position will be onsite in Columbus, Georgia at the customer’s data center.

Responsibilities
  • Advanced Diagnostics & Problem Solving: Exercise independent technical judgment to lead advanced hardware diagnostics, isolate root causes, and own the resolution of complex AI/HPC issues (e.g., GPU kernel hangs, interconnect anomalies) ensuring optimal cluster stability.
  • Technical Authority & Ownership: Serve as the definitive technical authority and primary escalation point within the client ticketing system, taking full ownership of complex hardware and network challenges to develop solutions rather than merely escalating them.
  • Continuous Process Improvement: Drive operational efficiencies by evaluating current workflows and implementing systemic process improvements. Lead the documentation and continuous refinement of operational strategies and Standard Operating Procedures (SOPs).
  • Technical Guidance & Mentorship: Provide expert technical guidance, mentorship, and training to junior staff on complex ticket resolution, physical interventions, and safety protocols, heavily influencing team development and decisions.
  • Risk Assessment & RCA: Identify risks, assess systemic impacts, and lead cross-functional Root Cause Analysis (RCA) investigations for recurrent hardware or facility failures, recommending and implementing corrective actions to prevent future outages.
  • Cross-Functional Collaboration: Collaborate directly with infrastructure engineering teams to maintain overarching cluster health, apply specialized expertise to optimize node uptime, and guide the execution of strategic data center maintenance.
  • Vendor Strategy & Management: Serve as the primary technical liaison with third-party vendors, using technical judgment to evaluate options, manage advanced RMA escalations, and ensure SLA compliance for hardware replacements.
  • Advanced Network Remediation: Evaluate complex fabric topologies and remediate physical layer outages across HPC cluster networks, applying deep expertise in InfiniBand and high-bandwidth optical networks.
  • Strategic Incident Response: Respond decisively to critical facility, network, and server events, evaluating impacts and ensuring the physical environment aligns with strict AI/HPC workload requirements (including after-hours support and on-call rotation).
  • Security & Compliance Leadership: Enforce physical Security Best Practices, safety guidelines, and compliance standards, proactively identifying vulnerabilities and recommending operational enhancements.
Qualifications
  • 5+ years of experience as a data center technician or in a similar complex IT infrastructure environment.
  • Demonstrated capability in independent analysis, exercising technical judgment, and making strategic decisions to address complex infrastructure issues.
  • Proven track record of taking full ownership of technical solutions, driving process improvements, and contributing to long-term operational strategy.
  • Extensive specialized expertise in installing, monitoring, and maintaining high-density data center equipment, particularly in AI/HPC environments.
  • Strong ability to provide technical guidance, evaluate complex options, and influence outcomes across engineering and support teams.
  • Exceptional skills in identifying risks, assessing broad impacts, and successfully implementing corrective actions.
  • Excellent English communication skills to clearly articulate complex technical guidance, risks, and strategies to stakeholders and clients.
  • NCA-AIIO, CompTIA ServerPlus, CompTIA Network, or CCNP certification is a plus.
Location

Onsite in Columbus, Georgia

Travel

None

Compensation & Benefits

The base pay range that the Company reasonably expects to pay for this position in Columbus, Georgia is $76,000 - $94,000; the pay ultimately offered may vary based on business considerations, including job-related knowledge, skills, experience, and education. The position is bonus-eligible, and there are medical, dental, and vision benefits available. There is a 401k saving plan and other benefits, such as Paid Time Off, Life Insurance, and an Employee Assistance Plan.

Inclusion & Belonging Statement

We are committed to creating an inclusive environment that embraces differences and fosters belonging for all.

Equal Opportunity Statement

We are an Aff

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Center Technician
Sr. Data Center Technician

sghcorp.com • Georgia

On-site
USD 76,000 - 94,000
Medical insurance
Dental insurance
Vision insurance
+5
Data Center Technician (GA, US)
Data Center Technician (GA, US)

Penguin Computing • Georgia

On-site
USD 60,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+4
Supervisor, Technical Operations (GA, US)
Supervisor, Technical Operations (GA, US)

Penguin Computing • Georgia

On-site
USD 95,000 - 118,000
Medical, dental, vision benefits
401k saving plan
Paid Time Off
+2
Supervisor, Technical Operations
Supervisor, Technical Operations

sghcorp.com • Georgia

On-site
USD 95,000 - 118,000
Bonus eligible
Medical, dental, and vision benefits
401(k) plan
+2
Data Center Technician
Data Center Technician

sghcorp.com • Georgia

On-site
USD 65,000 - 110,000
Senior Onsite Data Center Engineer — AI/HPC Ops
Senior Onsite Data Center Engineer — AI/HPC Ops

sghcorp.com • Georgia

On-site
USD 76,000 - 94,000
Medical insurance
Dental insurance
Vision insurance
+5
AI/HPC Data Center Engineer (Onsite)
AI/HPC Data Center Engineer (Onsite)

Penguin Computing • Georgia

On-site
USD 60,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+4
Senior Onsite Data Center Engineer for AI/HPC
Senior Onsite Data Center Engineer for AI/HPC

Penguin Computing • Georgia

On-site
USD 76,000 - 94,000
Bonus eligibility
Medical, dental and vision benefits
401k plan
+3
Manager, Software Engineering (Durham, NC, US, 27703)
Manager, Software Engineering (Durham, NC, US, 27703)

Penguin Computing • Durham (NC)

Hybrid
USD 165,000 - 195,000
Hardware Systems Engineer (Maynard, MA, US, 01754)
Hardware Systems Engineer (Maynard, MA, US, 01754)

Penguin Computing • Maynard (MA)

On-site
USD 92,000 - 112,000
Medical benefits
401k Savings Plan
Paid Time Off