Hardware Engineer

Yotta Infrastructure

Dadri, Navi Mumbai

On-site

INR 1,500,000 - 2,100,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Yotta Infrastructure in Uttar Pradesh's data centers seeks a Senior Hardware Engineer to own end-to-end server hardware operations for AI/GPU and traditional enterprise servers.

You will diagnose faults, perform break-fix, manage spare parts, coordinate RMAs with OEMs/ODMs, and ensure 24x7 availability in mission-critical environments.

This role requires 5+ years of hands-on experience and strong collaboration with multiple technical teams to uphold SLAs and reliability.

Qualifications

  • 5+ years in server hardware operations, data center support or enterprise infrastructure.
  • Hands-on server troubleshooting and break-fix experience.
  • Strong understanding of server architecture and components.
  • Experience diagnosing/replacing GPUs, CPUs, motherboards, memory, NICs, PSUs.
  • Spare-parts management and hardware inventory experience.
  • Faulty parts handling and RMA coordination.
  • Coordination with OEMs/ODMs and hardware service partners.
  • Knowledge of incident management and SLA-driven support.
  • Experience in 24x7 mission-critical environments.
  • Ability to troubleshoot independently and coordinate with multiple teams.

Responsibilities

  • Server hardware operations, maintenance and repair for AI/GPU and enterprise servers.
  • Maintain high availability and reliability across infrastructure.
  • Monitor faults, degradation and recurring failure patterns.
  • Diagnose faults and identify components for replacement.
  • Test and restore servers after repair or maintenance.
  • Perform FRU replacements and RMA coordination.
  • Support NVIDIA GPU-based servers and HPC compute nodes.
  • Manage spare parts inventory and on-site stock.
  • Coordinate with OEMs/ODMs and service partners.
  • Document incidents, resolutions and SLAs.

Skills

Server hardware operations
Hardware troubleshooting
Break-fix operations
Spare-parts management
RMA management
OEM/ODM coordination
24x7 data center support

Education

Bachelor's degree or diploma in Computer Science/Electronics/Electrical Engineering/IT

Tools

Server hardware diagnostics tools

Job description

Job Description Hardware Engineer
Experience: 5+ Years
Role Overview:

We are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.

The ideal candidate should have strong experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams.

Key Responsibilities:
1. Server Hardware Operations & Maintenance
  • Perform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers.
  • Ensure high availability and reliability of server hardware deployed across the infrastructure.
  • Monitor and identify hardware faults, degradation and recurring failure patterns.
  • Perform hardware diagnosis and identify faulty components for replacement.
  • Ensure servers are tested and operational after repair, replacement or maintenance activities.
  • Work within defined SLAs to ensure timely resolution of hardware-related incidents.
2. Server Hardware Break-Fix
  • Handle the complete hardware break-fix lifecycle, including:
    • Fault detection and diagnosis
    • Faulty component identification
    • Spare allocation
    • Onsite component replacement
    • Server testing and service restoration
    • Faulty part removal and tagging
    • RMA coordination and closure
  • Troubleshoot and replace server Field Replaceable Units (FRUs), including:
    • GPUs
    • CPUs
    • Motherboards
    • Memory
    • NICs
    • Power Supply Units (PSUs)
    • Fans
    • Storage and other server components
3. AI/GPU and Enterprise Server Support
  • Provide hardware support for NVIDIA GPU-based servers and compute nodes.
  • Work on high-density AI and HPC server environments.
  • Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms.
  • Perform troubleshooting and replacement of GPU cards and associated server components.
  • Support conventional enterprise and CPU-based servers, including:
    • General-purpose compute servers
    • Application and database servers
    • Virtualization servers
    • High-performance compute servers
4. Spare Parts & Inventory Management
  • Maintain and manage server hardware spares required for break-fix activities.
  • Ensure proper receipt, inspection, storage and issue of server components.
  • Maintain accurate records of spare inventory and component movement.
  • Monitor availability of critical server FRUs and escalates requirements for replenishment.
  • Support inventory reconciliation of replaced, repaired and available spare components.
  • Ensure critical AI/GPU server components are available as per operational requirements.
5. Faulty Part & RMA Management
  • Identify, tag and maintain proper records of faulty or replaced components.
  • Follow the defined process for segregation and storage of faulty hardware.
  • Coordinate with OEMs/ODMs for raising and tracking RMA cases.
  • Prepare faulty components for shipment to designated OEM/ODM service centers.
  • Track repair and replacement status of faulty components.
  • Ensure repaired or replacement components are received and appropriately updated in the inventory.
  • Maintain accurate documentation to avoid unaccounted or misplaced components.
6. OEM/ODM & Vendor Coordination
  • Coordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution.
  • Follow up on delayed parts, replacement requests and unresolved hardware issues.
  • Support warranty and service-related activities.
  • Escalate critical or recurring hardware failures to the relevant internal and external teams.
  • Coordinate with central engineering teams and onsite support teams for timely issue resolution.
7. Incident Management & Documentation
  • Respond to server hardware incidents within defined response and resolution timelines.
  • Maintain detailed records of hardware faults, repairs, replacements and RMA activities.
  • Document troubleshooting steps and resolutions for recurring issues.
  • Follow established operational processes, escalation procedures and SLAs.
  • Participate in shift or standby support for critical 24x7 environments, as required.
Required Skills & Experience:
  • 5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure.
  • Strong hands-on experience in server hardware troubleshooting and break-fix operations.
  • Strong understanding of server hardware architecture and components.
  • Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs.
  • Experience with spare-parts management and hardware inventory.
  • Experience in faulty part handling and RMA management.
  • Experience coordinating with OEMs, ODMs and hardware service partners.
  • Understanding of hardware incident management and SLA-driven support environments.
  • Experience working in 24x7 mission-critical data center or enterprise infrastructure environments.
  • Ability to troubleshoot issues independently and coordinate with multiple technical teams.
Preferred Skills:

Experience with one or more of the following will be an added advantage:

  • NVIDIA GPU-based servers and compute infrastructure
  • NVIDIA H100, H200, B200, B300, GB200 or GB300 platforms
  • NVIDIA HGX/DGX architecture
  • AI/HPC server environments
  • High-density compute infrastructure
  • Supermicro
  • ASUS
  • Gigabyte
  • Dell
  • HPE
  • GPUaaS, AI Cloud or large-scale AI data center environments
Educational Qualification:

Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.

Relevant hardware, server or OEM certifications will be an added advantage.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Project Engineer
Infrastructure Project Engineer

Netweb Technologies India • Dadri, New Delhi, Faridabad District

On-site
INR 900,000 - 1,300,000
Technical Support Engineer
Technical Support Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 600,000 - 900,000
Server Engineer
Server Engineer

E2E Networks • Chennai District, Greater Noida

On-site
INR 700,000 - 1,100,000
Infra Project Engineer
Infra Project Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 700,000 - 1,200,000
Field Application Engineer
Field Application Engineer

Netweb Technologies India Ltd. • Faridabad District

Hybrid
INR 900,000 - 1,300,000
Server Administrator
Server Administrator

Digital Edge DC • Navi Mumbai

On-site
INR 1,800,000 - 2,800,000
Executive
Executive

AH International Private Limited • Jaipur

On-site
INR 250,000 - 420,000
Senior Field Applications Engineer
Senior Field Applications Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Computer Hardware Technical Support Specialist - Voice process
Computer Hardware Technical Support Specialist - Voice process

Shashwath Solution • Dadri

On-site
INR 250,000 - 400,000
Technology Evangelist
Technology Evangelist

Vvdn Technologies • Bengaluru, Gurugram District

On-site
INR 1,500,000 - 2,100,000