Senior Data Center Engineer - GPU/AI Infra

Nebius B.V.

Vineland (NJ)

On-site

USD 85,000 - 140,000

Full time

46 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career growth
Flexibility and ownership
Collaborative culture
AI projects exposure
International environment

Job summary

Nebius is seeking a Data Center Support Engineer to support large-scale bare-metal GPU infrastructure powering AI workloads. In this role you will operate, troubleshoot, and maintain production infrastructure across servers, Linux systems, networking, firmware, drivers, and data center components.

You will serve as a senior escalation point, collaborating with data center technicians and engineering teams to diagnose issues, document procedures, and drive incident response.

Qualifications

  • 5+ years of experience in data center operations or related roles.
  • Hands-on experience with bare-metal servers, Linux, hardware, and networking in production.
  • Intermediate Linux CLI proficiency for validation, logs, and diagnostics.
  • Experience with server hardware diagnostics and BMC tools (iDRAC/ILO/IPMI/Redfish).
  • Strong networking fundamentals (TCP/IP, VLANs, DNS, DHCP).

Responsibilities

  • Support and troubleshoot production bare-metal GPU infrastructure across Nebius data center environments.
  • Diagnose high-priority issues across Linux, hardware, firmware, BIOS, drivers, networking, storage, optics, cabling, and physical infrastructure.
  • Perform hardware diagnostics, component replacement, firmware updates, BIOS configuration, break/fix, and infrastructure maintenance.
  • Use Linux CLI, logs, BMC data, and health signals to identify root cause.
  • Coordinate remote troubleshooting, incident response, RCA, and operational recovery with teams.
  • Troubleshoot network connectivity issues (TCP/IP, VLANs, DNS, DHCP, switches, optics, cabling).
  • Create troubleshooting guides, runbooks, escalation notes, and knowledge base docs.
  • Use scripting or automation to improve workflows and reduce manual intervention.
  • Participate in on-call or after-hours support.

Skills

Data center ops
Infrastructure support
Systems administration
Hardware support
Linux troubleshooting
Networking fundamentals
Incident response
Documentation
Cross-functional collaboration
Automation basics

Tools

iDRAC
iLO
IPMI
Redfish
BMC tools
Jira
Confluence
Grafana
Prometheus
Kubernetes
Docker

Job description

Nebius is seeking a Data Center Support Engineer to support large-scale bare-metal GPU infrastructure powering AI workloads. In this role you will operate, troubleshoot, and maintain production infrastructure across servers, Linux systems, networking, firmware, drivers, and data center components.

You will serve as a senior escalation point, collaborating with data center technicians and engineering teams to diagnose issues, document procedures, and drive incident response.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Center GPU Infrastructure Engineer
Senior Data Center GPU Infrastructure Engineer

Nebius • New Jersey

On-site
USD 85,000 - 140,000
Career growth opportunities
Flexible work environment
Collaborative culture
+1
Senior Data Center Network Architect for AI Cloud
Senior Data Center Network Architect for AI Cloud

Nebius • United States

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan with company contribution
Paid time off
Senior GPU Data Center Deployment Engineer
Senior GPU Data Center Deployment Engineer

Nebius • Independence (MO)

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan with company match
Parental leave
+2
Data Center Support Engineer
Data Center Support Engineer

Nebius B.V. • Vineland (NJ)

On-site
USD 85,000 - 140,000
Career growth
Flexibility and ownership
Collaborative culture
+2
Data Center Technician: GPU Infra & Hardware Ops
Data Center Technician: GPU Infra & Hardware Ops

Nebius • New Jersey

On-site
Health Insurance
401(k) Plan
Parental Leave
+3
Data Center Support Engineer
Data Center Support Engineer

Nebius • New Jersey

On-site
USD 85,000 - 140,000
Career growth opportunities
Flexible work environment
Collaborative culture
+1
GPU Data Center IT Technician & Project Lead
GPU Data Center IT Technician & Project Lead

jobr.pro • Oklahoma

On-site
USD 62,000 - 94,000
Career growth and learning opportunities
Flexibility and work-life balance
Collaborative and innovative culture
+1
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
AI Data Center Project Manager & Infra Lead
AI Data Center Project Manager & Infra Lead

Nebius • Oklahoma

On-site
USD 147,000 - 184,000
Competitive pay
Career growth
Flexible work
+3
Senior GPU Deployment & Field Infra Lead
Senior GPU Deployment & Field Infra Lead

Nebius B.V. • Independence (MO)

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2