Remote Technical Operations & Deployment Engineer (GPU/AI/Cloud Infrastructure) at Radian Arc

Radian Arc Limited

Polska

Hybrid

PLN 240,000 - 360,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Radian Arc Limited seeks a hands-on Infrastructure Engineer to deploy and operate GPU cloud deployments in data centers. You will coordinate with datacenter providers, engineers, and logistics to bring up rack-ready systems, validate network and storage, and ensure platform readiness.

The role spans hardware bring-up, network validation, and platform installation with strong emphasis on practical field work and cross-team collaboration in a fast‑growing AI infrastructure environment.

Qualifications

  • Hands-on experience deploying and maintaining datacenter infrastructure.
  • Experience with GPU, HPC, AI cloud deployments in production.
  • Familiarity with NVL72-style rack architectures and high-density compute requirements.
  • Comfort working across physical infrastructure, Linux, networking, and platform software.

Responsibilities

  • Coordinate physical deployment activities with datacenter providers and internal teams.
  • Validate rack layouts, power, cabling, and labeling before installation.
  • Bring up GPU servers, platform servers, storage nodes, and networking gear.
  • Validate BIOS, firmware, NIC, DPU, GPU, and storage components.
  • Support network validation including BGP/EVPN/VXLAN configurations.
  • Install and validate platform stack (Docker, Kubernetes, KVM/QEMU).
  • Perform acceptance tests and handover documentation.
  • Participate in on-call rotations and field escalations.

Skills

GPU datacenter deployment
Linux system administration
Networking & routing
Kubernetes / KubeVirt
NVIDIA drivers & DCGM
Redfish / IPMI
Rack cabling & power
Hardware validation
DC/Platform automation

Tools

NVIDIA CUDA drivers
KVM/QEMU
StorPool / Weka
VyOS / Networking OS
Docker / containerd
Prometheus / Grafana / Zabbix

Job description

This Full time on site position offers great opportunities for career growth.

About Radian Arc.

We're specialists in outcome-optimized AI infrastructure - deploying, orchestrating and monetizing GPU compute where data, users and demand actually meet: inside telco networks, at the edge, and in core data centers. Not a generic AI platform. Not a consultancy. We're the bridge between raw silicon and real-world results.

What impact you will have
Mission:

Install, validate, operate, and maintain regional and core GPU cloud deployments across datacenter environments. This role owns the practical deployment and operational readiness of the platform stack in the field. It bridges infrastructure engineering, datacenter operations, networking, host systems, storage, and platform operations. The engineer is responsible for taking a validated architecture and BOM and turning it into a working production environment, including physical deployment coordination, rack and cabling validation, switch and host bring‑up, firmware and BIOS validation, operating‑system installation, GPU and DPU validation, storage integration, platform stack installation, acceptance testing, and operational handover. The role is intentionally hands‑on and cross‑domain. It is not a pure datacenter technician role, and it is not a pure platform engineering role. It is the person who can work in a datacenter, understand cabling and optics, debug host and network issues, validate GPU servers, support platform installation, and coordinate with engineering when problems arise. The role is especially important as the platform evolves from regional deployments toward HGX‑based GPU systems, east‑west fabrics, and core AI infrastructure. This profile is focused on actual installation, commissioning, maintenance, and operational readiness.

What you'll do
Datacenter deployment and commissioning
  • Coordinate physical deployment activities with datacenter providers, integrators, logistics teams, and internal engineering.
  • Validate rack layouts, elevations, power feeds, airflow assumptions, cable paths, and labeling before installation.
  • Support rack‑and‑stack activities for GPU nodes, CPU nodes, storage nodes, switches, firewalls, routers, serial/OOB equipment, PDUs, and supporting infrastructure.
  • Validate fibre and copper cabling against the deployment design, including OOB, north‑south, east‑west, storage, and management networks.
  • Validate optics, transceivers, link speeds, breakout cables, port mappings, and redundancy assumptions.
  • Maintain accurate as‑built documentation, including rack elevations, cable maps, port maps, serial numbers, asset records, IP allocations, and change records.
Host and hardware bring‑up
  • Bring up GPU servers, platform servers, storage nodes, and supporting infrastructure.
  • Validate BIOS, BMC, firmware, NIC, DPU, GPU, NVMe, RAID/HBA, and platform firmware versions.
  • Configure and validate BMC access using Redfish/IPMI and OOB management networks.
  • Validate GPU visibility, PCIe topology, NUMA layout, thermals, power behavior, and hardware health.
  • Run hardware acceptance tests, burn‑in tests, GPU stress tests, network tests, and storage validation before handover.
  • Troubleshoot hardware issues across servers, GPUs, DPUs, NICs, optics, cables, disks, memory, firmware, and BIOS.
Network deployment support
  • Support deployment and validation of OOB, north‑south, storage, and east‑west networking.
  • Work with networking engineering to apply and validate switch configurations.
  • Validate BGP, ECMP, VLAN/VRF segmentation, EVPN/VXLAN where applicable, OVS/OVN integration, and routing reachability.
  • Validate VyOS routers, OOB firewalls, transit routers, Citrix NetScaler/WAF, and customer connectivity.
  • Support RoCE/RDMA fabric validation for distributed AI workloads where applicable.
  • Troubleshoot practical network issues such as link flaps, optics issues, incorrect polarity, MTU mismatches, route leaks, VLAN errors, packet loss, PFC/ECN issues, and fabric congestion.
  • Support integration with NVIDIA Cumulus / Spectrum-X environments, and assist with Cisco or SONiC‑based alternatives if those become part of the roadmap.
Platform stack installation and validation
  • Support installation and validation of the platform stack across regional and core deployments.
  • Install and validate host operating systems, kernel versions, NVIDIA drivers, Mellanox/NVIDIA OFED or inbox drivers, CUDA compatibility, Docker/containerd, KVM/QEMU, and platform agents.
  • Support CloudStack‑based deployments and Kubernetes/KubeVirt‑based deployments.
  • Validate GPU passthrough, SR‑IOV, BlueField NIC/DPU behavior, VM networking, and container networking.
  • Support Kubernetes node registration, GPU Operator validation, CSI validation, CNI validation, and node lifecycle workflows.
  • Support storage integration with StorPool, Weka, local NVMe, or other supported storage platforms.
  • Execute acceptance tests and produce deployment readiness reports.
Operational maintenance and Day‑2 support
  • Perform controlled maintenance activities such as firmware upgrades, switch upgrades, host OS updates, GPU driver updates, BIOS changes, and hardware replacements.
  • Support incident response for infrastructure issues affecting GPU nodes, hosts, networking, storage, or platform components.
  • Perform root‑cause analysis for deployment and operational failures.
  • Maintain runbooks for installation, validation, upgrade, rollback, troubleshooting, and handover.
  • Work with engineering to turn repeated operational issues into automation, better validation, or platform improvements.
  • Participate in on‑call or escalation rotations for regional and core environments where appropriate.
Platform observability and validation
  • Ensure telemetry is correctly configured for hosts, GPUs, DPUs, switches, storage, OOB devices, and platform components.
  • Validate Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, NVIDIA NetQ or equivalent telemetry sources.
  • Confirm that deployment health checks, hardware alerts, performance dashboards, and operational alarms work before production handover.
  • Support performance baseline testing for GPU, network, storage, and host layers.
  • Assist with NCP‑related validation and benchmarking.
Cross‑team coordination
  • Work closely with the Senior Director of Infrastructure Operations.
  • Work closely with Staff Network, Staff Storage, Sr Hardware/Infrastructure, Sr Platform, Sr Fleet Automation, Observability, Product Engineering, Sales Engineering, and Service Delivery roles.
  • Provide field feedback into reference architectures, BOMs, rack layouts, cabling standards, deployment playbooks, and validation procedures.
  • Coordinate with external vendors including datacenter providers, systems integrators, server vendors, storage vendors, NVIDIA, and networking vendors.
  • Act as the practical field escalation point when architecture, BOM, datacenter conditions, and platform implementation do not align.
Technical Stack

Hardware and datacenter GPU servers: L40S, RTX 6000 Pro, H200, B200/B300-class systems, HGX systems, and future NVL72-style rack-scale systems.

CPU/platform servers.

Storage nodes and JBODs.

PDUs, BMCs, serial/OOB, firewalls, routers, switches.

Rack layouts, power feeds, airflow, cold/hot aisle containment, cabling, optics.

DTC/DLC cooling.

Host and systems Ubuntu Linux.

Linux networking.

BIOS/BMC/firmware lifecycle.

Redfish, IPMI.

NVIDIA drivers, CUDA, DCGM/NVML.

Mellanox/NVIDIA NICs, BlueField DPUs.

KVM/QEMU, VFIO, PCI passthrough.

Docker/containerd.

Networking NVIDIA Spectrum/Cumulus.

VyOS.

OVS/OVN.

BGP, ECMP, VLAN, VRF, EVPN/VXLAN.

RoCE/RDMA.

SR‑IOV.

Citrix NetScaler / WAF.

OOB and break‑glass access.

Platform Kubernetes.

KubeVirt.

NVIDIA GPU Operator.

CSI/CNI integrations.

StorPool, Weka, local NVMe.

Observability Zabbix.

Prometheus.

Grafana.

Loki.

DCGM Exporter.

NVIDIA NetQ or equivalent.

Logs, metrics, hardware health, and deployment validation dashboards.

What you'll need
  • Strong hands‑on experience deploying and maintaining datacenter infrastructure.
  • Experience with GPU, HPC, AI cloud, private cloud, or high‑density compute environments, including both air‑cooled and liquid‑cooled (nice to have) deployments.
  • Familiarity with NVL72‑style rack‑scale architecture, NVLink/NVSwitch domains, in‑rack networking, high‑density power delivery, and OEM/NVIDIA validation requirements.
  • Comfortable working across physical infrastructure, Linux hosts, networking, storage, and platform software.
  • Experience bringing up servers from bare metal through production readiness.
  • Experience with rack layouts, cabling, optics, transceivers, power, OOB management, and datacenter handover.
  • Strong Linux troubleshooting skills.
  • Practical networking knowledge across VLANs, VRFs, BGP, ECMP, OVS/OVN, and routing.
  • Experience with GPU servers, NVIDIA drivers, firmware, PCIe topology, and hardware validation.
  • Strong documentation discipline and ability to produce accurate as‑built records and runbooks.
  • Ability to validate practical datacenter readiness for current and next‑generation AI infrastructure, including power density, cooling model, rack depth/width, floor loading, containment, serviceability, and maintenance access.
  • Personal qualities: Very hands‑on and practical. Comfortable working in datacenters and remotely with smart‑hands teams. Strong troubleshooting mindset across hardware, network, host, and platform layers. High attention to detail around cabling, labeling, asset records, and validation. Calm under pressure during deployment windows and customer‑impacting incidents. Able to distinguish between a temporary field workaround and a permanent engineering fix. Good communicator who can explain field issues clearly to engineering and leadership.
What we offer

Attractive compensation package reflecting your expertise and experience.

A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid‑friendly approach. You'll be part of a fast‑growing scale‑up with a mission to make a positive impact, offering an exciting career evolution. Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Datacenter GPU Deployment Engineer
Datacenter GPU Deployment Engineer

Radian Arc Limited • Polska

Hybrid
PLN 240,000 - 360,000
Datacenter Infrastructure Specialist
Datacenter Infrastructure Specialist

The POD Network • Polska

On-site
PLN 461,000 - 614,000
Equity
Medical, dental & vision plans
Flexible PTO
+2
Technical Lead - GPU Infrastructure
Technical Lead - GPU Infrastructure

Tether • Warszawa

On-site
PLN 300,000 - 520,000
Remote-friendly
Global collaboration
Senior Forward Deployed Solution Engineer (Poland) Engineering · Wroclaw · Onsite
Senior Forward Deployed Solution Engineer (Poland) Engineering · Wroclaw · Onsite

Spectro Cloud • Wrocław

On-site
PLN 180,000 - 300,000
Hardware Engineer, GPU Infrastructure
Hardware Engineer, GPU Infrastructure

CoreWeave Europe • Warszawa

On-site
PLN 262,000 - 350,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)
Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis • Poznań

On-site
PLN 180,000 - 300,000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA Gruppe • Warszawa

On-site
PLN 221,250 - 507,000
Senior Platform Engineer
Senior Platform Engineer

STN Inc • Poland

On-site
PLN 530,705 - 720,242
RDMA Performance engineer
RDMA Performance engineer

Codilime • Warszawa

Remote
PLN 120,000 - 170,000
Flexible work arrangements
Training budget
Senior Software Engineer, Cloud Automation
Senior Software Engineer, Cloud Automation

NVIDIA Gruppe • Warszawa

On-site
PLN 183,750 - 416,000