Technical Writer (Infrastructure L3 Support Team)

Nebius B.V.

Amsterdam

On-site

EUR 65,000 - 90,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Travel to data centers
Competitive compensation
On-site work

Job summary

Nebius B.V. is building an Infrastructure L3 Support function to document and support GPU server platforms, firmware, and Linux diagnostics across our global data center fleet.

You will translate engineering investigations into repeatable procedures, create runbooks and SOPs, and ensure knowledge is accessible to L1–L3 audiences. Travel to data centers in EMEA is expected as part of this role.

Qualifications

  • Hands‑on experience in data center, server, infra, production operations, or SRE.
  • Working knowledge of Linux, server hardware, firmware, and out‑of‑band management (IPMI, BMC, OpenBMC, Redfish).
  • Proven record creating runbooks, SOPs, or troubleshooting guides used by operations teams.
  • Experience translating complex engineering investigations into repeatable procedures.
  • Strong organizational skills and the ability to structure incomplete or complex information.
  • Fluent written and spoken English.
  • Willingness to travel regularly to data center locations.
  • Experience with documentation‑as‑code, Git‑based workflows, or large‑scale knowledge bases.

Responsibilities

  • Build and own the L3 knowledge base for GPU server platforms, server hardware, firmware, OOB management, and Linux diagnostics.
  • Create runbooks, SOPs, troubleshooting guides, error-code documentation, onboarding materials, and platform references.
  • Translate engineering investigations into safe, repeatable operational procedures.
  • Write for L1–L3 audiences with clear, structured language.
  • Test and improve documentation by validating procedures end to end.
  • Establish documentation governance: templates, ownership, approvals, versioning, review cycles.
  • Support new platform readiness with complete documentation packages.
  • Travel to data centers in EMEA and other locations as needed.

Skills

Linux knowledge
Technical writing
Runbooks & SOPs
Cross-functional collaboration
Travel readiness
Structured English
Data center operations

Tools

nvidia-smi
DCGM
dcgmi
log correlation tooling
Git-based workflows
documentation-as-code
Bash
Python

Job description

We are building our Infrastructure L3 Support function: the center of expertise for server hardware, GPU platforms, firmware—including BIOS and BMC—and deep Linux diagnostics across our global data center fleet. This function can only scale if its operational knowledge is accurate, accessible, and usable. We therefore treat documentation as a product with defined users, quality standards, ownership, feedback, and a managed lifecycle. That product will be yours to build and operate.

Your responsibilities will include:
  • Build and own the L3 knowledge system: Build and maintain the L3 knowledge base for GPU server platforms, server hardware, firmware, out-of-band management, and Linux-level diagnostics. Create runbooks, SOPs, troubleshooting guides, error-code documentation, onboarding materials, and platform reference documentation.
  • Translate engineering knowledge into operational procedures: Work with L3 and R&D engineers during and after investigations to capture symptoms, root causes, diagnostic evidence, resolutions, and preventive actions.
  • Write for L1, L2, and L3 audiences: Explain complex concepts in clear, structured language appropriate to the reader’s experience level.
  • Test and improve documentation: Validate procedures end to end in an appropriate environment before publication.
  • Establish documentation governance: Define and maintain documentation templates, quality standards, metadata, ownership rules, approval workflows, and review requirements.
  • Support new platform readiness: Create complete documentation packages for new hardware platform introductions.
  • Work where the hardware lives: Travel as needed to our data centers in EMEA, and occasionally to other locations.
We expect you to have:
  • Hands‑on experience in data center, server, infrastructure, production operations, or site reliability engineering.
  • Working knowledge of Linux, server hardware, firmware, and out-of-band management technologies such as IPMI, BMC, OpenBMC, or Redfish.
  • A demonstrated record of creating runbooks, SOPs, or troubleshooting guides that operations teams used successfully.
  • Experience translating complex engineering investigations into safe, repeatable operational procedures.
  • Strong organizational skills and the ability to structure incomplete or complex information.
  • The ability to write clear, precise, and structured English for readers with different levels of technical experience.
  • Experience validating documentation with its intended users and improving it through feedback.
  • An understanding of document governance, including ownership, approvals, versioning, review cycles, deprecation, and archiving.
  • The judgment to distinguish verified information from assumptions and to challenge incomplete or unclear technical input.
  • Strong collaboration skills and the confidence to work closely with engineers, technicians, and knowledge‑management stakeholders.
  • Fluent written and spoken English.
  • Willingness to travel regularly to data center locations.
It will be an added bonus if you have:
  • Experience with NVIDIA GPU server platforms and tools such as nvidia-smi, DCGM, dcgmi, and log‑correlation tooling.
  • Experience with HGX or other large‑scale AI infrastructure platforms.
  • Exposure to OCP‑based platforms or ODM manufacturing ecosystems.
  • Experience using Bash or Python for log collection, diagnostics, or operational automation.
  • Experience with documentation‑as‑code, Git‑based workflows, wikis, or large‑scale knowledge‑base platforms.
  • Experience defining documentation metrics or using incident and escalation data to prioritize improvements.
  • Experience supporting hardware‑platform introductions across multiple data centers.
  • Hands‑on experience in data center, server, infrastructure, production operations, or site reliability engineering, Working knowledge of Linux, server hardware, firmware, and out-of-band management (IPMI, BMC, OpenBMC, Redfish), Proven record of creating runbooks, SOPs, or troubleshooting guides, Experience translating complex engineering investigations into repeatable operational procedures, Ability to write clear, precise, and structured English for various technical levels, Experience validating documentation with users and improving it via feedback, Understanding of document governance (ownership, approvals, versioning, review cycles), Fluent written and spoken English, Willingness to travel regularly to data center locations
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Writer (Infrastructure L3 Support Team)
Technical Writer (Infrastructure L3 Support Team)

Nebius • Amsterdam

On-site
EUR 70,000 - 110,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
L3 Infrastructure Documentation Engineer
L3 Infrastructure Documentation Engineer

Nebius • Amsterdam

On-site
EUR 70,000 - 110,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
L3 Infrastructure Documentation Engineer
L3 Infrastructure Documentation Engineer

Nebius B.V. • Amsterdam

On-site
EUR 65,000 - 90,000
Travel to data centers
Competitive compensation
On-site work
Senior Technical Writer
Senior Technical Writer

Jobgether • Netherlands

On-site
EUR 90,000 - 130,000
International teams
Own documentation standards
Impactful work
+3
Data Center Lead
Data Center Lead

Webhosting • Rotterdam

Hybrid
EUR 70,000 - 110,000
IT System Engineer
IT System Engineer

Nearfield Instruments • Rotterdam

On-site
EUR 70,000 - 100,000
Product System Administrator (OT)
Product System Administrator (OT)

Nearfield Instruments • Rotterdam

Hybrid
EUR 65,000 - 90,000
System Engineer
System Engineer

CSC • Amsterdam

On-site
EUR 55,000 - 65,000
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in work
+1
Service Delivery Manager
Service Delivery Manager

Nebius B.V. • Amsterdam

On-site
EUR 90,000 - 120,000