Nebius is a Nasdaq-listed AI cloud infrastructure company building full-stack platforms for GPU orchestration and AI deployment. This role is a Technical Writer for the Infrastructure L3 Support team, responsible for building and maintaining the knowledge system for server hardware, GPU platforms, firmware, and Linux diagnostics across the global data center fleet.
What You’ll Do
- Build and maintain the L3 knowledge base for GPU server platforms, server hardware, firmware, out-of-band management, and Linux-level diagnostics, including runbooks, SOPs, troubleshooting guides, and error-code documentation
- Work with L3 and R&D engineers during investigations to capture symptoms, root causes, and resolutions, translating complex technical findings into clear, repeatable operational procedures for L1 and L2 technicians
- Write layered documentation for audiences at different technical levels, explaining complex concepts with appropriate terminology, diagrams, examples, and clear prerequisites, warnings, decision points, and success criteria
- Validate procedures end-to-end in appropriate environments and test with representative L1 and L2 users, using their feedback and escalation patterns to continuously improve the knowledge base
- Define and maintain documentation governance including templates, quality standards, metadata, ownership rules, approval workflows, review cycles, and document status tracking
- Support new hardware platform readiness by creating complete documentation packages and traveling to data centers in EMEA and other locations to observe procedures and capture operational knowledge
What You Need
- Hands-on experience in data center, server infrastructure, production operations, or site reliability engineering
- Working knowledge of Linux, server hardware, firmware, and out-of-band management technologies such as IPMI, BMC, OpenBMC, or Redfish
- Demonstrated record of creating runbooks, SOPs, or troubleshooting guides successfully used by operations teams
- Experience translating complex engineering investigations into safe, repeatable operational procedures with strong organizational and writing skills
- Fluent written and spoken English and willingness to travel regularly to data center locations
Nice to Have
- Experience with NVIDIA GPU server platforms and tools such as nvidia-smi, DCGM, dcgmi, and log-correlation tooling
- Experience with HGX or other large-scale AI infrastructure platforms
- Exposure to OCP-based platforms or ODM manufacturing ecosystems
- Experience using Bash or Python for log collection, diagnostics, or operational automation
- Experience with documentation-as-code, Git-based workflows, wikis, or large-scale knowledge-base platforms
- Experience defining documentation metrics or using incident and escalation data to prioritize improvements
Competitive compensation, career growth and learning opportunities, flexibility and ownership, collaborative and innovative culture