Turn this role into an interview — a resume and cover letter built around what this employer wants.
Cloudjobs seeks a software engineer to design and build systems for managing a fleet of GPU servers and data center infrastructure. You will drive observability, automation, and tooling for fault remediation, validation, and facilities management of power and liquid cooling systems.
Candidates should have strong distributed systems experience and expertise with cloud platforms like GCP and Kubernetes, with proficiency in Go, Python, Java, or Rust, and the ability to set technical direction for
Develop software for managing a fleet of GPU servers and data center infrastructure, focusing on advanced diagnostics, observability, and automation. Create tooling for hardware fault remediation, post-repair validation, and facilities management for power and liquid cooling systems.
Requirements: Requires professional software engineering experience with expertise in distributed systems, cloud platforms like GCP and Kubernetes, and proficiency in Go, Python, Java, or Rust. Candidates should be able to independently develop scalable solutions and set technical direction for complex projects.
Key Skills: Distributed Systems, Kubernetes, Infrastructure as Code, GCP, Go, Python, Java, Rust, GPU Fleet Management, Hardware Diagnostics, Observability Tooling, Automation, Reliability Engineering, Cloud Platforms, Analytical Problem Solving, Collaboration
Benefits: Restricted Stock Units, Health insurance (HDHP and PPO), Vision insurance, Dental insurance, Employer contributions to HSA accounts, Paid Parental Leave, Paid life insurance, Short-term disability, Long-term disability, Teladoc, 401(k) with 100% match up to 4% of salary, Paid time off, Holiday schedule, Cell phone reimbursement, Tuition reimbursement, Subscription to the Calm app, MetLife Legal, Company paid commuter benefit