Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Cloudjobs in New York seeks a reliability engineer to ensure uptime and reliability of a large-scale compute fleet. You will build automation for provisioning, health monitoring, and incident response to minimize outages.
The role requires deep system knowledge, strong Linux and networking fundamentals, experience with server hardware, and proficiency in Python or Go. Collaboration with cross-functional teams is essential to maintain highly available infrastructure.
The role focuses on ensuring the reliability and uptime of OpenAI's compute fleet by minimizing hardware failures. Responsibilities include building automation for server provisioning, monitoring health, and fixing performance bottlenecks.
Requirements: Candidates should have experience managing large-scale server environments and proficiency in Python or Go. Strong knowledge of Linux, networking, and server hardware is required, with a capacity for deep system-level investigation.
Key Skills: Python, Go, Linux, SQL, PromQL, Pandas, HPC, Distributed Systems, Server Hardware, Networking, Automation, Prometheus, Grafana, IPMI, Redfish, PCIe