A leading cloud technology firm located in Bellevue, Washington is seeking an ambitious individual for the Fleet Reliability Operations team. The role involves managing and troubleshooting supercomputing clusters, ensuring optimal performance, and supporting overall operations. The ideal candidate will have a solid foundation in Linux systems, troubleshooting skills, and some programming experience. The position offers competitive compensation ranging from $83,000 to $110,000, with a comprehensive benefits package including medical insurance and flexible PTO.
Qualifications
Strong understanding of Linux system administration and internals.
Ability to troubleshoot hardware and software issues.
Experience with software development or scripting languages like bash, python, powershell.
Responsibilities
Configure and maintain large-scale supercomputing clusters.
Troubleshoot hardware and software issues.
Monitor and analyze system performance.
Skills
Linux system administration
Troubleshooting hardware and software
Software development or scripting
Education
Bachelor's degree in a related field or equivalent experience
Tools
Grafana
Prometheus
Kubernetes
Job description
A leading cloud technology firm located in Bellevue, Washington is seeking an ambitious individual for the Fleet Reliability Operations team. The role involves managing and troubleshooting supercomputing clusters, ensuring optimal performance, and supporting overall operations. The ideal candidate will have a solid foundation in Linux systems, troubleshooting skills, and some programming experience. The position offers competitive compensation ranging from $83,000 to $110,000, with a comprehensive benefits package including medical insurance and flexible PTO.