A leading technology company is seeking a Senior SRE to join their Compute Farm team in Durham, North Carolina. You will be responsible for owning SRE solutions end-to-end, ensuring system uptime, and delivering services across a multi-cloud environment. The ideal candidate will hold a B.S. degree in Computer Science or equivalent, have over 5 years of experience, and strong proficiency in HPC clusters and modern CI/CD techniques. The role provides competitive compensation, equity, and benefits.
Qualifications
5+ years professional experience in supporting critical services.
Experience with large-scale HPC clusters.
Proficiency in modern CI/CD techniques.
5+ years coding experience in at least two high-level programming languages.
Responsibilities
Own SRE solutions from design to continuous improvement.
Automate provisioning using Infrastructure as Code.
Ensure uptime and quality of service.
Conduct capacity management and planning.
Skills
HPC (High-Performance Computing)
Python
Infrastructure as Code (IaC)
Container management
CI/CD techniques
Coding/Scripting
Debugging skills
Strong communication
Education
B.S. degree in Computer Science or related technical field
Tools
Slurm
LSF
Kubernetes
Job description
A leading technology company is seeking a Senior SRE to join their Compute Farm team in Durham, North Carolina. You will be responsible for owning SRE solutions end-to-end, ensuring system uptime, and delivering services across a multi-cloud environment. The ideal candidate will hold a B.S. degree in Computer Science or equivalent, have over 5 years of experience, and strong proficiency in HPC clusters and modern CI/CD techniques. The role provides competitive compensation, equity, and benefits.