About this role
We are seeking a foundational member for the Cloud infrastructure team at Writer. This role involves contributing to the development and implementation of our Site Reliability Engineering (SRE) program. The ideal candidate will ensure the reliability, scalability, performance, and security of Writer’s critical systems, proactively ensuring our high-ROI products reach customers seamlessly.
Your responsibilities:
- Lead the design, implementation, and maintenance of Writer, Inc.’s cloud infrastructure to ensure high availability and performance.
- Design and implement scalable cloud automation to support seamless deployment for enterprise customers.
- Automate infrastructure provisioning and management using Terraform & Python.
- Collaborate with development teams to optimize cloud resources and enhance system reliability.
- Develop and maintain monitoring and alerting systems to proactively identify and resolve issues.
- Conduct post-mortem analyses of system failures to identify root causes and implement preventive measures.
- Optimize and scale cloud infrastructure to support growth and ensure cost efficiency.
- Ensure security and compliance of systems, adhering to industry standards and regulations.
- Mentor and guide junior engineers, fostering a culture of reliability and continuous improvement.
- Stay current with emerging technologies and industry trends to improve SRE practices.
Is this you?
- Proven expertise in Site Reliability Engineering with at least 7 years of experience.
- Deep understanding of system architecture and infrastructure design for high availability and performance.
- Bachelor’s degree in Computer Science, Engineering, or related field.
- Strong proficiency in programming languages such as Python, Java, or Go for automation and monitoring.
- Experience with cloud platforms like AWS, Azure, or GCP and their services for scalable systems.
- Expertise in containerization (Docker, Kubernetes) and orchestration tools.
- Knowledge of monitoring and logging tools (Prometheus, Grafana, ELK Stack).
- Ability to lead and mentor junior engineers in reliability best practices.
- Excellent communication skills for effective collaboration.
- Proactive in identifying and mitigating system failures and bottlenecks.
Preferred skills & experience:
- Software engineering expertise
- Terraform
- Python
- Kubernetes
- Scala
- AWS/GCP
Benefits & perks (UK full-time employees):
- Generous PTO and company holidays
- Comprehensive medical and dental insurance
- Paid parental leave (12 weeks)
- Fertility and family planning support
- Early-detection cancer testing through Galleri
- Competitive pension scheme and company contributions
- Annual stipends for home office setup, wellness, and learning
- Company and team off-sites
- Competitive salary and stock options