We are looking for a Senior Site Reliability Engineer who is interested in an opportunity to work for an innovative hospitality company with cutting-edge technologies, with new development activities and challenges ahead. Our international team members share a common desire to develop brilliant products on reliable and resilient systems, along with their own skills. We run our services in Azure and traditional data centers. Take a chance to make a valuable contribution and enhance your professional skills.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that cloud services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer needs, and a fast rate of improvement. Additionally, SREs will keep an ever-watchful eye on our systems' capacity and performance.
Responsibilities:
- Develop and improve the whole lifecycle of services
- Establish and improve monitoring capabilities to reduce outage frequency and duration
- Create sustainable systems through automation and uplifts
- Develop and scale systems sustainably through mechanisms such as automation, and evolve systems by pushing for changes that improve reliability and velocity.
- Lead designs of major software components, systems, and features to improve the availability, scalability, latency, and efficiency of our services
- Analyze and support services before they go live via system design consulting, developing software platforms and frameworks, capacity planning
- Conduct post-incident analysis and reviews with an attitude of continuous improvement
Requirements:
- Ideally, strong experience in Azure Services and capabilities, but other cloud services (AWS, Google Cloud Platform etc.) will be considered
- Confidence and strong experience with KubernetesRecent and fluent Terraform and (Chef platform experience nice to have)
- Extensive expertise in software development/testing, development operations, and site reliability engineering
- Experience of Unix/Linux administration - an appreciation of systems internals (e.g., filesystems, system calls) is a bonus
- Experience with Continuous Integration and Deployment (CI/CD) and release orchestration and Configuration Management of VMs
- Cloud-agnostic approach, with flexibility to work across various cloud platforms
- Experience programming in one or more of the following languages: C#,, C++, Java, Python, JavaScript, Go, Perl, or Ruby
Nice to have:
- Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience
- Experience in distributed systems, storage systems, or databases
- Experience designing, analyzing, and troubleshooting large-scale distributed systems
- Systematic problem-solving approach, combined with excellent communication skills and a sense of ownership and drive
- Experience in configuring application monitoring with Azure Monitor and Application Insight
- Experience with Service Mesh
- Previous experience as a DevOps engineer is preferred
We offer*:
- Flexible working format - remote, office-based or flexible
- A competitive salary and good compensation package
- Personalized career growth
- Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
- Active tech communities with regular knowledge sharing
- Education reimbursement
- Memorable anniversary presents
- Corporate events and team buildings
- Other location-specific benefits
- not applicable for freelancers