A leading global IT services company is seeking a Reliability Engineer to design and manage monitoring and alerting systems. Ideal candidates will have over 3 years of experience with cloud platforms like AWS, Azure, or GCP. Responsibilities include proactively identifying issues and executing timely incident responses. Strong knowledge of monitoring tools and automation skills in Python or Bash are required. Join a team committed to innovation and employee growth in a diverse and collaborative environment.
Qualifications
3+ years' experience with cloud platforms in a production environment.
Solid understanding of monitoring and logging tools.
Strong scripting and automation skills.
Responsibilities
Design, implement, and manage monitoring and alerting systems.
Proactively identify issues and execute incident response.
Skills
Cloud platforms (AWS, Azure, GCP)
Monitoring tools (Datadog, CloudWatch)
Containerization (Docker, Kubernetes)
Scripting (Python, Bash)
Analytical skills
Troubleshooting
Job description
A leading global IT services company is seeking a Reliability Engineer to design and manage monitoring and alerting systems. Ideal candidates will have over 3 years of experience with cloud platforms like AWS, Azure, or GCP. Responsibilities include proactively identifying issues and executing timely incident responses. Strong knowledge of monitoring tools and automation skills in Python or Bash are required. Join a team committed to innovation and employee growth in a diverse and collaborative environment.