A leading data management company in the United States is seeking a Site Reliability Engineer (SRE) to enhance the reliability and scalability of their cloud platform. The role requires deep expertise in system reliability, Kubernetes, and monitoring tools. Candidates should possess strong programming skills, proficiency in cloud platforms, and excellent communication skills. This position offers the opportunity to work collaboratively with cross-functional teams and play a key role in operational excellence.
Qualifications
Minimum five years of experience supporting complex distributed systems.
Experience with monitoring and debugging tools like Prometheus or Grafana.
Proficiency in at least one major cloud platform (AWS, GCP, Azure).
Strong experience with Linux systems including performance tuning.
Excellent written and verbal communication skills.
Responsibilities
Deploy and maintain a reliable fleet of Kubernetes clusters.
Design and implement systems to enhance service availability.
Develop monitoring and incident response strategies.
Automate repetitive tasks for operational efficiency.
Work closely with teams to integrate reliability practices.
Skills
SRE Expertise
Observability Tools
Cloud Platforms
Database Knowledge
Programming Skills
Linux Expertise
Communication Skills
Job description
A leading data management company in the United States is seeking a Site Reliability Engineer (SRE) to enhance the reliability and scalability of their cloud platform. The role requires deep expertise in system reliability, Kubernetes, and monitoring tools. Candidates should possess strong programming skills, proficiency in cloud platforms, and excellent communication skills. This position offers the opportunity to work collaboratively with cross-functional teams and play a key role in operational excellence.