A consulting services company located in Wisconsin is seeking a Site Reliability Engineer with proven experience in managing Kubernetes and cloud platforms. Ideal candidates will have strong scripting skills and expertise in designing monitoring systems. This position emphasizes incident response and postmortem analysis expertise. Competitive salary and benefits offered.
Qualifications
Proven experience in Site Reliability Engineering, DevOps, or a similar role.
Strong understanding of distributed systems and microservices architecture.
Skills
Proficiency in programming languages
Kubernetes management
Cloud platforms experience
Monitoring systems design
Incident response
Tools
Python
Kubernetes
Azure
Docker
Prometheus
Grafana
Job description
Job Description:
Proven experience in Site Reliability Engineering, DevOps, or a similar role.
Skills:
Proficiency in programming and scripting languages (e.g., Python, Go, Bash).
Proven experience managing Kubernetes in production environments.
Experience with cloud platforms (Azure) and container orchestration systems (Kubernetes, Docker)
Experience in Prometheus, Grafana, Datadog, AppDynamics, New Relic, Elasticsearch etc.
Strong understanding of distributed systems, microservices architecture, and networking.
Expertise in designing monitoring systems with KPIs, SLOs, and SLIs.
Experience with incident response, postmortem analysis, and reliability reporting.