A technology services firm in Chicago seeks a Site Reliability Engineer to oversee monitoring of dashboards, handle production issues, and represent the SRE team in client interactions. The ideal candidate will possess experience with GCP, Kubernetes, and Dynatrace, along with a background in support projects. This role offers a dynamic work environment with a focus on continuous improvement.
Qualifications
Experience working on support projects is essential.
Familiarity with GCP, Kubernetes, and Dynatrace is required.
Knowledge of Splunk/Sumologic log monitoring is a plus.
Responsibilities
Monitor dashboards and alert for production issues.
Represent SRE in client calls and manage tickets.
Create alerts and new dashboards for features.
Skills
GCP
Kubernetes
Dynatrace
Log monitoring
Job description
Job description
Responsibilities
Monitoring all the key dashboards and timely alerting, find RCA for all production issues, help team to debug potential issues, handling daily standup calls and all other client calls, run through all the JIRA tickets and have the ticket updated with latest findings/RCA
Represent SRE in all client calls and have the deep knowledge on all the tickets/issues created by the team
Create alerts based on production issues
Create new dashboards for any new features, support onsite team with providing various details from different tools to help with the investigation, always validate if the team is following the SOPs or the process defined for an alerts/ issue
Contact external vendors if their integrations fail
Measure the front-end metrics for the site with various tools available
Qualifications
Must have worked on support projects
Must know GCP, Kubernetes and Dynatrace
Knowledge of Splunk / Sumologic log monitoring is an advantage