Job Description:
Act as software detectives, provide a dynamic service identifying and solving issues within multiple components of critical business systems.
Deep Technical Support & Root Cause Analysis
- Handle complex customer issues, going deep into root cause analysis that spans customer configuration, product bugs, and underlying infrastructure.
Proactive Reliability Improvement
- Analyse patterns of customer issues, support cases, and outage impacts to identify systemic weaknesses.
- Design and implement solutions, automation, or monitoring to prevent future occurrences (involving coding, configuration changes, or proposing architectural improvements).
Bridging Customer Impact and Engineering
- Act as a liaison between the customer-facing support teams and the core Support/Development teams.
- Translate customer pain into technical requirements and SLOs, and explain technical constraints and incident impacts back to support teams.
Improving Supportability
- Develop tools, playbooks, and dashboards to help front-line support and diagnose and resolve issues more quickly.
- Feedback into the product development lifecycle to ensure new features are designed with supportability and reliability in mind.
Customer-Centric Service Level Objectives (SLOs)
- Contribute to defining and refining SLOs to better reflect actual customer pain and perceived performance, not just server-side metrics.
Incident Management & Postmortems
- Participate in incident response, bringing a strong understanding of customer impact.
- Contribute significantly to postmortems, ensuring preventative actions address both the technical root cause and the customer's experience.
Proactive Customer Engagement (for Key Customers)
- Engage with key customers to understand their critical workloads, review their architecture, and provide guidance on reliability best practices.
Essential Experience & Qualifications
- 4–8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Systems Engineering.
- Experience supporting enterprise platforms.
- Experience with cloud-native technologies and automation.
- Experience with AI / ML
Desirable
- Certified Kubernetes Administrator (CKA) / Certified Kubernetes Application Developer (CKAD)
- Google Professional Cloud DevOps Engineer
- Google Professional Cloud Network Engineer / VMware Certified Professional - Network Virtualization (2V0-41.24)
- Linux Foundation Certified System Administrator (LFCS) / Linux Foundation Certified IT Associate (LFCA) / Linux - ------- Professional Institute LPIC-3 Mixed Environments
- Prometheus Certified Associate (PCA)
Required Technical Skills
- Infrastructure & Platforms
- Kubernetes (essential)
- Linux (high priority)
- Networking (high priority)
- Service Mesh Basics (Istio) (high priority)
- PKI (high priority)
- AI / ML (high priority)
- Google Compute
- Storage Technologies (optional)
- Observability & Monitoring
- Prometheus (high priority)
- Grafana (high priority)
- Loki (high priority)
- Splunk
- -PH
Requirements: