Site Reliability Engineer (SRE) Wiz / CNAPP
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in Wiz/CNAPP, Kubernetes, cloud platforms, and monitoring & observability. The role will be responsible for ensuring the reliability, scalability, security, and performance of cloud-native environments while leveraging Wiz capabilities to identify and manage cloud security risks.
The ideal candidate will have hands-on experience with AWS, Azure, and/or GCP, Kubernetes-based platforms, infrastructure automation, incident management, and modern observability tools. Experience with Wiz CNAPP, cloud security posture management, vulnerability management, and cloud workload security is highly desirable.
Key Responsibilities
- Design, implement, and maintain highly available, scalable, and resilient cloud infrastructure.
- Support and manage Wiz/CNAPP capabilities across cloud environments.
- Configure and monitor Wiz for cloud security posture, vulnerability, workload, identity, and risk management.
- Integrate Wiz findings with existing SRE, DevOps, SIEM, ticketing, and incident-management workflows.
- Manage and troubleshoot Kubernetes clusters, workloads, deployments, networking, and associated infrastructure.
- Work with AWS, Azure, and/or GCP services to maintain reliable cloud platforms.
- Implement and maintain monitoring, logging, alerting, and observability solutions.
- Define and improve SLIs, SLOs, SLAs, and error budgets for critical services.
- Perform root-cause analysis for production incidents and drive permanent remediation.
- Participate in on-call support, incident response, and problem management.
- Automate repetitive operational tasks using scripting and Infrastructure as Code.
- Collaborate with DevOps, Cloud, Security, Development, and Infrastructure teams to improve platform reliability and security.
- Identify performance bottlenecks, reliability risks, and security gaps and recommend improvements.
- Develop operational runbooks, technical documentation, and incident post-mortems.
- Continuously improve cloud infrastructure through automation, standardization, and engineering best practices.
Required Technical Skills
Wiz / CNAPP
- Hands-on experience with Wiz or similar CNAPP/cloud security platforms.
- Understanding of CNAPP, CSPM, CWPP, CIEM, vulnerability management, and cloud workload security.
- Ability to analyse cloud security findings and drive remediation.
- Experience integrating Wiz/security findings with operational and security workflows.
Site Reliability Engineering
- Strong understanding of SRE principles and practices.
- Experience with incident management, root-cause analysis, capacity planning, performance engineering, and reliability improvement.
- Good understanding of SLI, SLO, SLA, and error-budget concepts.
- Experience supporting production environments and participating in on-call rotations.
Kubernetes
- Strong hands-on experience with Kubernetes.
- Knowledge of pods, deployments, services, ingress, ConfigMaps, Secrets, RBAC, namespaces, networking, and persistent storage.
- Experience troubleshooting Kubernetes performance, availability, and deployment issues.
- Exposure to Helm and Kubernetes ecosystem tools is preferred.
Cloud Platforms
Hands-on experience with one or more:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
Understanding of cloud networking, IAM, compute, storage, containers, security controls, and highly available architectures.
Monitoring & Observability
- Experience with modern monitoring and observability platforms.
- Strong understanding of metrics, logs, traces, dashboards, and alerting.
- Experience with tools such as Prometheus, Grafana, Datadog, Splunk, ELK/Elastic, Open Telemetry, or equivalent.
- Ability to design meaningful alerts and reduce alert noise.
Automation / DevOps
- Proficiency in at least one scripting/programming language such as Python, Bash, or Go.
- Experience with Terraform or other Infrastructure-as-Code tools.
- Understanding of CI/CD pipelines and DevOps practices.
- Experience with Git-based development and configuration management.