Get more replies from employers
Send a job-specific resume in minutes.
NCS Philippines is seeking a Site Reliability Engineer to define SLOs/SLIs, build and maintain monitoring, logging, and alerting systems, and embed reliability into CI/CD workflows.
You will drive automated remediation, lead post-incident RCAs, and partner with cross-functional teams to improve scalability, fault tolerance, and system resilience across cloud platforms.
Define, implement, and enforce SLOs and SLIs to balance innovation speed and system stability.
Develop and maintain monitoring, logging, and alerting systems for early detection of anomalies using industry-standard tools.
Establish metrics-driven approaches to performance monitoring, ensuring continuous tracking of availability, latency, error rates, and resource utilization.
Integrate SLOs and monitoring within CI/CD workflows to ensure new releases meet reliability standards.
Implement automated remediation for recurring incidents, optimizing incident management to reduce Mean Time to Resolution (MTTR).
Conduct post-incident RCAs, recommending and implementing preventive measures.
Continuously refine alerting policies to ensure they are actionable, relevant, and reduce alert fatigue.
Participate in on-call rotations and elevate issues following established protocols.
Collaborate with cross-functional teams to improve scalability, fault tolerance, and system resilience.
Bachelor’s degree in Computer Science, Engineering, or related technical field.
Master’s degree or relevant certifications (e.g., Google Cloud Professional DevOps Engineer, AWS Certified DevOps Engineer) preferred.
4+ years in a Site Reliability, DevOps, or similar role managing production systems at scale.
Proven experience defining SLOs, SLIs, and implementing monitoring/alerting systems in cloud environments.
Experience with incident management, RCA, and improving system resilience.
Distributed systems, microservices architecture, and cloud platforms (AWS, GCP, Azure).
System performance metrics and availability management principles.
Incident response processes and tools (PagerDuty, Opsgenie).
Monitoring and observability tools (Prometheus, Grafana, Datadog, New Relic).
CI/CD workflows with integrated monitoring.
Alerting strategies, escalation processes, and on-call best practices.
Scripting and automation tools (Python, Bash, Terraform, Ansible).
Performance optimization, fault tolerance, and scalability techniques.