JOB SUMMARY
Come and join us at Avid as a Site Reliability Engineer (Remote, Philippines), where you will play a key role in ensuring the reliability, performance, and scalability of our cloud infrastructure and production systems. You’ll work closely with cross‑functional engineering teams to design resilient architectures, automate deployments, and deliver a highly available platform.
WHAT YOU WILL DO
- Champion and continuously improve platform reliability, observability, and DevOps culture across the engineering organization.
- Define and track SLAs, SLOs, and SLIs to drive reliability goals and monitor service health across the platform.
- Design, implement, and tune application and component monitoring, alerting, and dashboards using Prometheus, Grafana, CloudWatch, and Elastic.
- Improve and harden core systems in conjunction with the larger cloud engineering team, including:
- Design, operate, and optimize Kubernetes workloads on Amazon EKS, managing containerized applications across multiple environments.
- Implement and maintain Istio service mesh for secure, resilient, and observable service‑to‑service communication within Kubernetes.
- Build and manage GitOps pipelines using ArgoCD, ensuring Kubernetes manifests and Helm charts are deployed and audited correctly.
- Automate CI/CD workflows with GitHub Actions, enabling fast and safe software delivery.
- Automate infrastructure provisioning with Terraform, enabling consistent, repeatable, and auditable AWS deployments.
- Participate in a 24/7 on‑call rotation, handle incident response, perform post‑mortems, and maintain up‑to‑date runbooks.
- Secure applications and infrastructure with tools like Snyk and follow security best practices.
- Manage edge and DNS configurations with Cloudflare and Route 53, ensuring performance and global availability.
- Operate and tune AWS services such as RDS, OpenSearch, and IAM, supporting data and identity needs.
WHAT YOU CAN DELIVER
Minimum Requirements
- Bachelor’s degree in Information Technology, Computer Science, Software Engineering, or related fields.
- 5+ years of experience in Site Reliability Engineering, DevOps, and/or equivalent.
- Strong proficiency with Kubernetes (preferably Amazon EKS) and containerized application deployments.
- Proficiency with observability stacks (Prometheus, Grafana, ELK) and alerting best practices.
- Strong scripting skills (Bash, Python, or similar) for automation and tooling.
Preferred Skills, Experience, Capabilities
- Infrastructure‑as‑code tools such as Terraform, and GitOps workflows using ArgoCD.
- AWS services: RDS, IAM, OpenSearch, CloudWatch, Route 53.
- CI/CD automation using GitHub Actions or similar.
- Service mesh technologies such as Istio.
- Security: vulnerability scanning, remediation.
- Problem‑solving in distributed, cloud‑native systems.
- Excellent communication skills and cross‑functional collaboration.
WHAT TO LOOK FORWARD TO
Join a global team and experience a dynamic, collaborative work environment that fosters innovation and growth. Remote work model offering flexibility to balance work and life. Access to development programs with strong support and mentoring to help you grow and advance within the company. Attractive benefits package including health & life insurance, referral rewards, and generous leave policies to ensure a healthy work‑life balance.
Avid is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.