DevOps / Site Reliability Engineer (SRE)
DevOps / Site Reliability Engineer (SRE)
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 5 to 9 years
- Relevant Experience Required: 4+ years of dedicated experience in infrastructure automation, cloud orchestration, and CI/CD pipelines
Job Summary
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, automate, and scale our cloud-native infrastructure pipelines. The ideal candidate will bridge the gap between software development and systems operations, building highly available deployment pipelines, writing infrastructure-as-code (IaC), and optimizing cluster scaling to maximize application uptime, system reliability, and performance.
- Architect and manage infrastructure-as-code (IaC) templates using Terraform or OpenTofu to provision secure, modular, and repeatable multi-environment architectures.
- Orchestrate containerized production workloads, configuring cluster scaling, service meshes, network routing policies, and deployment strategies on Kubernetes (EKS/AKS/GKE).
- Implement automated monitoring, logging, and alerting systems utilizing observability tools (e.g., Prometheus , Grafana , Datadog , ELK stack) to actively track platform performance metrics.
- Drive system high-availability and fault tolerance efforts, designing disaster recovery plans, automated load balancing parameters, and self-healing cluster scripts.
- Manage centralized configuration and secret management systems , securely vaulting database credentials, API tokens, and certificate profiles (e.g., HashiCorp Vault , AWS Secrets Manager ).
- Participate in on-call rotations and lead incident root-cause analysis (RCA), systematically diagnosing runtime infrastructure failures, performance bottlenecks, and resource leaks.
Requirements
- 5 to 9 years of core systems engineering or software development experience, with 4+ dedicated years actively designing, building, and operating cloud-native production platforms.
- Strong technical mastery of Kubernetes cluster administration, Terraform automation layouts, Linux system internals, shell scripting (Bash, Python, or Go), and network protocols.
- Deep structural understanding of microservices design, caching mechanics, database scaling limits, and cloud provider API governance.
- Prior experience implementing DevSecOps controls (e.g., integrating SAST/DAST tools directly into container build phases).
- Experience with GitOps methodologies and progressive delivery mechanisms (e.g., Canary or Blue/Green deployments using Flagger or Istio).
- 12+ years of total IT software engineering or operational management background, with 6+ dedicated years acting as a CISO, Director of Security, or Principal Enterprise GRC Advisor.
- Strong visionary mastery of modern security trends, zero-trust target states, risk calculation paradigms, and multi-cloud information landscape parameters.
- Deep communication execution skills, with a proven history of negotiating security budgets, steering board panels, and handling high-pressure public communication events.
- Mandatory certification: CISM or CISSP.
Preferred Qualifications
- Certified in the Governance of Enterprise IT (CGEIT) or Certified in Risk and Information Systems Control (CRISC) credential.
- Prior experience steering complex post-merger information platform integrations or stabilizing security posture profiles during major corporate equity restructurings.