An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Legend Holding Group Ltd is seeking a Senior DevOps SRE Engineer with 5-6 years of hands-on experience to own reliability and cost efficiency across the tech platform. The role emphasizes AWS, Kubernetes (EKS), CI/CD, IaC, monitoring, security, and cost optimization, with a strong on-call and incident response component.
You will own production infrastructure in AWS, drive automated deployments, and lead RCA/postmortems, while building scalable observability and FinOps capabilities across
JD - Senior DevOps SRE EngineerExperience 5 6 Years Department Technology Engineering Employment Type Full-Time Reporting To Engineering Manager IT Manager
We are looking for a Senior DevOps SRE Engineer with 5 6 years of experience to own the reliability scalability observability security automation and cost efficiency of our technology platform The role is highly hands-on and requires strong experience with AWS Kubernetes EKS CI CD Infrastructure as Code monitoring incident management and cloud cost optimization
Design manage and optimize production infrastructure on AWS Manage services including EKS EC2 RDS S3 ECR ElastiCache VPC IAM Load Balancers Route 53 CloudFront and CloudWatch Manage production and non-production environments Implement highly available scalable secure and cost-efficient infrastructure Troubleshoot infrastructure networking and production issues
Strong hands-on ownership of Kubernetes AWS EKS Manage deployments services ingress HPA RBAC secrets namespaces and node groups Troubleshoot pod node networking resource and scheduling issues Use Helm for application deployments Build and maintain CI CD pipelines using GitHub Actions Jenkins or similar tools Implement automated deployments rollbacks and release processes
Define and monitor SLIs SLOs and SLAs for critical services Improve system availability resilience performance and scalability Implement proactive monitoring alerting health checks and automated recovery Reduce operational toil through automation Participate in production incident management and on-call activities Lead RCA and postmortems for major incidents Drive corrective and preventive actions
Build and maintain monitoring logging and distributed tracing Work with tools such as Datadog Prometheus Grafana CloudWatch OpenTelemetry Fluent Bit or ELK OpenSearch Create actionable dashboards and alerts Monitor latency traffic errors saturation infrastructure health and service availability
Manage infrastructure using Terraform or equivalent IaC tools Automate infrastructure provisioning deployments scaling backups and operational processes Build reusable infrastructure and deployment components Minimize manual operational activities
Implement AWS and Kubernetes security best practices Manage IAM secrets access controls and infrastructure security Implement backup disaster recovery RTO RPO and high-availability strategies Regularly test recovery and failover procedures
Own BU-level cloud cost visibility allocation and optimization Implement resource tagging by BU application environment and project Monitor AWS spend and budget vs actuals Identify cost anomalies and unnecessary underutilized resources Optimize EKS EC2 RDS Redis S3 logging and data-transfer costs Build cost dashboards and reports for Business and Finance teams Drive measurable cloud cost savings without compromising reliability or performance
5 6 years of hands-on DevOps SRE Cloud Engineering experience Strong AWS experience Strong Kubernetes EKS experience Strong Docker and Linux experience Strong CI CD experience Experience with Terraform and Helm Strong scripting skills in Bash Python or Go Strong monitoring and observability experience Experience handling production incidents and conducting RCA Good understanding of networking security and cloud architecture Understanding of SRE principles SLI SLO SLA Error Budgets MTTR capacity planning and incident management Experience with cloud cost management FinOps is highly preferred Good to HaveDatadogPrometheus GrafanaOpenTelemetryKafkaRedis ElastiCachePostgreSQL RDSArgoCD GitOpsOpenCostAWS certifications
Production availability and SLO achievementReduction in MTTR and recurring incidentsCI CD reliability and deployment successInfrastructure automationObservability and alert coverageKubernetes health and resource utilizationDisaster recovery readinessBU-level cloud cost visibility and optimizationMeasurable cloud cost savings
A hands-on engineer who thinks beyond deployment and focuses on: Reliability + Automation + Observability + Security + Scalability + Cost Efficiency Someone who can own the platform from AWS Kubernetes CI/CD Monitoring Incident Response Automation Cost Optimization.