Team Lead - Application DevSecOps & SRE

Dialog Group Berhad

Petaling Jaya

On-site

MYR 240,000 - 360,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dialog Group Berhad is seeking a Lead Engineer to head a small DevSecOps and SRE team responsible for the speed, reliability and security of major external-facing applications built on AWS with Kubernetes. You will own environments, write IaC, and drive end-to-end CI/CD and zero-downtime deployments across microservices.

The role requires 6+ years in cloud infrastructure or SRE with senior/Mentor experience, deep AWS expertise, and the ability to set technical direction while remaining calm

Qualifications

  • 6+ years in cloud infrastructure, DevOps, DevSecOps or SRE with production ownership
  • Deep hands-on AWS experience across compute, databases, storage and IAM
  • Strong IaC skills with Terraform or CloudFormation and CI/CD ownership

Responsibilities

  • Own end-to-end DevSecOps and SRE for core applications
  • Lead a small team, set technical direction and mentor engineers
  • Plan capacity, performance testing and cost governance with central infra
  • Collaborate with infrastructure and application teams to align on strategy

Skills

AWS
Kubernetes
Terraform
CloudFormation
CI/CD
Observability
SRE
Incident handling
Security
Cost optimization
Linux
Communication
Leadership

Tools

Terraform
CloudFormation
Jenkins
GitHub Actions
GitLab CI

Job description

  • We are hiring a Lead Engineer to head a small, high-ownership DevSecOps and SRE team of two to three engineers.
  • The team owns the speed, reliability and security of our major external-facing applications — distributed microservices platforms running on managed Kubernetes in AWS. The immediate focus is one significant system, and as the team matures and capacity allows, that scope is expected to extend to further business-critical applications.
  • You are not a shared services function stretched thin across a sprawling portfolio. You are the product-focused infrastructure owner for a small number of systems that genuinely matter, operating on Site Reliability Engineering principles.
  • This is a genuine lead position. We are looking for the seniority and judgement to set technical direction, to keep a small team focused and effective, and to stay composed when a production system is misbehaving and several people want an answer at once.
  • You will work in close partnership with our central infrastructure team and our application engineering team. Neither is a supervisor handing down instructions — technical direction is reached collaboratively, and your reasoning carries real weight in it.
Scope of Ownership
Application Environments

Full ownership of the Dev, UAT, Staging and Production environments for the applications in your team’s scope. Writing and maintaining the application-specific Infrastructure-as-Code covering their resources — Kubernetes clusters, databases, queues and caches.

Cluster lifecycle management: creation, version upgrades, security patching, and scaling of both control plane and worker nodes. Cluster cost optimisation through node autoscaling, Reserved and Spot instance strategy, and pod scheduling efficiency. Internal cluster networking — CNI configuration, service mesh, and ingress/egress controllers, operating within the address space allocated by the central infrastructure team.

Design, build and maintenance of end-to-end CI/CD pipelines for the applications in scope. Implementation of advanced zero-downtime deployment strategies — blue/green and canary releases — across the microservices estate. Management of application configuration and feature flags in production.

Site Reliability Engineering

Implementation and tuning of the monitoring, logging and distributed tracing stack for the systems you own. Definition, tracking and reporting of Service Level Objectives and Indicators. Owning the post-incident process — root cause analysis and corrective actions that actually get closed out.

DevSecOps Integration

Integration of security scanning (SAST/DAST) directly into the delivery pipeline. Vulnerability management across application dependencies and runtime environments, with a rapid patching cadence. Handling of application secrets, keys and certificates.

Capacity & Performance

Planning and execution of regular performance and load testing in close partnership with the QA team. Configuration and optimisation of application-level auto-scaling across Kubernetes and compute resources.

  • Experience: 6+ years in cloud infrastructure, DevOps, DevSecOps or SRE, including meaningful production ownership. Prior technical lead experience, or senior experience with demonstrated mentoring and direction-setting.
  • Cloud: Deep hands-on AWS. Comfortable across compute, managed relational databases, managed caching, object storage, IAM and networking primitives.
  • Kubernetes: Production ownership of a managed Kubernetes service (EKS, AKS or GKE) — cluster upgrades, node group management, ingress and load balancer routing, namespace segregation, and workload resource definitions.
  • Infrastructure as Code: Strong proficiency with a declarative IaC toolchain (Terraform or CloudFormation) — authoring reusable modules, managing remote state and locking, debugging provider behaviour, and enforcing plan review gates.
  • CI/CD: Demonstrated end-to-end ownership of pipelines on an enterprise platform (Azure DevOps, GitLab CI, GitHub Actions or Jenkins), including progressive and zero-downtime release strategies.
  • Observability & SRE practice: Building and tuning monitoring, logging and tracing stacks. Practical experience defining SLIs and SLOs and using error budgets to inform release decisions.
  • Incident handling: Able to lead calmly during a production incident and to run a blameless post-mortem that produces corrective actions people actually complete.
  • Security: Integrating SAST/DAST and container image scanning into delivery. Managed secret handling. Least-privilege identity design, including workload identity federation (IAM roles for service accounts or equivalent).
  • Cost governance: Practical FinOps — right-sizing, autoscaling policy, Spot and Reserved capacity strategy, and identifying structural waste.
  • Linux: Strong native Linux and WSL2 proficiency, with terminal-first operational habits.
  • Communication: Able to explain a technical trade-off to a non-specialist stakeholder without either oversimplifying it or hiding behind detail.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
Senior DevOps Engineer
Senior DevOps Engineer

Aventra Group • Kuala Lumpur

On-site
MYR 120,000 - 160,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000
SENIOR DEVOPS TECHNICAL LEAD - DevOps CI/CD pipeline
SENIOR DEVOPS TECHNICAL LEAD - DevOps CI/CD pipeline

Happiest Minds Technologies • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Lead DevOps Engineer
Lead DevOps Engineer

RinggitPlus • Malaysia

Hybrid
MYR 180,000 - 300,000
DevOps Engineer (Cloud Hosting & Operations)
DevOps Engineer (Cloud Hosting & Operations)

AGENSI PEKERJAAN LINKTRIX CONSULTANTS SDN. BHD. • Kuala Lumpur

On-site
MYR 180,000 - 360,000
Lead DevSecOps & SRE Engineer (Kubernetes/AWS)
Lead DevSecOps & SRE Engineer (Kubernetes/AWS)

Dialog Group Berhad • Petaling Jaya

On-site
MYR 240,000 - 360,000
Engineering Manager – Platform & SRE
Engineering Manager – Platform & SRE

INSCALE • Kuala Lumpur

On-site
MYR 350,000 - 650,000
Site Reliability Engineer
Site Reliability Engineer

SLB • Kuching

On-site
MYR 90,000 - 150,000