Lead Site Reliability Engineer ELK

Horizontal Talent

Kuala Lumpur

On-site

MYR 240,000 - 420,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Horizontal Talent is seeking a Lead Site Reliability Engineer to own the ELK and Kafka platforms in Kuala Lumpur, ensuring reliability at scale. You will influence architecture, drive operational excellence, and mentor engineers while hands-on troubleshooting critical systems.

The role emphasizes deep ELK expertise in production environments, strong Linux skills, and IaC-driven automation. Collaboration with Architecture, Security, and Operations is essential.

Qualifications

  • Bachelor's degree in CS/Engineering/IT or related field.
  • 15+ years in infrastructure/platform/site reliability engineering or similar.
  • Proven ownership of large-scale ELK environments (Elasticsearch, Logstash, Kibana) and Kafka.
  • Strong Linux administration and automation with IaC experience.
  • Experience leading data platforms in enterprise, regulated, or banking contexts is preferred.

Responsibilities

  • Lead design, operation and evolution of enterprise ELK/Kibana and Kafka platforms.
  • Drive architecture, scalability, capacity planning, and reliability improvements.
  • Manage cluster health, tuning, ILM, DR planning, and upgrades.
  • Mentor teams, drive best practices, and communicate with stakeholders.
  • Own incident response, post-incident reviews, and automation tooling.

Skills

Elasticsearch
Logstash
Kibana
Kafka
Linux
Scripting
Infrastructure as Code
Architectural leadership
Observability
Distributed systems

Education

Bachelor's Degree in Computer Science/Engineering/IT

Tools

Databricks
AWS/Azure/GCP

Job description

Lead Site Reliability Engineer ELK
Kuala Lumpur, Malaysia

About Horizontal

Established since 2003 in the US, Horizontal solves complex challenges across two distinct businesses: Horizontal Digital and Horizontal Talent. We are consistently recognized for being a top workplace and one of the fastest-growing private companies. Horizontal Talent specializes in staffing for IT, Digital & Creative, and Business & Strategy markets. We have global offices in US, UAE, India, Australia and Malaysia.

Job Summary

We are seeking a highly experienced Lead Site Reliability Engineer to own and drive the reliability, scalability, and evolution of our enterprise observability platform. This is not a traditional DevOps, Cloud Engineering, or Kubernetes-focused role. We are specifically looking for an engineer with deep hands-on expertise operating large-scale Elasticsearch, Logstash, Kibana (ELK) and Kafka platforms in mission-critical production environments. The successful candidate will serve as a technical leader, influencing architecture decisions, driving operational excellence, and mentoring engineering teams while remaining actively involved in troubleshooting and platform engineering. You will work closely with Architecture, Engineering, Security, and Operations teams to ensure our observability and data ingestion platforms remain reliable, scalable, and performant.

Key Responsibilities
  • Lead the design, operation, maintenance, and continuous evolution of enterprise-scale Elasticsearch, Logstash, and Kibana platforms.
  • Drive platform architecture, scalability planning, capacity management, and reliability improvements.
  • Manage Elasticsearch cluster health, performance tuning, shard allocation strategies, index lifecycle management, and storage optimization.
  • Plan and execute platform upgrades, migrations, patching activities, and disaster recovery initiatives.
  • Investigate and resolve critical production incidents related to data ingestion, indexing performance, search latency, and cluster stability.
  • Establish operational standards, monitoring frameworks, and best practices for observability platform operations.
Kafka & Data Ingestion Engineering
  • Design, operate, and optimize Kafka-based data ingestion pipelines supporting large-scale enterprise workloads.
  • Troubleshoot data flow issues across Kafka, Logstash, Elasticsearch, and downstream consumers.
  • Manage ingestion performance, throughput optimization, data retention, replay strategies, and platform resilience.
  • Identify and eliminate bottlenecks across distributed ingestion architectures.
  • Ensure end-to-end data reliability, consistency, and availability across production environments.
Technical Leadership
  • Act as a technical authority for observability and platform engineering initiatives.
  • Lead architecture reviews and challenge designs when necessary to ensure operational excellence.
  • Drive technical decision-making through influence, collaboration, and engineering best practices.
  • Partner with cross-functional teams to define scalable and maintainable platform solutions.
  • Mentor engineers and contribute to the growth of technical capabilities across the organization.
  • Simplify and communicate complex technical concepts to both technical and non-technical stakeholders.
Reliability & Operations
  • Lead root cause analysis for complex production incidents and implement sustainable preventive measures.
  • Develop automation, monitoring, self-healing capabilities, and operational tooling.
  • Drive continuous improvement initiatives across reliability, performance, security, and observability.
  • Participate in critical incident response and post-incident reviews.
  • Establish robust operational processes, documentation, and knowledge-sharing practices.
Required Skills & Qualifications
  • Bachelor's Degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 15+ years of experience in infrastructure engineering, platform engineering, site reliability engineering, or related enterprise technology environments.
  • Demonstrated experience owning and operating large-scale production ELK platforms.
  • Deep expertise in:
    • Elasticsearch
    • Logstash
    • Kibana
    • Kafka
  • Strong understanding of:
    • Distributed systems architecture
    • High-availability platform design
    • Capacity planning
    • Performance tuning
    • Disaster recovery
    • Enterprise observability platforms
  • Proven experience troubleshooting and optimizing large-scale Kafka ingestion pipelines
  • Strong Linux administration and systems engineering knowledge.
  • Experience automating operational processes using scripting and Infrastructure-as-Code practices.
  • Ability to lead technical discussions across Architecture, Engineering, Operations, and Security teams.
  • Excellent analytical, troubleshooting, and communication skills.
Experience
  • 15+ years of overall experience in data engineering, data architecture, or analytics platform roles.
  • 3+ years of experience in a senior or lead capacity, with accountability for architecture decisions and technical delivery.
  • Demonstrated experience leading data engineering teams or acting as the primary technical authority for data platforms.
  • Prior experience supporting banking, financial services, or regulated environments is strongly preferred.
  • Proven track record of working with cloud-based data platforms (Databricks, AWS / Azure / GCP), including production-grade systems.
  • Bachelor’s or Master’s degree in Computer Science, IT, Engineering, or a related technical field.
  • Relevant certifications in cloud platforms, data engineering, or data architecture is an advantage

The above description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ELK & Kafka Platform Lead
Senior ELK & Kafka Platform Lead

Horizontal Talent • Kuala Lumpur

On-site
MYR 240,000 - 420,000
Senior Data Engineer - Data Lakehouse
Senior Data Engineer - Data Lakehouse

ONL Biz Solutions Sdn Bhd • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Senior Platform Engineer
Senior Platform Engineer

Monroe Consulting Group • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Splunk Enterprise Security Engineer
Splunk Enterprise Security Engineer

NTT DATA Business Solutions • Cyberjaya

On-site
MYR 150,000 - 210,000
Senior Platform Engineer
Senior Platform Engineer

Businesslist • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Cloud Engineer
Cloud Engineer

DIGITAL BRIGHT TECHNOLOGIES • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Splunk Enterprise Security Engineer
Splunk Enterprise Security Engineer

NTT DATA Business Solutions • Malaysia

Hybrid
MYR 120,000 - 180,000
Health Insurance
Dental
Medical checkup
+1
Ops Engineer
Ops Engineer

Millennium Technology Services • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 70,000 - 90,000