Senior Site Reliability Engineer

NEXT Ventures

Kuala Lumpur

On-site

MYR 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Global engineering team
Competitive compensation

Job summary

NEXT Ventures is a global fintech group powering FundedNext and FNmarkets, with offices across multiple regions. The Platform Engineering team owns the infrastructure, reliability, and observability backbone that supports product squads at scale.

As a Site Reliability Engineer, you will own centralized log management, drive service-level optimization, and perform log analysis across Linux and Windows environments to keep incidents short and signals clean.

Qualifications

  • 5–7 years of professional engineering experience with at least 3 years in SRE or reliability-focused DevOps.
  • Proficient with centralized log platforms (ELK/OpenSearch, Datadog Logs, Loki).
  • Designs and operates SLOs/SLIs and incident-response processes.

Responsibilities

  • Own centralized log management, ingestion, parsing, and retention.
  • Build and tune alerting with tiered thresholds to reduce noise.
  • Perform log analysis across Linux and Windows to diagnose incidents.
  • Drive MTTD under 15 minutes with improved signals and runbooks.
  • Define and own SLOs/SLIs and error budgets for customer-facing services.
  • Lead root-cause analysis and durable fixes across stack layers.
  • Collaborate with product squads to instrument services and surface signals.
  • Maintain runbooks, dashboards, and operational procedures.

Skills

SRE experience
Log analysis
Automation
Python/Bash/Go
Cross-team collaboration
Observability

Tools

ELK/OpenSearch
Datadog
Loki
Terraform

Job description

NEXT Ventures is a global fintech group powering FundedNext - one of the world's fastest-growing proprietary trading platforms - and FNmarkets, a regulated CFD brokerage. Across offices in Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai, we build and operate the technology that lets traders access global markets at scale. Our Platform Engineering team owns the infrastructure, reliability, and observability backbone that every product squad depends on.

Your Role in Our Mission

As our Site Reliability Engineer, you are the dedicated specialist who keeps our services observable, fast, and resilient. You own the centralized logging and alerting backbone, drive service-level optimization across the stack, and perform log analysis across both Linux and Windows environments. Working within the Platform Engineering squad, you execute the reliability initiatives that free the squad lead to focus on architecture - and you are the reason incidents are short, signals are clean, and detection is fast.

This role is distinct from our DevSecOps Engineer: DevSecOps builds and secures the platform foundation; you measure, detect, and optimize on top of it.

How You’ll Make an Impact

Own and operate the centralized log management platform - ingestion, parsing, structured logging standards, and retention across all services.

Build and tune alerting with tiered thresholds - catching real problems early while minimizing noise and alert fatigue.

Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.

Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.

Service-Level Optimization & Reliability

Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers - profile before guessing, measure every fix.

Define, implement, and own SLOs, SLIs, and error budgets for critical customer-facing services, and drive improvements against them.

Build deep observability with Datadog - APM, dashboards, monitors, log management, and SLO tracking.

Lead reliability and performance root-cause analysis and drive durable fixes.

Support load testing and capacity planning - identify breaking points before traffic growth causes production issues.

Operations & Cross-Team Support

Participate in a shared on-call rotation with solid runbooks and blameless post-incident reviews.

Continuously reduce manual toil through automation and better tooling.

Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals.

Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance.

What You Bring

5-7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work.

Strong hands-on experience with centralized log management platforms - ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.

Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.

Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.

Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.

Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services.

Can isolate and resolve latency and inefficiency across edge, application, and backend layers - profiles before guessing, measures every fix.

Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.

Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team.

Evidence-driven - profiles and measures before guessing; validates every optimization against before/after data.

Reliability-oriented - treats detection speed and signal quality as first-class engineering problems.

Good communicator - works with product squads to define SLIs and explains reliability constraints clearly.

You actively use modern AI agentic workflows daily - not limited to Copilot autocomplete.

You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools.

You are comfortable with project-level AI configuration (CLAUDE.md, rules files), agentic task delegation, and AI-driven code review.

You think in terms of 5-10x productivity through AI-augmented development - and you can demonstrate it.

Your Journey After Applying

Stage 1 - TA Interview

Stage 4 - Head of IT Interview

Why Join NEXT

Work on reliability challenges at real scale - 100M+ row data stores, multi-region traffic, high-frequency trading infrastructure.

A team that treats observability and detection speed as first-class engineering problems, not afterthoughts.

Flat structure - your work directly shapes how the platform operates, not filtered through layers of process.

Offices across Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai - a genuinely global engineering team.

Competitive compensation benchmarked to your market, with room to grow as the team scales.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Mastering Observability & Reliability at Scale
Senior SRE: Mastering Observability & Reliability at Scale

NEXT Ventures • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Global engineering team
Competitive compensation
Software Engineer, Trading Platform (AI-Native)
Software Engineer, Trading Platform (AI-Native)

NEXT Ventures • Kuala Lumpur

On-site
MYR 50,000 - 80,000
Senior Full Stack Engineering Lead-Brokerage Industry
Senior Full Stack Engineering Lead-Brokerage Industry

NEXT Ventures • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Food and Beverage
Entertainment
Learning Opportunities
+2
Software Engineer, Traders Experience (AI-Native)
Software Engineer, Traders Experience (AI-Native)

NEXT Ventures • Kuala Lumpur

On-site
MYR 120,000 - 180,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
SRE Lead
SRE Lead

Chubb • Malaysia

On-site
MYR 300,000 - 420,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Full Stack Developer - Partner Experience (AI Native)
Full Stack Developer - Partner Experience (AI Native)

NEXT Ventures • Kuala Lumpur

On-site
MYR 240,000 - 320,000
SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
Product Manager — Trading Platform, Rules & Risk
Product Manager — Trading Platform, Rules & Risk

NEXT Ventures • Kuala Lumpur

On-site
MYR 120,000 - 240,000