Senior DevOps Engineer (AI & Production Infrastructure)

Deriv

Cyberjaya

On-site

MYR 66,960 - 133,920

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deriv is seeking an experienced production infrastructure and reliability engineer to own end-to-end systems that run millions of trades. You’ll design, build, and operate cloud, container, and database environments with AI-assisted tooling, automate delivery, and harden security alongside a dedicated security team.

You’ll lead incident response, implement self-healing patterns, and continuously improve availability while advancing autonomous AI-based operational capabilities.

Qualifications

  • 4+ years of DevOps, SRE, or production operations experience.
  • Hands-on experience shipping and supporting real production infrastructure.
  • Strong background in incident response, on-call operations, escalation handling, and high-availability web services.
  • Infrastructure automation and configuration management experience: Terraform, CloudFormation, Ansible, or equivalent.
  • Practical containerization experience with Docker and Kubernetes.
  • Cloud infrastructure experience with AWS, GCP, Azure, or similar, at scale.
  • Working knowledge of Linux and Windows Server environments, networking, databases, security, monitoring, and CI/CD.
  • Scripting or programming ability in Bash, Python, Go, PowerShell, or similar.
  • AI-first operating style using tools to build, troubleshoot, document, and improve infrastructure.
  • Ability to move fast while protecting reliability: prototype, test with real workloads, deploy carefully, learn from incidents.

Responsibilities

  • Production Infrastructure — Cloud, container, database, monitoring, and CI/CD environments for high-availability services; design it, not just patch it.
  • AI-Native Delivery — Automation design, infrastructure-as-code generation, testing, refactoring, documentation, and runbook creation with AI tooling.
  • Incident Response & Resilience — Alerting logic, remediation scripts, self-healing patterns, circuit breakers, fault-tolerant architecture.
  • Security Operations — Hardening, intrusion detection, configuration audits, and AI-enhanced threat detection with security team.
  • End-to-End Ownership — Define problem, architect, implement, deploy safely, monitor, and iterate.
  • Monitoring & Observability — Maintain Datadog, Grafana, and custom observability systems with AI-assisted analysis.
  • Incident Diagnosis — Diagnose production incidents, coordinate with developers, and drive durable infra improvements.
  • Autonomous Operations — Prototype and deploy autonomous AI systems for operations, including always-on agents for remediation.

Skills

DevOps
SRE practices
Incident response
On-call operations
Automation

Tools

Terraform
CloudFormation
Ansible
Docker
Kubernetes
AWS
GCP
Azure
Datadog
Grafana
PostgreSQL
Redis
Linux
Windows Server
Python
Go
PowerShell

Job description

We're hiring someone who wants production infrastructure that has to hold at 3am, every time, for millions of traders who never stop trading. Real uptime targets. Real incidents. Real consequences when something breaks. You'll own that infrastructure end to end, not just the parts that are already stable.

Why This Matters

Trading for Anyone, Anywhere, Anytime means services that don't sleep, across time zones and regulatory regimes. That scale doesn't run on tribal knowledge and manual runbooks. We're already running AI-assisted monitoring that catches problems before they page anyone, automation that remediates known failure patterns on its own, and security workflows that flag threats faster than a human scanning dashboards ever could. Not experiments. Systems carrying production traffic right now. You won't be maintaining this from a distance. You'll be designing it, breaking it, fixing it, and deciding what gets built next.

Why Deriv

We're in production, not planning.

  • Autonomous security analysts already triaging alerts and correlating threats against historical patterns
  • Dozens of fraud detection models running continuously against real transactions
  • Automated security review on every pull request, every day
  • Infrastructure-as-code, monitoring logic, and runbooks increasingly generated, tested, and documented with AI as a normal part of the workflow, not a side experiment

We share what we learn. Deriv is where we write about AI in production, including what breaks and what we figured out the hard way. You'll own systems that are already handling real transactions, not prototypes waiting for product-market fit.

What You’ll Do

This role owns outcomes across production infrastructure and reliability engineering, with regular cross-functional work alongside the security team:

  • Production Infrastructure — Cloud, container, database, monitoring, and CI/CD environments for high-availability services. You design it, not just patch it.
  • AI-Native Delivery — Automation design, infrastructure-as-code generation, testing, refactoring, documentation, and runbook creation, with AI tooling built into how the work gets done.
  • Incident Response & Resilience — Alerting logic, remediation scripts, self-healing patterns, circuit breakers, and fault-tolerant architecture for systems that can't afford to go down.
  • Security Operations — Hardening, intrusion detection, configuration audits, and AI-enhanced threat detection, built jointly with the security team.
  • End-to-End Ownership — Take a vague operational problem, design the architecture, implement the solution, deploy it safely, monitor the outcome, and keep iterating once it's live.
  • Monitoring & Observability — Maintain and extend Datadog, Grafana, and custom observability systems, with AI-assisted analysis to catch anomalies and potential failures early.
  • Incident Diagnosis — Diagnose and resolve production incidents across complex systems, coordinate with developers, and turn what you learn into durable infrastructure improvements, not one-off fixes.
  • Autonomous Operations — Explore, prototype, and deploy autonomous AI systems for operations, including always‑on agents that maintain context, investigate anomalies, and take approved remediation actions.
Who You Are

You work across all three paradigms that make production infrastructure actually reliable:

  • Deterministic systems — the Terraform, the CI/CD pipeline, the database that has to be right every time
  • Predictive systems — the anomaly detection that flags "this looks wrong" before a human would notice
  • Agentic systems — AI tools that draft the automation, the docs, the first pass at a fix, so your judgment goes toward the decisions that actually need it
Required Experience
  • 4+ years of DevOps, SRE, infrastructure engineering, or production operations experience
  • Hands‑on experience shipping and supporting real production infrastructure — not only proofs of concept or sandbox environments
  • Strong practical background in SRE practices, incident response, on‑call operations, escalation handling, and high‑availability web service architecture
  • Infrastructure automation and configuration management experience: Terraform, CloudFormation, Ansible, or equivalent
  • Practical containerization experience with Docker and Kubernetes
  • Cloud infrastructure experience with AWS, GCP, Azure, or similar, at scale
  • Working knowledge of Linux and Windows Server environments, networking, databases, security, monitoring, and CI/CD
  • Scripting or programming ability in Bash, Python, Go, PowerShell, or similar
  • A proven AI‑first operating style using tools such as Claude Code, Cursor, Codex, or Kiro to build, troubleshoot, document, and improve infrastructure — not just to autocomplete a function, but to design and reason through a system
  • The ability to move fast while protecting reliability: prototype, test with real workloads, deploy carefully, learn from incidents, and improve continuously
Bonus Points
  • Database operations experience with PostgreSQL, Redis, Aurora, CloudSQL, Supabase, or MS‑SQL
  • Linux system hardening, Windows Server administration, IIS, or MS‑SQL administration
  • Information security experience: data protection, firewalls, IDS/IPS, DDoS protection, security standards
  • CI tooling experience with Jenkins, Travis CI, CircleCI, or similar
  • Experience building monitoring, alerting, or remediation agents
  • Experiments with autonomous agent frameworks such as the Claude Agent SDK
Tech Stack

Terraform, CloudFormation, Ansible, Docker, Kubernetes, AWS/GCP/Azure, Datadog, Grafana, PostgreSQL/Redis, Linux and Windows Server, Bash/Python/Go/PowerShell, Claude Code and similar AI‑assisted development tools.

The Honest Reality

This is demanding work. You'll own outcomes where the fix isn't obvious and the clock is running. You'll make deploy‑or‑wait calls with incomplete information and live with what happens next. You'll build automation that works perfectly until it meets a production edge case nobody predicted.

But you'll build infrastructure that's actually load‑bearing, not a proof of concept waiting for approval. You'll set the standards for how AI‑assisted infrastructure work gets done here, and that work will outlast whatever ticket it started as.

If you want a role with clear boundaries and someone else owning the pager, this isn't it. If you want real ownership over systems that can't afford to fail, it might be.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineering Tech Lead
Data Engineering Tech Lead

Deriv • Cyberjaya

On-site
MYR 100,000 - 150,000
Senior Engineer (Infra)
Senior Engineer (Infra)

Pertama Partners • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Offensive Security Engineer
Senior Offensive Security Engineer

Deriv • Cyberjaya

Hybrid
MYR 180,000 - 280,000
Tech Lead - Disaster Recovery and Incident Management
Tech Lead - Disaster Recovery and Incident Management

Deriv • Cyberjaya

On-site
MYR 180,000 - 260,000
Senior DevOps Engineer: AI-Powered Production Infra
Senior DevOps Engineer: AI-Powered Production Infra

Deriv • Cyberjaya

On-site
Staff Applied AI Engineer
Staff Applied AI Engineer

Deriv.com • Cyberjaya

On-site
MYR 180,000 - 300,000
Team Lead - Application DevSecOps & SRE
Team Lead - Application DevSecOps & SRE

Dialog Group Berhad • Petaling Jaya

On-site
MYR 240,000 - 360,000
Senior Data Engineer
Senior Data Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 70,000 - 90,000
Lead DevOps Engineer
Lead DevOps Engineer

RinggitPlus • Malaysia

Hybrid
MYR 180,000 - 300,000
AI Product Designer (Builder)
AI Product Designer (Builder)

Deriv • Cyberjaya

On-site
MYR 120,000 - 180,000