DevOps Engineer

Manus AI

Singapore

On-site

SGD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity & Shared success
Unlimited Manus Tokens
Frontier AI projects

Job summary

Manus AI is seeking a hands-on DevOps/SRE to operate and build our multi-cloud infrastructure, including a self-hosted E2B sandbox. You will manage Kubernetes/Docker clusters, Nomad/Consul, and Firecracker, ensuring availability, scaling, and cost-efficiency.

You will implement CI/CD pipelines (Jenkins, Argo CD, GitHub Actions), monitoring (Prometheus/Grafana/ELK), and IaC (Terraform, Ansible). On-call rotation included for incident response.

Qualifications

  • 3–4+ years hands-on Linux system operations, DevOps, or SRE experience.
  • Strong Linux fundamentals, networking basics, and security best practices.
  • Deep containerization expertise (Kubernetes, Docker) with production clusters.
  • Experience building/managing CI/CD pipelines (Jenkins, GitHub Actions, Argo CD).
  • Familiarity with Nginx, MySQL, Redis, Kafka, Elasticsearch.
  • Proficiency in Shell and Python for automation.
  • Hands-on with at least two public clouds (AWS, Azure, GCP).
  • Familiarity with Prometheus, Grafana, ELK for monitoring/logging.
  • Terraform experience for infrastructure as code.
  • Ability to communicate in Mandarin and English for day-to-day technical work.

Responsibilities

  • Cluster Operations & Management across container clusters and middleware.
  • Own the self-hosted E2B cluster, including Nomad/Consul and Firecracker.
  • Ensure performance, scalability, and reliability of distributed systems.
  • Linux & Cloud Infrastructure: manage day-to-day ops and cloud resources (AWS/Azure/GCP).
  • Infrastructure Platform Development: CI/CD pipelines, monitoring, logging, and automation tooling.
  • Drive high availability with 24/7 on-call, incident response, RCA, and postmortems.
  • Automation & Process Improvement: self-service tools and IaC practices (Terraform/Ansible).
  • Maintain documentation, SOPs, and runbooks for operations.

Skills

Linux fundamentals
SRE/DevOps practices
Shell & Python
Cloud platforms experience
Mandarin/English communication

Tools

Kubernetes
Docker
Jenkins
Argo CD
GitHub Actions
Terraform
Prometheus
Grafana
ELK
Nginx
MySQL
Redis
Kafka
Elasticsearch
Nomad
Consul
Firecracker

Job description

Location: Beijing/Singapore

About the Role

Operate and build the platform for our multi-cloud infrastructure and container clusters, including our self-hosted E2B (AI agent sandbox) cluster, and keep production services highly available, scalable, and cost-efficient through automation and platform engineering.

Responsibilities

  • Cluster Operations & Management
  • Manage and maintain container clusters (Kubernetes, Docker) and open-source middleware clusters (Kafka, Redis, Elasticsearch) across multiple business units.
  • Own the self-hosted E2B (AI agent sandbox) cluster, including the Nomad / Consul orchestration layer and the Firecracker microVM runtime.
  • Ensure the performance, scalability, and reliability of distributed systems.
  • Linux & Cloud Infrastructure
  • Handle day-to-day operations of Linux servers, including hardening, patch management, backups, and performance tuning.
  • Manage cloud resources across AWS, Azure, and GCP, ensuring reliability, security, and cost efficiency.
  • Infrastructure Platform Development
  • Design, build, and continuously improve our infrastructure operations platform.
  • Develop and maintain infrastructure management, CI/CD GitOps pipelines (Jenkins / Argo CD), monitoring and alerting (Prometheus / Grafana), and centralized logging (ELK).
  • Drive platform standardization and automation.
  • High Availability & Reliability
  • Ensure the highest availability of production services through proactive monitoring and incident response.
  • Take part in a 24/7 on-call rotation, and lead incident response, root cause analysis, and postmortems.
  • Implement and maintain SLA/SLO frameworks and reliability engineering practices.
  • Automation & Process Improvement
  • Build self-service tools and workflows to improve team productivity.
  • Establish best practices for infrastructure as code (Terraform / Ansible) and configuration management.
  • Maintain thorough documentation, SOPs, and runbooks.

Requirements

  • 3–4+ years of hands-on experience in Linux system operations, DevOps, or SRE (no degree requirement; ability comes first).
  • Strong Linux fundamentals, with knowledge of networking basics and system security best practices.
  • Deep expertise in containerization (Kubernetes, Docker), with experience operating production-grade clusters.
  • Experience building and managing CI/CD pipelines (Jenkins, GitHub Actions / Argo CD).
  • Familiarity with common infrastructure components: Nginx, MySQL, Redis, Kafka, and Elasticsearch.
  • Proficiency in Shell and Python for scripting and automation.
  • Hands-on operations experience with at least two public clouds among AWS, Azure, and GCP.
  • Familiarity with infrastructure monitoring, logging, and observability tools (Prometheus, Grafana, ELK).
  • Hands-on experience with Terraform.
  • Ability to collaborate on technical work in Mandarin and to read, write and participate in day-to-day technical communication in English.

Preferred Qualifications

  • Familiarity with or exposure to E2B / AI agent sandbox infrastructure (Firecracker microVM, Nomad, Consul).
  • Experience with service mesh architecture (Istio) and eBPF.
  • CKA/CKAD or AWS/Azure/GCP professional certifications.

Manus excels at various tasks in work and life, getting everything done while you rest at Manus AI.

What we offer:

  • Build at the Frontier of AI Agents - Work with a fast-paced team to turn frontier AI breakthroughs into real-world impact.
  • Equity & Shared success - Our share incentive plan enables employees to participate in Manus’s long-term growth and share in the value we create together.
  • Unlimited Manus Tokens - Enjoy unlimited Manus tokens to experiment, build, and supercharge your productivity.

If you're passionate about cutting-edge technology and making a real impact, we’d love to hear from you!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multi-Cloud DevOps Engineer — Kubernetes, SRE & Automation
Multi-Cloud DevOps Engineer — Kubernetes, SRE & Automation

Manus AI • Singapore

On-site
SGD 120,000 - 180,000
Equity & Shared success
Unlimited Manus Tokens
Frontier AI projects
DevOps / SRE Engineer – AI Cloud
DevOps / SRE Engineer – AI Cloud

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 130,000
DevSecOps Engineer
DevSecOps Engineer

CAPGEMINI SINGAPORE PTE. LTD. • Singapore

On-site
SGD 80,000 - 120,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

Clera • Singapore

On-site
SGD 191,000 - 319,000
Visa sponsorship
On-site in Singapore
DevSecOps Engineer - #1634
DevSecOps Engineer - #1634

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 140,000 - 210,000
Backend Engineer, AI (Agent Systems)
Backend Engineer, AI (Agent Systems)

ActAI • Singapore

On-site
SGD 120,000 - 180,000
DevOps Engineer – Investment Technology (SG/HK)
DevOps Engineer – Investment Technology (SG/HK)

Io Tech Solutions Limited • Singapore

On-site
SGD 90,000 - 140,000
Agent Evaluation Engineer
Agent Evaluation Engineer

Manus AI • Singapore

On-site
SGD 90,000 - 150,000
Equity plan
Unlimited Manus Tokens
Frontier AI projects
DevOps Engineer – Investment Technology
DevOps Engineer – Investment Technology

IO TECH SOLUTIONS LIMITED • Singapore

On-site
SGD 120,000 - 160,000
#EG Cloud Engineer / Architect – AI Infrastructure
#EG Cloud Engineer / Architect – AI Infrastructure

NCS • Singapore

On-site
SGD 180,000 - 240,000