A complete application in a minute — tailored resume and cover letter, ready to send.
Lenovo Malaysia seeks a skilled DevOps/SRE engineer to deploy, manage, monitor, and troubleshoot Kubernetes clusters and containerized apps in production. You will optimize CI/CD pipelines, develop automation scripts, and work with infra as code to improve reliability and security.
The role requires 3+ years in DevOps or SRE with cloud-native experience and collaboration with development teams. Experience with AI/ML infrastructure is a plus.
Jora Malaysia will close on 9th September 2026. Thank you for being with us, we are cheering you on as you continue your career journey.
Deploy, manage, monitor, and troubleshoot Kubernetes (K8s) clusters and containerized applications in production environments.
Support and optimize CI/CD pipelines for software deployment and release management.
Develop automation scripts and tools to eliminate manual operational tasks and improve efficiency.
Manage day-to-day operational incidents, troubleshooting, and root cause analysis (RCA).
Monitor application and infrastructure performance, identify bottlenecks, and implement optimization plans.
Build and maintain observability solutions including logging, metrics collection, monitoring, and APM platforms.
Manage and maintain databases, including performance tuning, backup/recovery, and fault resolution.
Collaborate with development teams to improve deployment processes, system reliability, and operational excellence.
Support cloud-native infrastructure and AI/ML-related services.
Administer Linux-based environments and web platforms while ensuring high availability and security.
Implement infrastructure automation using modern DevOps tools and Infrastructure-as-Code principles.
Work with networking and security teams to maintain reliable and secure services.
Continuously evaluate emerging technologies and recommend improvements to systems and processes.
Experience supporting AI/ML platforms or AI-related infrastructure.
Experience building end-to-end observability solutions.
Strong understanding of GitOps methodologies.
Knowledge of Infrastructure as Code (Terraform, Ansible).
Experience in enterprise application operations and production support.
Familiarity with modern microservices and cloud-native environments.
Deploy, manage, monitor, and troubleshoot Kubernetes (K8s) clusters and containerized applications in production environments.
Support and optimize CI/CD pipelines for software deployment and release management.
Develop automation scripts and tools to eliminate manual operational tasks and improve efficiency.
Manage day-to-day operational incidents, troubleshooting, and root cause analysis (RCA).
Monitor application and infrastructure performance, identify bottlenecks, and implement optimization plans.
Build and maintain observability solutions including logging, metrics collection, monitoring, and APM platforms.
Manage and maintain databases, including performance tuning, backup/recovery, and fault resolution.
Collaborate with development teams to improve deployment processes, system reliability, and operational excellence.
Support cloud-native infrastructure and AI/ML-related services.
Administer Linux-based environments and web platforms while ensuring high availability and security.
Implement infrastructure automation using modern DevOps tools and Infrastructure-as-Code principles.
Work with networking and security teams to maintain reliable and secure services.
Continuously evaluate emerging technologies and recommend improvements to systems and processes.
Preferred Qualifications
Experience supporting AI/ML platforms or AI-related infrastructure.
Experience building end-to-end observability solutions.
Strong understanding of GitOps methodologies.
Knowledge of Infrastructure as Code (Terraform, Ansible).
Experience in enterprise application operations and production support.
Familiarity with modern microservices and cloud-native environments.
Strong troubleshooting and analytical skills.
Excellent problem-solving capabilities.
Ability to work independently in a fast-paced environment.
Strong communication and collaboration skills.
Self-motivated with a passion for learning new technologies.
Good time management and prioritization skills.
Customer-focused mindset with a strong sense of ownership.
Bachelor's degree in computer science, Information Technology, Engineering, or related discipline.
3+ years of experience in DevOps, Site Reliability Engineering (SRE), Cloud Operations, Platform Engineering, or IT Operations.
Hands-on experience supporting production applications and cloud-native infrastructure.