Service Delivery Lead - L3 AWS Cloud

PCCW SOLUTIONS INSYS PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

PCCW SOLUTIONS INSYS PTE. LTD. in Singapore seeks a Senior SME for multi-cloud infrastructure operations. You will own operations across AWS, Azure, and GCP, lead incident response, and drive automation to meet 24/7 availability and strict SLAs.

The role requires hands-on expertise with patching, security, and DevSecOps practices, plus mentoring of L2 engineers and close coordination with security, compliance, and development teams in a fast-paced, government-ready environment.

Responsibilities

  • Operate, maintain, and improve cloud-native production environments across AWS, Azure, and GCP.
  • Provide hands-on technical leadership across cloud services including but not limited to AWS, Azure, GCP.
  • Lead incident response, root-cause analysis, and resolution of complex infra and app issues.
  • Oversee OS patching and lifecycle management for RHEL and Windows Server.
  • Coordinate patch approvals with security, compliance, and stakeholders.
  • Deploy, operate, and troubleshoot applications across Windows and Linux in cloud/hybrid environments.
  • Implement and enhance monitoring, logging, and alerting frameworks.

Job description

Responsibilities:
Multi-Cloud Infrastructure Operations
  • Operate, maintain, and continuously improve cloud-native production environments acrossAmazon Web Services (AWS),Microsoft Azure, and Google Cloud Platform (GCP).
  • Provide hands-on technical leadership across a broad range of cloud services, including but not limited to:
  • AWS: Lambda, ECS/EKS, FSx, Glue, SES, GuardDuty, WAF, Shield Advanced, Security Hub, KMS, Secrets Manager, SNS, SQS, EventBridge, API Gateway, EC2, S3, CloudWatch, Systems Manager
  • Azure: Virtual Machines, Azure Kubernetes Service (AKS), Azure Functions, Azure Storage, Azure Monitor
  • GCP: Compute Engine, Google Kubernetes Engine (GKE), Cloud Functions, Cloud Storage, Cloud Monitoring
  • Monitor, analyze, and troubleshoot infrastructure performance, availability, scalability, and cost efficiency across all cloud platforms.
  • Support both production and staging environments, ensuringadherence to 24/7 high-availability and reliability objectives, including strict SLA and SLO commitments.
  • Participate in a 24/7 shift rotation to provide round-the-clock operational coverage.
  • Provide hands-on technical support andguidance to L2 engineers, leading incident response, root-cause analysis, and resolution of complex infrastructure and application issues.
Operating System Lifecycle & Patch Management
  • Lead and overseeoperating system patching and lifecycle managementacross RHEL (v8–v10) andWindows Server (2016–2025)environments using tools such as AWS Systems Manager Patch Manager, Azure Update Management, WSUS, SCCM, and YUM/DNF.
  • Maintain strong foundational knowledge of Linux system administration, complemented by deep expertise in Windows (Wintel) operating system patching, hardening, and lifecycle management.
  • Plan, schedule, automate, and track patch deployments across development, staging, and production environments, ensuring consistency and repeatability.
  • Coordinate patch approvals with security, compliance, and business stakeholders to ensure alignment with organizational policies, risk frameworks, and audit requirements.
  • Execute monthly and quarterlypatching cycleswith minimal service disruption, adhering to defined change management and maintenance windows.
  • Performpost-patch validation, health checks, and remediation activities to confirm system stability, security posture, and operational readiness.
Application Deployment & Troubleshooting
  • Deploy, operate, and troubleshoot applications acrossWindows and Linux operating systemsin cloud-based and hybrid environments.
  • Provide OS-level diagnostics, performance tuning, and stability support to application teams, including CPU, memory, disk, network, and process-level analysis.
  • Partner closely with development and DevOps teams to identify, isolate, and resolveinfrastructure-, platform-, and OS-related application issuesthroughout the application lifecycle.
  • Implement, maintain, and continuouslyenhance application monitoring, logging, and alerting frameworksto ensure early issue detection and rapid incident response in production environments.
Security & Compliance
  • Execute and manageCIS (Center for Internet Security) control implementations and remediationsacross multi-cloud environments to strengthen security posture.
  • Performsecurity hardeningin accordance withCIS Benchmarks, industry best practices, andgovernment-mandated security baselines.
  • Conduct continuousvulnerability identification, assessment, and remediationusing tools such asTrend Micro Vision One,Qualys,Tenable, andAWS Config, ensuring timely risk mitigation.
  • Track, manage, and renewSSL/TLS certificatesacross all environments to prevent service disruptions and maintain secure communications.
  • Proactively identify and remediateEnd-of-Life (EOL) and End-of-Support (EOS)components, including operating systems, middleware, andAWS Lambda runtimes, to reduce security and compliance risks.
  • Support and maintain compliance withgovernment-grade security, audit, and regulatory requirements, including evidence collection, audit readiness, and remediation tracking.
Container & DevSecOps
  • Demonstrate strong working knowledge ofcontainer and orchestration technologies, includingDocker,Kubernetes, and managed container platforms such asAWS ECS/EKS,Azure AKS, andGoogle GKE.
  • Apply familiarity with DevSecOps principles and practices, including exposure to SHIP-HATS (Secure Hybrid Integration Pipeline – Hive Agile Testing Solutions) within the Singapore Government technology ecosystem.
  • Support and maintainCI/CD pipeline operations, ensuring seamless integration withsecurity scanning, vulnerability assessment, and compliance validation toolsacross the software delivery lifecycle.
ITIL & Service Management
  • Adhere to ITIL-based service management processes, including Incident, Problem, Change, and Request Management, ensuring consistent and controlled service delivery.
  • Manage, prioritize, and resolveITSM ticketsusing platforms such asServiceNow,Jira, or equivalent tools, meeting defined service commitments and response targets.
  • Drive timely and effectiveticket escalation and coordinationbetween engineering teams, service owners, and stakeholders to ensure prompt issue resolution.
  • Coordinate and governchange management activities, including preparing change documentation and participating inChange Advisory Board (CAB)reviews, providing guidance and oversight to junior engineers.
  • Monitor, maintain, and report againstService Level Agreements (SLAs)andOperational Level Agreements (OLAs)to ensure service performance, accountability, and continuous improvement.
Documentation & Knowledge Management
  • Create, maintain, and continuously update comprehensive infrastructure runbooks, system documentation, architecture design artefacts, and change-tracking logs for assigned applications and platforms.
  • Develop and standardizeStandard Operating Procedures (SOPs), operational guidelines, andknowledge base articlesto support consistent service delivery and efficient incident resolution.
  • Ensureaudit readinessthrough disciplined documentation practices, including version control, traceability, and alignment with security and compliance requirements.
  • Maintain accurateConfiguration Management Databases (CMDB)andasset inventories, ensuring alignment with deployed infrastructure and operational states.
Leadership & Mentorship
  • Providetechnical leadership, guidance, and mentorshiptoLevel 2 and junior engineers, fostering skill development, accountability, and operational excellence.
  • Lead and facilitatetechnical discussions, design reviews, and architecture governance forums, ensuring solutions align with organizational standards, security requirements, and best practices.
  • Plan and deliverknowledge transfer sessions, technical training, and operational walkthroughsto uplift team capability and reduce single points of failure.
  • Act as theprimary escalation pointfor complex or high-impact technical issues, driving root-cause analysis and sustainable long-term remediation.
  • Championcontinuous improvement initiatives, automation adoption, and operational best practices to enhance service reliability, efficiency, and team maturity.
Soft Skills & Competencies
  • Problem Solving– Demonstrates advanced troubleshooting and analytical skills to diagnose and resolve complex issues across multi-cloud and hybrid environments.
  • Communication– Communicates clearly and effectively with technical and non-technical audiences, including engineers, stakeholders, and senior management.
  • Leadership– Provides direction and influence to guide teams, drive technical initiatives, and deliver high-quality outcomes.
  • Collaboration– Works effectively across engineering, security, operations, and business teams to achieve shared objectives.
  • Adaptability– Remains responsive and effective in fast-changing, dynamic, and high-pressure environments.
  • Accountability & Attention to Detail– Takes ownership of service delivery and outcomes, ensuring accuracy, reliability, security, and compliance in all implementations.
  • Customer Focus– Maintains a service-oriented mindset with strong stakeholder management and a commitment to meeting business and customer needs.
  • Continuous Learning– Proactively stays current with evolving cloud technologies, security standards, and industry best practices.
  • Resilience– Performs effectively under pressure, particularly during incidents, outages, and critical operational situations.
  • Mentorship– Actively develops, coaches, and supports junior engineers to build team capability and long-term sustainability.

This Subject Matter Expert (SME) role requires the individual to consistently demonstrate the following behaviors and capabilities:

  • Deep proficiency in Amazon Web Services (AWS), with solid working knowledge ofMicrosoft AzureandGoogle Cloud Platform (GCP)to support and guide multi-cloud operations is a must.
  • Proven ability to operate withinuptime-critical, security-sensitive, and compliance-driven environments, maintaining service reliability and operational excellence.
  • Strongtechnical leadership and mentorship capabilities, providing guidance, oversight, and skills development for junior and mid-level engineers.
  • A proactive mindset focused onincident prevention, continuous improvement, and adoption of best practices to enhance system stability and resilience.
  • Acalm, structured, and methodical approach to incident management, with strict adherence tochange management,incident response, and escalation procedures.
  • Anaudit-readiness mindset, supported by rigorous documentation, traceability, and evidence-based operational practices.
  • Ability todrive technical escalations, coordinate cross-functional resolution efforts, and manageclear, timely stakeholder communicationsduring incidents and service-impacting events.
  • Demonstrated experience working withinSingapore Government technology frameworks, policies, and regulatory standards, including alignment with public-sector governance and security requirements.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

[LPS] Service Delivery Lead - Cloud
[LPS] Service Delivery Lead - Cloud

LPS • Singapore

On-site
SGD 75,000 - 100,000
Cloud Engineer
Cloud Engineer

WORLD PARTNERS SOLUTION PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Service Delivery Director - GCC/GCC+ (AWS)
Service Delivery Director - GCC/GCC+ (AWS)

PCCW SOLUTIONS INSYS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
L2 Cloud Linux Engineer
L2 Cloud Linux Engineer

ENGGSOL PTE. LTD. • Singapore

On-site
SGD 70,000 - 100,000
Cloud Operations Support Engineer
Cloud Operations Support Engineer

SEDHA CONSULTING PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Project Manager (Operations)
Project Manager (Operations)

UFINITY Pte Ltd. • Singapore

On-site
SGD 90,000 - 130,000
Senior Software Engineer (Operations)
Senior Software Engineer (Operations)

UFINITY PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
Project Manager (Operations)
Project Manager (Operations)

ufinity pte ltd • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

ARYAN SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 60,000 - 90,000
Technical Manager
Technical Manager

COMBUILDER PTE LTD • Singapore

On-site
SGD 120,000 - 190,000