SmartX provides leading hyperconvergence, storage and management products to help our customers build the enterprise cloud infrastructure. We help achieve lower TCO, simpler architecture, and more stability in virtualization, private cloud, disaster recovery, VDI, Dev & Test, remote branch offices, edge and other scenarios.
Since 2013, SmartX is dedicated to building a powerful and easy-to-use next-generation IT infrastructure, to provide a solid foundation for enterprise digital transformation. Our products are empowering multiple industries across finance, healthcare, manufacturing, retail and real estate, serving industry leaders including Bank of Communications, Taikang Insurance Group, Guotai Junan Securities, Haier, BOE, Evergrande Real Estate, SBS, and Cafe24.
SmartX has built a wide services and sales network across China, Korea, South-East Asia, Middle East, and Europe. Its products and services have gained recognition from Gartner, IDC, Forrester, and GigaOm: SmartX is Gartner Peer Insights Customer’s Choice in the APAC HCI software market, and also No.1 in China HCI software market according to IDC.
To learn more please visit: https://www.smartx.com/hk-mo/
About the Role
We are looking for a highly capable Customer Support Engineer to provide advanced technical support for customers operating SmartX’s enterprise cloud and hyperconverged infrastructure platform. This role owns complex technical issues across virtualization, distributed storage, networking, security, backup, disaster recovery, and platform operations. The successful candidate combines Professional Services delivery capability with deeper troubleshooting expertise, strong incident ownership, and the ability to manage high-priority production issues.
Key Responsibilities
- Provide advanced remote and on-site technical support for SmartXcustomers across Singapore and the wider region.
- Take technical ownership of complex customer issues from investigation through resolution, root-cause analysis, and closure.
- Diagnose issues across compute, virtualization, distributed storage, networking, security, management plane, backup, replication, disaster recovery, and hardware integration layers.
- Handle high-severity production incidents involving service degradation, platform unavailability, data availability, performance, migration, or upgrade failures.
- Perform incident triage, impact assessment, evidence collection, containment, recovery planning, escalation, and customer communications.
- Analyse logs, alerts, metrics, configuration data, topology, version details, and operational history to isolate fault domains and identify root causes.
- Work with Engineering and Product teams to reproduce defects, validate fixes, assess product behaviour, and track resolution progress.
- Provide guidance on configuration, upgrades, patching, capacity planning, performance optimisation, security hardening, and operational best practices.
- Support complex upgrade, expansion, migration, and DR activities including pre-checks, risk reviews, change planning, validation, and post-change verification.
- Produce detailed case notes, RCA reports, knowledge-base articles, troubleshooting guides, and customer-facing technical summaries.
- Mentor Resident Engineers and contribute to the technical development of Professional Services and partner teams.
- Identify recurring product, process, documentation, and usability issues and provide actionable feedback internally.
Technical Requirements
Core Platform and Infrastructure Expertise
- Strong hands-on expertise in virtualization, hyperconverged infrastructure, private cloud, distributed storage, and enterprise data-centre operations.
- Deep understanding of KVM, VMware vSphere, OpenStack, HCI, or similar enterprise virtualization and cloud infrastructure platforms.
- Strong Linux troubleshooting capability across logs, services, CPU, memory, file systems, disk and network diagnostics, kernel messages, and resource contention.
- Proven ability to isolate faults across hosts, hypervisors, virtual machines, storage, networks, management services, external integrations, and hardware components.
Storage, Performance, and Resilience
- Advanced knowledge of replication, striping, consistency, failure domains, rebuild processes, snapshots, cloning, thin provisioning, and capacity management.
- Strong capability in diagnosing storage performance using IOPS, throughput, latency, queue depth, disk utilisation, workload patterns, and network behaviour.
- Experience troubleshooting cluster health, node failures, storage availability, virtual volumes, capacity anomalies, and performance degradation.
- Strong understanding of backup, restore, replication, failover, failback, DR testing, and business-continuity processes.
- Experience with synchronous or asynchronous replication, stretched clusters, active-active architectures, or VM-level DR is highly desirable.
Networking, Security, and Support Engineering
- Advanced troubleshooting across TCP/IP, VLANs, routing, MTU, bonding/LACP, DNS, firewalls, virtual switching, traffic isolation, and network connectivity.
- Familiarity with RDMA, RoCE, SR-IOV, vGPU, GPU virtualization, multi-path networking, NIC bonding, and high-performance network configuration is highly desirable.
- Understanding of secure access, SSH hardening, encryption in transit and at rest, key management, auditability, traffic monitoring, and security baseline enforcement.
- Strong technical case-management discipline covering issue classification, priority assessment, escalation, evidence collection, customer updates, and resolution tracking.
- Proven ability to prepare complete and actionable escalation packages for Engineering teams.
- Strong Shell scripting; Python, Ansible, APIs, log analytics, monitoring automation, ITSM, incident management, problem management, and change management experience is preferred.
Experience and Qualifications
- Bachelor’s Degree in Computer Science, Information Technology, Engineering, or a related discipline; equivalent practical experience will also be considered.
- 5-8+ years of experience in enterprise technical support, infrastructure engineering, virtualization, private cloud, HCI, storage, cloud platforms, or data-centre operations.
- At least 3 years of hands‑on experience resolving complex issues in production enterprise environments.
- Demonstrated experience managing critical incidents involving availability, performance, storage, networking, virtualization, or disaster recovery.
- Prior Professional Services, implementation, systems integration, or technical consulting experience is strongly preferred.
- Experience supporting enterprise customers under defined SLAs and incident‑management processes is highly desirable.
- VMware VCP/VCAP, RHCE, CCNP, HCIP, cloud infrastructure, or storage certifications are advantageous.
- Working proficiency in English, including clear customer communication and professional technical documentation; additional Asian language capability is advantageous.
Desired Competencies
- Strong ownership mindset and commitment to customer outcomes.
- Excellent analytical and troubleshooting skills with an evidence-based decision-making approach.
- Ability to communicate calmly, clearly, and professionally during high‑pressure incidents.
- Ability to explain complex technical risks and remediation options to varied stakeholder groups.
- Strong cross‑functional collaboration with Engineering, Product, QA, Professional Services, Sales Engineering, and external partners.
- Ability to mentor junior engineers and contribute to technical excellence and knowledge sharing.
- Willingness to participate in support rotations and respond to critical customer issues outside standard business hours when required.
Your application will include the following questions:
- Which of the following statements best describes your right to work in Singapore?
- What's your expected monthly basic salary?
- How many years' experience do you have as a Customer Support Engineer?
- Which of the following types of qualifications do you have?
- Which of the following languages are you fluent in?