Arcfra is looking for a Sustaining Software Engineer withexperience in virtualization, cloud infrastructure, and softwaredevelopment using Python or Go.
You will be responsible for maintaining and improving ourvirtualization and cloud platform products. The role focuses ontroubleshooting production issues, resolving customer-reporteddefects, developing software fixes, and improving the reliabilityand serviceability of the platform.
This is not a traditional technical support role. You will workdirectly with source code, system logs, APIs, operating systems,virtualization components, and distributed cloud services. Youwill collaborate with support engineers, field teams, QA, andcore development teams to diagnose complex issues and deliverproduction-quality solutions.
Key Responsibilities
- Maintain and improve virtualization and cloud platformproducts.
- Investigate customer-reported issues involving virtualmachines, hosts, clusters, networking, storage, orchestration,and management services.
- Reproduce production issues and perform detailed root-causeanalysis.
- Develop, test, review, and deliver bug fixes using Python or Go.
- Troubleshoot problems across application, service, operatingsystem, virtualization, networking, and storage layers.
- Analyze logs, traces, metrics, API requests, service states,database records, and system configurations.
- Diagnose issues involving virtual machine lifecycle management,scheduling, high availability, live migration, resourcemanagement, and cluster operations.
- Work closely with technical support and field engineering teamsto resolve complex customer escalations.
- Provide technical guidance, workarounds, diagnostic procedures,and corrective actions for production issues.
- Collaborate with core engineering teams on architectural defectsand long-term product improvements.
- Develop scripts, diagnostic tools, automated tests, andtroubleshooting utilities.
- Improve product monitoring, logging, observability,upgradeability, and operational reliability.
- Participate in release validation and safely backport fixes tosupported product versions.
- Write clear technical documentation, root-cause analysisreports, and internal knowledge-base articles.
Required Qualifications
- Bachelor's degree or above in Computer Science, ComputerEngineering, Software Engineering, or a related field.
- Professional software development experience using Python orGo.
- Strong Linux administration and troubleshooting skills.
- Solid understanding of operating systems, processes, threads,networking, storage, and distributed services.
- Experience with virtualization technologies or cloudinfrastructure.
- Familiarity with one or more of the following:
- KVM and QEMU
- libvirt
- VMware vSphere or ESXi
- OpenStack
- Container runtimes
- Cloud management platforms
- Hyper-converged infrastructure
- Experience troubleshooting REST APIs, background services, ordistributed control-plane components.
- Strong debugging, problem-solving, and root-cause analysisskills.
- Ability to understand and modify an existing productioncodebase.
- Ability to communicate technical findings clearly in writtenand spoken English and Chinese.
- Willingness to work with customer-facing teams on complexproduction incidents.
- Must be a Singapore citizen, permanent resident, or hold avalid Singapore work permit.
Preferred Qualifications
- Experience developing or maintaining virtualization, cloudmanagement, or infrastructure software.
- Understanding of virtual machine lifecycle management, CPU andmemory virtualization, virtual networking, and virtual storage.
- Familiarity with high availability, live migration, clusterscheduling, resource management, and failure recovery.
- Familiarity with Linux networking, Open vSwitch, bridges,VLANs, VXLAN, routing, and software-defined networking.
- Experience with storage protocols or systems used byvirtualization platforms.
- Experience with diagnostic and observability tools such as GDB,strace, tcpdump, perf, Prometheus, Grafana, or distributedtracing systems.
- Experience writing automated tests, integration tests, andtroubleshooting tools.
- Previous experience in sustaining engineering, escalationengineering, product support engineering, site reliabilityengineering, or cloud operations.
What We Value
- Strong ownership of customer-impacting technical problems.
- The ability to troubleshoot systematically across multiplelayers.
- A software engineering approach to product maintenance andcustomer support.
- Attention to backward compatibility, upgrade safety, andproduction reliability.
- Clear communication during complex or high-priority incidents.
- The ability to balance short-term customer recovery withlong-term engineering improvements.