Arcfra is looking for a Sustaining Software Engineer with a strongbackground in distributed systems, storage technologies, andC/C++ development.
In this role, you will be responsible for maintaining andimproving our production-grade storage products. You willinvestigate complex software defects, troubleshootcustomer-reported issues, develop reliable fixes, and workclosely with support, QA, and engineering teams to ensure productstability.
This is a hands-on engineering position. It combines softwaredevelopment, production troubleshooting, root-cause analysis, andcustomer issue resolution. The ideal candidate enjoys working onmature and complex systems, understands how software behaves inreal-world production environments, and is comfortable debuggingproblems across multiple layers of the system.
Key Responsibilities
- Maintain and improve distributed storage and data infrastructureproducts.
- Investigate and resolve customer-reported product issues,including performance degradation, service interruptions,data-path failures, and unexpected system behavior.
- Reproduce production issues in laboratory or test environmentsand identify their root causes.
- Develop, review, test, and deliver high-quality bug fixes usingC or C++.
- Analyze system logs, core dumps, stack traces, performancemetrics, and storage-related diagnostic information.
- Troubleshoot issues involving distributed systems, storageengines, networking, operating systems, concurrency, andhardware interactions.
- Work closely with customer support and field engineering teamsto collect technical information and provide troubleshootingguidance.
- Collaborate with product development teams on complex defects,architectural improvements, and product reliability.
- Create diagnostic tools, scripts, test cases, and internaldocumentation to improve troubleshooting efficiency.
- Participate in product release validation and ensure fixes aresafely backported to supported product versions.
- Contribute to improving product observability, serviceability,stability, and maintainability.
Required Qualifications
- Bachelor's degree or above in Computer Science, ComputerEngineering, Software Engineering, or a related field.
- Strong programming experience in C or C++.
- Solid understanding of data structures, algorithms,multithreading, memory management, and network programming.
- Practical experience with Linux systems and Linux debuggingtools.
- Good understanding of distributed systems concepts, such asreplication, consistency, consensus, fault tolerance,distributed state management, and failure recovery.
- Experience with one or more storage technologies, such as:
- Distributed storage systems
- Block, file, or object storage
- Storage virtualization
- RAID, snapshots, replication, or data protection
- Local file systems or storage engines
- NVMe, SCSI, iSCSI, Fibre Channel, or NVMe over Fabrics
- Strong troubleshooting and root-cause analysis skills.
- Ability to read and understand a large and mature codebase.
- Ability to communicate technical findings clearly in writtenand spoken English and Chinese.
- Willingness to work directly with customer-facing teams oncomplex production issues.
- Must be a Singapore citizen, permanent resident, or hold avalid Singapore work permit.
Preferred Qualifications
- Experience developing or maintaining distributed storageproducts.
- Experience with Ceph, SPDK, RocksDB, LevelDB, distributeddatabases, or similar infrastructure software.
- Familiarity with Linux kernel storage, device drivers, filesystems, or networking subsystems.
- Experience analyzing core dumps with GDB and diagnosing memorycorruption, deadlocks, race conditions, and performancebottlenecks.
- Experience with observability and diagnostic tools such as perf,eBPF, Valgrind, AddressSanitizer, strace, SystemTap, or similartools.
- Familiarity with Python, Bash, or other scripting languages fortesting and automation.
- Experience in enterprise infrastructure, cloud platforms,hyper-converged infrastructure, or data centre products.
- Previous experience in sustaining engineering, escalationengineering, product support engineering, or site reliabilityengineering.
What We Value
- A strong sense of ownership and responsibility for productquality.
- Patience and persistence when investigating difficult technicalproblems.
- The ability to distinguish symptoms from root causes.
- A practical engineering mindset focused on reliable andmaintainable solutions.
- The ability to work effectively across engineering, QA,support, and customer-facing teams.
- A willingness to understand both the product source code andthe customer's production environment.