Turn this role into an interview — a resume and cover letter built around what this employer wants.
Firmus Technologies in Singapore seeks a Senior AI Infrastructure Engineer, Observability, to define health signals and service-readiness criteria for GPU infrastructure powering customer and internal workloads.
You will publish dashboards, alerts, and runbooks, enable self-service knowledge, and work with operations, engineering, and telemetry teams to keep fleet healthy and prepare for future AI workloads.
Firmus Technologies is a globalleader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combiningcutting-edgetechnology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, andoperate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AIcomputeat scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to join our Engineering and Technology team. You will establish how we measure, validate and communicate the health of GPU infrastructure used for customer and internal workloads. You will define trusted health signals and service-readiness criteria, and turn them into reusable dashboards, alerts, queries, diagnostic checks and operational guidance. Your work will help commissioning, infrastructure and operations teams bring capacity online safely, identify degradation early and recover from failures quickly. You will also make knowledge self-service by publishing clear reference implementations, runbooks and AI-ready operational knowledge that other teams can use and extend.
Reference Implementations
Publish golden dashboards, alerts, PromQL/LogQL queries and health checks that other teams adopt and extend. Your output is the standard and the examples. The alerts must be specific, low-noise, and with a clear next action.
Diagnostics with Real Pass/Fail Criteria
DCGM checks, NCCL and bandwidth tests, stress and burn-in, and validation jobs that confirm a server matches expected performance. Others must be able to run them without you.
Fault Isolation
Separate a bad GPU from a cooling, host, network or power-limit problem using host and BMC/management-interface telemetry together, including when the same signature appears across many servers. A clean management view means nothing if the host is throwing faults.
Detection of the Failures that Don't Crash
Rising correctable ECC counts, NVLink retries, thermal slowdown, XID patterns, wrong results with no error. Keep the knowledge current: what each signal means, what to do next on the machine, and where the operation team must make the final decision.
Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
7+ years in GPU, HPC, AI infrastructure or closely related large-scale systems engineering environments, including ownership of health monitoring or diagnostics used by customers or internal teams.
Experience with production GPU fault diagnosis. You've found genuine faults through XID events, ECC/memory errors, NVLink issues, power or thermal limits and caught at least one before the job died.
Strong Linux and server fundamentals. You can debug on the machine and reason across GPU, CPU, memory, PCIe, power and cooling. You automate in Python or similar.
Hands-on with GPU diagnostics and validation such as DCGM, NCCL/collective tests, stress testing and you've turned them into checks other people run.
Experience analysing infrastructure telemetry, building trusted dashboards and alerts. Savvy with PromQL, LogQL, Grafana or comparable tooling. Able to follow an unexpected signal until there is a cause and can turn that into an alert or check others will trust.
Uses AI tools as a normal part of analysis and build work. Can structure health knowledge and build AI skills so operation team with AI assistants can use it effectively to recover from incidents.
Willing to take part in the incident-response on-call rotation.
Willing to travel overseas occasionally when the role requires it.
Clear and effective written and verbal communication in English.
Full-time
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions