Senior AI Infra Engineer, Observability

Firmus Technologies

Singapore

On-site

SGD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Firmus Technologies in Singapore seeks a Senior AI Infrastructure Engineer, Observability, to define health signals and service-readiness criteria for GPU infrastructure powering customer and internal workloads.

You will publish dashboards, alerts, and runbooks, enable self-service knowledge, and work with operations, engineering, and telemetry teams to keep fleet healthy and prepare for future AI workloads.

Qualifications

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • 7+ years in GPU, HPC, AI infrastructure or related large-scale systems engineering environments with health monitoring ownership.
  • Experience with production GPU fault diagnosis involving XID, ECC/memory errors, NVLink issues, power or thermal limits.
  • Strong Linux and server fundamentals; ability to debug on the machine and reason across GPU/CPU/memory/PCIe/power/cooling.
  • Hands-on with GPU diagnostics and validation (DCGM, NCCL tests, stress tests) and converting them to checks for others.
  • Experience analysing infrastructure telemetry and building trusted dashboards and alerts (PromQL, LogQL, Grafana).
  • Uses AI tools in analysis and has prepared operational knowledge for AI assistants.

Responsibilities

  • Define GPU and host health criteria and service-readiness gates for repair and capacity workflows.
  • Publish golden dashboards, alerts, queries and health checks for adoption by other teams.
  • Run DCGM checks, NCCL tests, stress/burn-in and validation jobs to confirm server performance.
  • Isolate faults by using telemetry to distinguish hardware, cooling, network or power-limit issues.
  • Maintain knowledge of signals, what they mean, and actions for incident recovery.

Skills

Health monitoring
GPU diagnostics
Linux fundamentals
Python automation
PromQL/LogQL
Telemetry dashboards

Education

Bachelor's degree in CS or related field

Tools

DCGM
NCCL/Tests
PromQL/LogQL
Grafana

Job description

Firmus Technologies in Singapore seeks a Senior AI Infrastructure Engineer, Observability, to define health signals and service-readiness criteria for GPU infrastructure powering customer and internal workloads.

You will publish dashboards, alerts, and runbooks, enable self-service knowledge, and work with operations, engineering, and telemetry teams to keep fleet healthy and prepare for future AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure & Observability Engineer
Senior AI Infrastructure & Observability Engineer

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infrastructure Engineer, Observability
Senior AI Infrastructure Engineer, Observability

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Senior AI Infrastructure Engineer, Observability
Senior AI Infrastructure Engineer, Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
Staff AI Infrastructure Systems Engineer
Staff AI Infrastructure Systems Engineer

Cloudera • Singapore

On-site
SGD 180,000 - 260,000
Generous PTO Policy
Flexible WFH Policy
Mental & Physical Wellness programs
+2
Senior Platform Engineer - AI Observability & Automation
Senior Platform Engineer - AI Observability & Automation

Firmus • Singapore

On-site
SGD 140,000 - 200,000
Senior Platform Engineer – AI/ML Observability
Senior Platform Engineer – AI/ML Observability

Firmus Technologies • Singapore

On-site
SGD 120,000 - 210,000
Senior InfraOps Engineer: GPU Cloud | Equity & PTO
Senior InfraOps Engineer: GPU Cloud | Equity & PTO

Lightning AI • Singapore

On-site
SGD 165,000 - 205,000
Health coverage
Retirement contributions
Unlimited PTO
+4
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Sustainable Metal Cloud • Singapore

On-site
SGD 180,000 - 240,000