Senior AI Infrastructure & Observability Engineer

Firmus

Singapore

On-site

SGD 180,000 - 260,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to define health signals and service-readiness gates for GPU infrastructure. You will publish reusable dashboards, alerts, and checks for fleet health and incident guidance.

You will collaborate with commissioning, operations and telemetry owners to produce production‑grade signals, runbooks, and AI‑ready knowledge that scales across teams. Strong English communication is required.

Qualifications

  • Bachelor’s degree in computer science or a related technical field.
  • 7+ years in GPU, HPC, AI infrastructure or closely related large‑scale systems engineering environments.
  • Experience with production GPU fault diagnosis. You’ve found genuine faults through XID events, ECC/memory errors, NVLink issues, power or thermal limits.
  • Strong Linux and server fundamentals. You can debug on the machine and reason across GPU, CPU, memory, PCIe, power and cooling. You automate in Python or similar.
  • Hands‑on with GPU diagnostics and validation such as DCGM, NCCL/collective tests, stress testing and you’ve turned them into checks other people run.
  • Uses AI tools as a normal part of analysis and build work. Can structure health knowledge to enable AI assistants for incident recovery.

Responsibilities

  • The Health Standard: Define GPU and host health criteria and service-readiness gates for repair and capacity workflows.
  • Reference Implementations: Publish golden dashboards, alerts, PromQL/LogQL queries and health checks that others adopt.
  • Diagnostics with Real Pass/Fail Criteria: DCGM checks, NCCL and bandwidth tests, stress and burn‑in, validation jobs.
  • Fault Isolation: Separate a bad GPU from cooling, host, network or power‑limit problems using telemetry.
  • Detection of the Failures that Don’t Crash: Monitor rising ECC counts, NVLink retries, thermal slowdown and latency patterns.

Skills

GPU diagnostics
DCGM checks
NCCL tests
Linux fundamentals
Python automation
Telemetry analysis
PromQL/LogQL
AI tooling
Incident on-call
Overseas travel
English communication

Education

Bachelor’s degree in computer science or a related technical field

Job description

Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to define health signals and service-readiness gates for GPU infrastructure. You will publish reusable dashboards, alerts, and checks for fleet health and incident guidance.

You will collaborate with commissioning, operations and telemetry owners to produce production‑grade signals, runbooks, and AI‑ready knowledge that scales across teams. Strong English communication is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer, Observability
Senior AI Infra Engineer, Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infrastructure Engineer, Observability
Senior AI Infrastructure Engineer, Observability

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
Senior Platform Engineer - AI Observability & Automation
Senior Platform Engineer - AI Observability & Automation

Firmus • Singapore

On-site
SGD 140,000 - 200,000
Senior AI Infrastructure Engineer, Observability
Senior AI Infrastructure Engineer, Observability

Firmus Technologies • Singapore

On-site
SGD 180,000 - 240,000
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Sustainable Metal Cloud • Singapore

On-site
SGD 180,000 - 240,000
Senior Platform Engineer – AI/ML Observability
Senior Platform Engineer – AI/ML Observability

Firmus Technologies • Singapore

On-site
SGD 120,000 - 210,000
Senior Software Engineer, Platform
Senior Software Engineer, Platform

Firmus • Singapore

On-site
SGD 140,000 - 200,000
Engineering Manager, AI Platform & Observability
Engineering Manager, AI Platform & Observability

Firmus • Singapore

On-site
SGD 180,000 - 260,000
Site Reliability Engineer — AI HPC Infrastructure
Site Reliability Engineer — AI HPC Infrastructure

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000