A complete application in a minute — tailored resume and cover letter, ready to send.
FIRMUS METAL INTERNATIONAL PTE. LTD. is seeking a Senior AI Engineer (Agents & Applications) to design, build, and operate agentic systems that coordinate, optimize, and automate model-to-grid workflows in a production setting.
You will implement reference architectures, build orchestration for single- and multi-agent systems, and develop retrieval-augmented pipelines with observability and safety controls across the design-build-operate lifecycle.
Role SummaryThe Senior AI Engineer (Agents & Applications) will design, build, and operate production-grade agentic systems that coordinate, optimize, and automate decision-making across the design-build-operate lifecycle of AI factories. The role is a core contributor to the AI & Applications team's Model-to-Grid product, connecting models, inference endpoints, benchmark intelligence, validated workload recipes, job-scheduler decisions, infrastructure telemetry, AI-factory operations, and grid-related constraints into safe, explainable, and measurable workflows.The role will build more than conversational co-pilots. It will create agentic applications that ingest and reason over time-series telemetry, logs, traces, events, scheduler state, benchmark results, configuration data, operational documentation, incident records, and multimodal sources where appropriate. These applications will help engineers, operators, and customers move from observation to diagnosis, recommendation, planning, simulation, controlled execution, verification, and continuous improvement.The engineer will define and implement the underlying agent architecture and engineering framework: orchestration, state and memory management, retrieval, tool use, specialized sub-agents, evaluation, safety controls, human approvals, observability, and deployment. The role will use fit-for-purpose self-hosted and external model endpoints, with close integration to the team's inference platform.
Design, build, and operate agentic applications supporting AI-factory planning, commissioning, validation, workload onboarding, benchmark analysis, model and recipe optimisation, scheduling, operations, maintenance, incident response, and continuous improvement.
Define reference architectures for single-agent, multi-agent, workflow-based, eventdriven, and human-in-the-loop agentic systems.
Build orchestration workflows using appropriate agent frameworks and libraries, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or equivalent custom-built frameworks.
Select the appropriate architecture for each use case rather than applying multi-agent patterns by default:Deterministic workflow and state-machine architectures for repeatable, high-confidence operational processes.Planner-executor architectures for decomposing complex investigation, planning, and remediation tasks.Supervisor-worker or manager-worker architectures for coordinating specialist domain agents.Router architectures for selecting the right model, tool, knowledge source, workflow, or specialist agent.Reflection, critic, verifier, or judge patterns for quality assurance, validation, and safety checks.Event-driven architectures for responding to telemetry anomalies, workload failures, scheduler events, benchmark regressions, and operational alerts.Human-in-the-loop architectures for high-impact recommendations, privileged actions, or changes to production environments.
Build specialist agents for relevant Model-to-Grid and AI-factory domains, such as:Benchmark-analysis and performance-diagnosis agents.Workload recipe and runtime-configuration recommendation agents.Inference-endpoint selection, capacity, and optimization agents.Kubernetes and job-scheduler diagnostic agents.GPU-topology, network, RDMA, storage, and utilization-analysis agents.AI-factory health, capacity, maintenance, and operational-triage agents.Documentation, knowledge, incident-review, and runbook-execution assistants.Thermal domain specific monitoring and optimization agents.Power domain specific monitoring and optimization agents.Grid-integration specific monitoring and optimization agents.
Develop the intelligent coordination layer for Model-to-Grid, enabling agents to reason across model characteristics, inference and training configuration, validated recipes, GPU resources, topology, scheduling policies, network and storage performance, capacity, power, thermal conditions, health signals, and operational constraints.
Build Retrieval-Augmented Generation (RAG) pipelines using a combination of vector retrieval, hybrid search, metadata filtering, reranking, structured-data queries, graphbased retrieval where valuable, source attribution, and permission-aware access controls.
Use tools such as pgvector, OpenSearch, Elasticsearch, Milvus, Weaviate, Pinecone, Qdrant, Neo4j, or equivalent data and retrieval platforms as appropriate to the product architecture and deployment environment.
Design knowledge-ingestion pipelines for documentation, runbooks, ticketing systems, configuration repositories, benchmark reports, experiment records, cluster state, telemetry catalogues, incident reports, and approved internal knowledge sources.
Build data and context pipelines that combine unstructured knowledge with structured operational data, including metrics, logs, traces, events, time-series databases, scheduler queues, job states, resour