Senior Software Engineer — Infra Agent Systems

Talanto

Amsterdam

Hybrid

EUR 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Talanto is seeking a Senior Software Engineer for Infra Agent Systems in Amsterdam. You will build production AI agents that diagnose hardware failures, investigate incidents, and automate remediation, while partnering on the platform and knowledge graphs that power these agents.

You will work across Infrastructure Agent Systems and Core Agent Platform, crafting scalable backend systems, API integrations, and tooling with a focus on reliability and continuous improvement.

Qualifications

  • 5+ years of experience building production backend systems or distributed infrastructure.
  • Strong systems design skills with ownership of large-scale projects.
  • Experience with Kubernetes and cloud platforms; familiarity with IaC.
  • Proficiency in Go, TypeScript, Python, or Rust.

Responsibilities

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues.
  • Develop knowledge graphs, retrieval systems, and orchestration for AI agents.
  • Build and operate the platform powering agents, including evaluation and tooling.
  • Own services end to end: architecture, implementation, testing, deployment, observability.
  • Improve agent performance via retrieval enhancements and production feedback loops.
  • Integrate with observability, incident management, ticketing, and internal systems via APIs.

Skills

Backend engineering
Kubernetes
Go
TypeScript
Python
Rust

Tools

ArgoCD
NATS
Kafka
Prometheus
Grafana

Job description

Senior Software Engineer — Infra Agent Systems

Important: if an employer asks you to log into their system via iCloud or Google, send a code, an SMS or Telegram password, run some code, or install software — refuse. These are signs of fraud.

About the Role

runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.

We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.

You’ll work across two areas:

Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.

Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.

We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.

This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.

responsible for delivering the software but also for operating and supporting it in production.

Why this Role

You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.

You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.

Hybrid in Amsterdam

Responsibilities

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
  • Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
  • Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
  • Turn what agents learn in production into reliable, reviewed software and automation.

Requirements

  • 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
  • Strong systems design skills and experience owning significant systems from design through production.
  • Depth in at least one of the following:
  • AI agent systems, orchestration, tool use, evaluation, or grounding
  • Knowledge graphs or graph data modeling
  • Search, retrieval, ranking, RAG, or semantic search systems
  • Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
  • Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
  • Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Experience in the following is a plus:

  • GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
  • Event-driven systems and messaging platforms such as NATS or Kafka
  • Observability platforms such as Prometheus and Grafana
  • Building evaluation frameworks or improving the quality and reliability of LLM-powered systems

the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together AI • Amsterdam

Hybrid
EUR 95,000 - 150,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Togetherai • Amsterdam

Hybrid
EUR 120,000 - 180,000
Senior Infra Agent Systems Engineer: AI Agents & Platform
Senior Infra Agent Systems Engineer: AI Agents & Platform

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Amsterdam

On-site
EUR 90,000 - 130,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Amsterdam

On-site
EUR 90,000 - 130,000
Senior Software Engineer Together Cloud Infrastructure
Senior Software Engineer Together Cloud Infrastructure

Together AI • Amsterdam

On-site
EUR 70,000 - 90,000
Senior AI Agent Infra Engineer – Production Systems
Senior AI Agent Infra Engineer – Production Systems

Together AI • Amsterdam

Hybrid
EUR 95,000 - 150,000
Principal Full Stack Engineer, AI Platform & Agents
Principal Full Stack Engineer, AI Platform & Agents

Wolters Kluwer • Alphen aan den Rijn

Hybrid
EUR 70,000 - 100,000
Senior ML Engineer (Token Factory)
Senior ML Engineer (Token Factory)

United States Digital Space LLC • Amsterdam

Hybrid
EUR 100,000 - 180,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3