Senior Software Engineer — Infra Agent Systems Remote India Together AI India

Neura Market

Hinoba-an

Remote

INR 3,000,000 - 5,400,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Together AI is building the next generation of autonomous AI infrastructure. The Infra Agent Systems team develops production AI agents that diagnose hardware failures, investigate incidents, and automate workflows across a massive GPU fleet.

The Core Agent Platform will power these agents with knowledge graphs, retrieval, orchestration, and evaluation tooling. You will work on two hard problems: making AI agents trustworthy for production, and building the data and systems that enable

Qualifications

  • 5+ years of experience building production backend systems.
  • Strong systems design skills with ownership from design to production.
  • Deep knowledge in AI agent systems, knowledge graphs, or semantic search.
  • Proficient backend engineering including API design and data modeling.
  • Experience with Kubernetes, GitOps, IaC, and cloud platforms.

Responsibilities

  • Design and build production AI agent systems to diagnose and remediate infrastructure issues.
  • Develop distributed services, orchestration, knowledge graphs, and retrieval systems.
  • Create fleet intelligence by combining telemetry, state, and incidents for better decisions.
  • Integrate with observability, incident management, and source control via APIs.
  • Own services end-to-end: architecture, implementation, testing, deployment, and operations.
  • Improve agent performance and reliability through feedback loops and tooling.
  • Turn learnings into reliable software and automation for production.

Skills

Backend systems
Distributed systems
Systems design
AI agents
Knowledge graphs
Search & retrieval
API design
Kubernetes
GitOps (ArgoCD)
Python
Go
TypeScript

Tools

Prometheus
Grafana
NATS
Kafka
Graph databases
ArgoCD
Terraform
Cloud platforms

Job description

About the Role

Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.

We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.

You’ll work across two areas:

Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.

Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.

We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.

This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.

responsible for delivering the software but also for operating and supporting it in production.

Why this Role

You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.

You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.

Remote based in India

Responsibilities
  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
  • Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
  • Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
  • Turn what agents learn in production into reliable, reviewed software and automation.
Requirements
  • 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
  • Strong systems design skills and experience owning significant systems from design through production.
  • Depth in at least one of the following:
    • AI agent systems, orchestration, tool use, evaluation, or grounding
    • Knowledge graphs or graph data modeling
    • Search, retrieval, ranking, RAG, or semantic search systems
  • Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
  • Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
  • Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Experience in the following is a plus:

  • GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
  • Graph databases
  • Event-driven systems and messaging platforms such as NATS or Kafka
  • Observability platforms such as Prometheus and Grafana
  • Building evaluation frameworks or improving the quality and reliability of LLM-powered systems
About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Software Engineer, AI Agents
Lead Software Engineer, AI Agents

V2 Solutions • Hinoba-an

Hybrid
PHP 928,000 - 1,724,000
Health insurance
Flexible vacation
Global team
Senior AI Infra Engineer — Build Self-Improving Agents
Senior AI Infra Engineer — Build Self-Improving Agents

Neura Market • Hinoba-an

Remote
INR 3,000,000 - 5,400,000
Senior Data Operations Engineer
Senior Data Operations Engineer

Field AI • Boston

On-site
PHP 7,421,000 - 11,132,000
AI Engineer
AI Engineer

Dry Ground • Philippines

On-site
PHP 2,995,000 - 4,794,000
Competitive salary and performance-based incentives
Flexible work environment
Collaborative innovation-driven culture
Senior Software Engineer AI Infrastructure
Senior Software Engineer AI Infrastructure

Alignerr • Manila

On-site
PHP 3,438,000 - 5,156,000
Senior AI Engineer
Senior AI Engineer

Adventure Consultancy Solutions Philippines Inc. • Metro Manila

On-site
PHP 2,200,000 - 4,000,000
Full Stack AI Engineer
Full Stack AI Engineer

Foss United • Hinoba-an

On-site
INR 4,000,000 - 7,000,000
Health insurance
ESOPs
Software Engineer – Backend & Platform
Software Engineer – Backend & Platform

Soket AI Labs • Hinoba-an

On-site
PHP 900,000 - 1,300,000
Competitive compensation
Equity participation opportunities
Flexible work arrangements
+2
AI Agent Engineer
AI Agent Engineer

Teoh Capital • Philippines

On-site
PHP 1,200,000 - 2,200,000
Lead Architect
Lead Architect

V2 Solutions • Hinoba-an

Hybrid
PHP 1,194,000 - 2,122,000