Agent Engineer - Evals and Harness

Story Terrace Inc.

Mumbai

Hybrid

INR 1,800,000 - 2,400,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Lexsi Labs is hiring engineers to design and implement the core agent loop, evaluation system, and shared tool protocols for autonomous AI systems. You will build scalable, sandboxed execution environments and robust observability tooling, working across research and engineering to advance evaluation methodologies and interpretability.

We value engineers who ship production services, understand concurrency at scale, and contribute to open-source projects.

Qualifications

  • Strong software engineering fundamentals and advanced Python.
  • Concurrency and distributed systems; async execution and retries.
  • Containers and sandboxing with Docker and OCI internals.
  • Test and CI infrastructure; large suites and flaky tests handling.
  • Measurement literacy: variance, benchmarking, SIP and significance.

Responsibilities

  • Develop the core agent loop and execution model including orchestration, concurrency, sub-agent coordination, cancellation, timeouts, checkpointing and resume.
  • Manage context and state for long-running tasks and reasoning over past work.
  • Define and maintain the shared tool protocol, schemas, versioning and internal libraries used by all agents.
  • Build the sandboxed execution substrate that is reproducible and resource-bounded across tasks.
  • Establish trace schema, debugging surface and provide robust observability tooling.
  • Collaborate across the three agent teams with research for evaluation design and interpretability.

Skills

Python
Concurrency
Distributed systems
Testing / CI
Observability
Debugging
Open source

Tools

Docker
OCI internals

Job description

Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission is to build safety aware autonomous systems in the extreme near term. Our research work spans areas like AI alignment methodologies, interpretability-led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences including ICLR, ICML, WWW, IJCNN, MICCAI and EurIPS. Our labs are located in India (Mumbai and remote), Paris, and London.

We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior.

The Role

Our current sprint on building autonomous systems for complex problems, across software engineering, data science, and AI research, involves building the harness, execution substrate and evaluation system, and each is designed to run inside a customer's environment rather than ours.

This role sits underneath all three agents. The coding agent, the data science agent and the AI engineering agent look different from the outside, but they are the same system underneath: a loop that plans, acts, observes, recovers, and knows when to stop. You will build that shared layer, and the evaluation system that tells us whether any change to it made things better. Both halves matter equally. A harness we cannot measure is a harness we cannot improve.

What you'll work on:

Harness

  • The core agent loop and its execution model, covering orchestration, concurrency, sub-agent coordination, cancellation, timeouts, checkpointing and resume.
  • Context and state management for long-running tasks, including retention, compaction, and how an agent reasons over what it has already done.
  • The shared tool protocol, its schemas, versioning, and the internal libraries every agent team builds against.
  • The execution substrate. Sandboxed environments that are reproducible, snapshot-able and resource-bounded, and that behave identically for a repository task, a training run and a data analysis.
  • Failure semantics. Retry policy, partial failure, idempotency, and the distinction between a recoverable error and a task that should stop and hand back.
  • The trace schema every agent emits, which serves at once as our debugging surface, our training signal, and the audit record our customers keep.

Evals

  • Task suites for each agent type, built from real work rather than synthetic benchmarks, and the infrastructure to run them at scale in parallel.
  • Verifiers and scoring. Deterministic checks where the task allows it, model-graded rubrics where it does not, and calibration of the graders themselves.
  • Regression gates that run on every harness change, with cost and latency accounted alongside quality.
  • The statistics to say a result is real. Seeds, variance, pass@k, confidence, and knowing how many runs are needed before anyone claims an improvement.
  • Turning observed production failures into permanent test cases.
  • Tooling for how we work. Trace inspection, replay, and diffing one run against another.

You will work across all three agent teams and closely with our research team on evaluation design, post-training and interpretability of agent behavior.

What We Are Looking For

This is a systems and infrastructure role. Most of the difficulty here is concurrency, state and measurement, not prompting.

  • Strong software engineering fundamentals and advanced Python. You have shipped and operated production services, not only written them.
  • Concurrency and distributed systems. Async execution, worker pools and queues, retries and idempotency, backpressure, cancellation, and reasoning clearly about what happens when a step fails halfway through.
  • Containers and sandboxing. Docker and OCI internals, resource isolation, reproducible environments, and an understanding of why a job that passes locally fails in a sandbox.
  • Test and CI infrastructure. You have built test harnesses, run large suites in parallel, and dealt with flakiness as an engineering problem rather than an annoyance.
  • Measurement literacy. Comfort with variance, sampling and significance, and healthy skepticism toward benchmark results including your own.
  • Observability instincts. Distributed tracing, structured logging, and building the inspection tooling that makes non-deterministic systems debuggable.
  • You debug systematically. You find out what actually happened rather than adjusting things until the symptom disappears.
  • You are comfortable when problems are loosely specified and ownership is assumed rather than assigned.
  • Experience with evaluation or benchmarking work for AI systems is a strong plus.
  • Experience with agentic systems, LLM inference and serving, or developer tooling is a plus.
  • Open source contributions we can read are a plus.

We are hiring several engineers for this team at a range of experience levels, including engineers early in their careers who have strong fundamentals and want to work on agents from the infrastructure side.

We move quickly and expect candidates to do the same. We value substance over polish and execution over rhetoric.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineering- Lead (AI Coding Agent, Platform Engineering, Loop Engineering, AWS)
AI Engineering- Lead (AI Coding Agent, Platform Engineering, Loop Engineering, AWS)

FICO • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Flexible work options
Benefits program
Parental leave
+1
Agent Infrastructure Engineer — Core Harness (Superagent)
Agent Infrastructure Engineer — Core Harness (Superagent)

ImagineArt • India

On-site
INR 4,000,000 - 7,000,000
AI Software Engineer, Agent Harness
AI Software Engineer, Agent Harness

AlleyCorp • India

Hybrid
INR 4,000,000 - 7,000,000
Sr AI Agent Engineer
Sr AI Agent Engineer

Story Terrace Inc. • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Agent Engineer | AI
Agent Engineer | AI

Stealth Video Generation Startup • India

On-site
INR 2,000,000 - 3,000,000
Competitive Pay
Work-Life Balance
Cutting-Edge AI Work
+1
Lead AI Engineer - Agentic Engineering
Lead AI Engineer - Agentic Engineering

Blend360 India • Hyderabad

Hybrid
INR 3,000,000 - 5,000,000
Engineer I - AI
Engineer I - AI

NewSpace Research and Technologies • Bengaluru

On-site
INR 900,000 - 1,500,000
Forward Deployed Engineer
Forward Deployed Engineer

Insight Global • Hyderabad

On-site
INR 3,500,000 - 6,500,000
Senior AI Engineer
Senior AI Engineer

Arcana • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Software Engineer (Junior/Mid) — Agentic AI (Quant Research Platform) | Chennai (on-site)
Software Engineer (Junior/Mid) — Agentic AI (Quant Research Platform) | Chennai (on-site)

Jnaara • Chennai District

On-site
INR 1,500,000 - 3,000,000
Relocation support available
Equity-based grant
Clear path up: comp/scope re-benchmark