Turn this role into an interview — a resume and cover letter built around what this employer wants.
Lexsi Labs is hiring engineers to design and implement the core agent loop, evaluation system, and shared tool protocols for autonomous AI systems. You will build scalable, sandboxed execution environments and robust observability tooling, working across research and engineering to advance evaluation methodologies and interpretability.
We value engineers who ship production services, understand concurrency at scale, and contribute to open-source projects.
Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission is to build safety aware autonomous systems in the extreme near term. Our research work spans areas like AI alignment methodologies, interpretability-led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences including ICLR, ICML, WWW, IJCNN, MICCAI and EurIPS. Our labs are located in India (Mumbai and remote), Paris, and London.
We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior.
Our current sprint on building autonomous systems for complex problems, across software engineering, data science, and AI research, involves building the harness, execution substrate and evaluation system, and each is designed to run inside a customer's environment rather than ours.
This role sits underneath all three agents. The coding agent, the data science agent and the AI engineering agent look different from the outside, but they are the same system underneath: a loop that plans, acts, observes, recovers, and knows when to stop. You will build that shared layer, and the evaluation system that tells us whether any change to it made things better. Both halves matter equally. A harness we cannot measure is a harness we cannot improve.
Harness
Evals
You will work across all three agent teams and closely with our research team on evaluation design, post-training and interpretability of agent behavior.
This is a systems and infrastructure role. Most of the difficulty here is concurrency, state and measurement, not prompting.
We are hiring several engineers for this team at a range of experience levels, including engineers early in their careers who have strong fundamentals and want to work on agents from the infrastructure side.
We move quickly and expect candidates to do the same. We value substance over polish and execution over rhetoric.