Stand out for this role — generate a tailored resume and cover letter in about a minute.
Delta Exchange in Bengaluru, India, is seeking a Senior QA Engineer - AI to own the quality lifecycle of our AI products used by live traders. You will design evaluation datasets, set scoring mechanisms, and lead the release gate to prevent regressions from reaching customers.
You will build end-to-end evaluation pipelines, triage failures, and define quality metrics for AI surfaces, working with a highly autonomous team that develops an API Copilot, a customer support bot, and an MCP server.
At Delta, we are reimagining and rebuilding the financial system. Join our team to make a positive impact on the future of finance.
Mission Driven: Re-imagine and rebuild the future of finance.
Most innovative cryptocurrency derivatives exchange. With a daily traded volume of ~$ 10 billion, and increasing. Delta is bigger than all the Indian crypto exchanges combined.
Offer the widest range of derivative products and have been serving traders all over the globe since 2018 and growing fast.
The founding team is comprised of IIT and ISB graduates. Business co-founders have previously worked with Citibank, UBS and GIC; and our tech co-founder is a serial entrepreneur who previously co-founded TinyOwl and Housing.com.
Funded by top crypto funds (Sino Global Capital, CoinFund, Gumi Cryptos) and crypto projects (Aave and Kyber Network).
Delta operates three AI products used by live traders: a customer support chatbot, an API Copilot that generates and executes trading scripts, and an MCP server. This role owns the quality of those products and is responsible for measuring it objectively.
Testing probabilistic systems differs fundamentally from testing deterministic ones. The same input produces different output on every run, an incorrect answer can be entirely fluent, and a prompt change in one flow can degrade another without any visible signal. The core of this role is establishing what correct means for these systems, building the datasets and scoring that measure it, and operating the release gate that prevents a regression from reaching customers.
This is an emerging discipline with no established playbook. We are defining the methodology as we build it, and the role carries a high degree of autonomy and ownership.
Most QA roles that mention AI mean using AI to do testing faster: generating test cases from a PRD, healing flaky selectors, exploring an app with an agent. Those are useful and we do them.
This role is the other thing. The system under test is itself an AI, and the hard problem is deciding whether its output is correct when the same question produces a different answer every time and a wrong answer reads as convincingly as a right one. That means building evaluation datasets, defining what correct looks like, scoring against it, and defending a number that decides whether a release ships.
If you have spent time on the second problem, this role is built for you.
Non-negotiable
Also required
Bonus