Turn this role into an interview — a resume and cover letter built around what this employer wants.
Infrrd, Bengaluru-based Enterprise AI company, seeks an engineer to build automated systems that measure, diagnose, and improve document extraction accuracy at scale. This role replaces brute-force prompt iterations with agentic evaluation pipelines, automated feedback loops, and internal tooling to accelerate the team.
You will design and productionize LLM-based extraction and classification pipelines, establish robust A/B testing, live dashboards, and cost-optimized deployments.
Hello there! Infrrd here — it’s pronounced In-fur-d.
We’re an Enterprise AI company that uses AI and Machine Learning to help global organisations automate data extraction from complex documents — invoices, contracts, insurance claims, and more. Our customers are some of the world’s leading enterprises in mortgage, insurance, and manufacturing, and we’ve been profitable and independent since 2016.
To build the automated systems that measure, diagnose, and improve document extraction and classification accuracy at scale. This role eliminates the manual bottleneck in the accuracy improvement cycle — replacing brute-force prompt iteration with agentic evaluation pipelines, automated feedback loops, and intelligent internal tooling. The engineer in this role makes the entire team faster without proportionally increasing headcount, and enables systematic accuracy improvement as a repeatable engineering capability rather than an ad-hoc effort.
BE / MTech in Computer Science, AI/ML, Computational Data Science (CDS), Computer Science & Automation (CSA), or related discipline.
8-10 years total; minimum 4-6 years building production LLM or AI systems; minimum 4-6 years in evaluation, quality measurement, or accuracy improvement work.
Python, FastAPI / Flask, MongoDB, Git, GitHub Actions / Jenkins, LLM APIs (OpenAI / Anthropic / Gemini or equivalent), LangChain / LlamaIndex, Pandas / Numpy, Pytest, Docker
NLP concepts, LLM prompt engineering patterns, REST APIs, RAG pipelines, vector databases, JSON data structures
Agentic workflow design and orchestration, LLM evaluation metrics (F1 / Precision / Recall, per-class analysis, confusion matrices), production Python systems (error handling, retries, logging, monitoring), NoSQL aggregations, systematic A/B testing for model changes, prompt optimization methodology
By submitting your application, you agree that your personal information and resume may be collected, processed, and stored by us for recruitment purposes, including consideration for future roles.