A complete application in a minute — tailored resume and cover letter, ready to send.
Convo is seeking a Prompt & Evaluation Engineer with 5+ years of experience in applied NLP/LLM engineering to design prompt chains, agent behaviors, and evaluation frameworks. The role centers on translating business requirements into reliable AI solutions using Python-based evaluation, structured outputs, regression testing and thorough failure analysis, collaborating with domain experts.
You will work with cross-functional teams to optimize prompts and maintain robust evaluation pipelines
Convo is looking for a Prompt & Evaluation Engineer with 5+ years of experience in applied NLP/LLM engineering to design and optimize prompt chains, agent behaviors, and evaluation frameworks. The ideal candidate should have strong expertise in prompt engineering, LLM evaluation, Python, structured outputs, regression testing, and failure analysis, with the ability to translate business requirements and commercial use cases into reliable, measurable AI solutions.
Engineer prompt chains, agent constitutions and evaluation suites that convert governed CPG semantics into reliable commercial-agent behavior.
Works with the CPG Domain Expert, Agent Runtime, Agentic Governance, application teams and AI Test & Dataset Engineering.
Great compensation package, medical benefit for you and your family, free lunch, annual performance-tied increments & performance recognition awards and a great lean and agile work culture!
Convo endorses a culture of diversity in all aspects and aims to build a diverse team of amazing individuals!