Senior Full Stack AI Engineer — Agentic
Location and working model: Pune, India — hybrid. Minimum three hours of daily overlap with US Central time.
Experience: 5+ years in engineering, including at least 1 year building LLM-backed products that ran in production.
Role overview
We is building enterprise-grade software solutions that combine modern web engineering, cloud platforms and production-grade AI capabilities. This is a senior hands-on role anchoring agentic development: you will design and build LLM-powered agents, establish the reusable engineering patterns other engineers build against, and set the standard for how non-deterministic systems are evaluated. Mentoring and design contribution beyond your own features are core to the role, not occasional extras.
Key responsibilities
- Anchor the hands-on build of LLM-powered agents — reasoning loops, planning, tool and function calling, retries, fallbacks and guardrails.
- Establish reusable agent engineering patterns and shared libraries that other engineers build against.
- Define and implement evaluation for agent behaviour: how quality is proven before release and how regression is detected afterwards.
- Build retrieval and structured data access over large corpora to support agent reasoning. Instrument agents for observability — tracing, token accounting and cost attribution.
- Contribute to design reviews across the wider engineering group and review other engineers' agent work.
- Mentor engineers on agentic patterns and on responsible AI-assisted development, defining review, testing and traceability standards for AI-generated changes so unverified or architecture-breaking code does not reach the codebase.
- Integrate agent capability with surrounding platform services, including administration, access control and workflow.
Required Skills And Experience
- Deep production Python — async patterns, queues, structured outputs, and the discipline to make non-deterministic systems testable.
- Hands-on experience with an agent framework (for example LangGraph, LlamaIndex, CrewAI or AutoGen, or an in-house equivalent) and with function and tool calling.
- Prompt design paired with evaluation: you can explain how you measured whether an agent worked, not only how you built it.
- Retrieval-augmented generation or structured retrieval over a large corpus.
- Observability for LLM systems — tracing, token accounting and regression suites for agent behaviour.
- Demonstrated mentoring: pairing, design review and visibly raising the standard of the engineers around you.
- Disciplined use of AI coding assistants within defined engineering and security guardrails: you independently validate generated code, verify it through unit, integration and regression tests, check maintainability, architecture alignment, security, privacy, licensing and performance, and never treat generated output as production-ready without human review. You can explain and defend any code you submit.
- Comfortable with ambiguity — agent behaviour is established empirically rather than specified up front.
Preferred Qualifications
- Full stack range — able to build the interfaces that expose agent capability.
- QA or test-automation domain knowledge.
- Experience with browser automation or DOM analysis libraries. Cost and performance optimization for LLM workloads.