Role Overview
We are looking for a highly skilled Senior Software Engineer / Lead Engineer to build and scale enterprise-grade Agentic AI platforms and GenAI solutions. The ideal candidate should have hands‑on experience designing, deploying, and operating production AI systems, with strong expertise in Python, Agentic AI, RAG, Distributed Systems, Cloud‑Native Engineering, and AI Governance.
Key Responsibilities
Agentic AI & Platform Engineering
- Design and develop production-grade Agentic AI workflows and multi-agent systems.
- Build reliable agent runtimes with planning, orchestration, tool execution, memory, approvals, and recovery mechanisms.
- Implement structured outputs, tool/function calling, workflow orchestration, and deterministic execution patterns.
- Define engineering best practices for scalable and governed AI systems.
RAG, Memory & Context Management
- Design and implement enterprise-scale RAG solutions.
- Build memory and context-management capabilities across AI workflows.
- Optimize retrieval quality using embeddings, hybrid search, reranking, and citation grounding.
AI Governance & Reliability
- Implement controls for prompt injection, hallucinations, PII protection, RBAC, tenant isolation, and policy enforcement.
- Establish monitoring, evaluation, observability, and auditability for production AI systems.
- Drive performance tuning, reliability, testing, incident management, and cost optimization.
Cloud & Platform Engineering
- Build and operate scalable AI services using Python, Docker, Kubernetes, and CI/CD.
- Develop enterprise integrations and MCP-based tool interfaces.
- Implement observability using logs, metrics, monitoring, and distributed tracing.
Mandatory Skills
- 6-8 years of software engineering experience with strong backend development expertise.
- 2+ years of hands‑on experience building and operating GenAI / Agentic AI applications in production.
- Strong Python development experience using FastAPI, Pydantic, AsyncIO, APIs, and microservices architecture.
- Experience with OpenAI, Azure OpenAI, Anthropic Claude, or Google Gemini.
- Strong experience in:
- Agentic AI Architectures
- Multi-Agent Systems
- Tool/Function Calling
- Workflow Orchestration
- MCP (Model Context Protocol)
- RAG Frameworks
- Hands‑on experience with:
- Embeddings
- Vector Databases
- Hybrid Retrieval & Reranking
- Memory & Context Management
- Strong understanding of:
- Distributed Systems
- Concurrency
- Idempotency
- Fault Tolerance
- Failure Recovery
- Experience with:
- PostgreSQL
- Redis and/or MongoDB
- Docker & Kubernetes
- CI/CD Pipelines (GitHub Actions, GitLab CI, etc.)
- Strong knowledge of:
- API Design
- Testing Automation
- Code Quality
- Observability
- Production Operations
Preferred Skills
- Experience with LangGraph, CrewAI, AutoGen, Semantic Kernel, OpenAI Agents SDK, DSPy, or Google ADK.
- Exposure to LLMOps and AI Observability tools such as Langfuse, TrueFoundry, Arize Phoenix, Ragas, DeepEval, or Promptfoo.
- Experience with AI Governance and Security frameworks such as Presidio, Guardrails AI, NeMo Guardrails, OPA, or Cedar.
- Experience with AWS, Azure, or GCP.
- Exposure to BFSI, Insurance, Healthcare, or other regulated industries.