Role: Lead Agentic AI Engineer
Experience: 5–10 years
Employment Type: Full-time
Role Summary
We are looking for a Lead Agentic AI Engineer who can own the end-to-end design and delivery of complex, production-grade agentic systems. You will be the go-to technical expert and the engine room of our most demanding AI initiatives — turning ambiguous client challenges into scalable, functional platforms. You will drive technical solutioning for client engagements, architect multi‑agent pipelines, and bridge AI engineering with business outcomes while elevating the capability of the team around you.
Responsibilities
Agentic Architecture & Engineering
- Architect multi-agent systems — orchestrator/sub-agent patterns, state machines, tool registries, using Microsoft Agent Framework, LangGraph, CrewAI, AutoGen, or Semantic Kernel.
- Design and optimize retrieval pipelines: hybrid search, re-ranking, query expansion, multi-hop reasoning, and knowledge graphs.
- Apply Quantization, PEFT/LoRA fine‑tuning, and prompt optimization techniques to adapt foundation models for client-specific tasks.
- Design and enforce comprehensive guardrail frameworks — output validation, factual grounding checks, prompt injection defenses, content filtering, and hallucination‑mitigation strategies (chain-of-verification, retrieval grounding, self-consistency) — for enterprise-grade deployments.
MLOps & Production Readiness
- Productionize AI services on AWS / Azure using Docker, Kubernetes, and CI/CD pipelines (GitHub Actions / Azure DevOps).
- Build comprehensive monitoring for LLM systems — tracking accuracy, hallucinations, latency, cost, and drift using LangSmith, Arize, or Phoenix.
- Define and implement LLM evaluation suites using RAGAS, G-Eval, TruLens, or custom metrics aligned to client KPIs.
- Drive down inference costs through token budgeting, prompt compression, KV-cache management, model routing, streaming strategies, and intelligent batching.
- Own and evolve CI/CD pipelines for ML systems, enforcing automated testing (unit, contract, and model-quality tests) as a standard across all engagements.
- Optimize model serving for high-throughput production using vLLM, DeepSpeed, or Triton Inference Server.
Client Solutioning & Leadership
- Lead technical discovery and proposal for AI engagements; translate ambiguous client problems into actionable AI solutions.
- Guide junior engineers, review architecture decisions, and build the team’s internal library of reusable AI patterns, accelerators, and playbooks.
- Present solution designs, demo prototypes, and communicate technical trade-offs clearly to client technical and business stakeholders.
MUST-HAVE QUALIFICATIONS
- 5–10 years in software engineering or data science, with at least 3 years in applied Gen AI/LLM engineering in a services or consulting context.
- Proven experience building production agents with LangGraph, CrewAI, AutoGen, or Semantic Kernel.
- Deep expertise in RAG architectures, vector databases (Pinecone, Qdrant, Weaviate), and embedding pipelines.
- Strong working knowledge of GPT‑4o, Claude 3.x/4.x, Gemini, and open-source models (Llama 3, Mistral).
- Hands‑on with AWS/Azure AI services, Docker, Kubernetes, and CI/CD workflows.
- Strong Python, FastAPI, SQL; software design patterns; a “software engineering first” approach to ML — with rigorous unit, integration, and model-quality testing.
- Proven track record taking LLM systems from prototype to production – owning deployment pipelines, observability, evaluation suites, guardrails, and ongoing model health in live client environments.
- B.Tech / B.E. / M.Tech in Computer Science or related discipline.
GOOD TO HAVE
- Experience with PEFT/LoRA fine-tuning workflows and serving optimized models.
- Hands‑on experience with vLLM, DeepSpeed, or Triton Inference Server for high-throughput model serving.
- Exposure to GraphRAG or ontology-based retrieval strategies.
- Experience with vision‑language models or multi‑modal agent pipelines.
- Certifications: AWS Solutions Architect, Azure AI Engineer Associate, or equivalent.