Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
HRS in Mohali, On-Site is seeking a Mid-Level AI Engineer to build production AI crews within the AI-Workflow programme. You will work with a seven-crew platform and deliver tested, instrumented, and documented outputs in a tight 10-day cycle, ensuring guardrails and human-in-the-loop gates are embedded.
Role focuses on DSPy-based orchestration, agent frameworks, and production integration with AWS Bedrock, LangGraph, and Terraform IaC. Strong Python and MLOps practices are required.
We are seeking a Mid-Level AI Agentic Engineer to join the AI-Workflow programme and build the autonomous crew systems that augment HRS operations across Finance, Controlling, Operations, Customer Service, Customer Experience, and HR. This is not a research role or a prototype environment — you will be building production AI crews that handle live operational workflows for real departments, with real outcomes measured from day one.
You will work within a "crews building crews" model: a platform of seven build agents (PM, Architect, Automation, QA, SRE, Documentation, Observability) scaffolds, tests, and documents the operational crews you build. Your job is to close the gap between agent scaffolded output and production-ready code — working directly with the Tech Lead, the PM, and the build agent platform to deliver tested, instrumented, and documented crews within a 10-day delivery lifecycle.
The mission is workforce augmentation. AI handles the volume. Humans handle the judgement. Every crew you build encodes that principle in every escalation boundary, every guardrail, and every human-in-the-loop gate.
Build operational AI crews from structured To-Be process descriptions using DSPy
typed signatures with assertion guards and agent workflow orchestration patterns
such as state machines, human-in-the-loop checkpoints, and resumable execution
Implement N8N workflow automation and JSON integration connectors linking
crews to operational systems including Zammad, Genesys, and enterprise back
office platforms
Work directly with the Automation Agent to scaffold DSPy modules and agent
workflows — extending and improving generated output, not accepting it verbatim
Design and encode specific, testable escalation boundaries for every crew before
shadow deployment — grounded in real process context, not generic confidence
thresholds
Deliver every crew with 100% unit test coverage, a complete runbook, and New
Relic instrumentation live before go-live — these are deployment gates, not
aspirational standards
Contribute reusable patterns to the shared crew library and peer-review modules
built by other engineers on the team
Implement agent memory management using explicit typed state, structured
context handling, and clear handoff boundaries to prevent context degradation
across multi-step operational workflows
Apply three-layer output validation — DSPy assertions, output validators, and policy
enforcer — on every crew module before merge
Build and validate test suites covering non-deterministic edge cases and failure
modes — not just happy paths — using the QA Agent's generated baseline as a
starting point
Integrate crews with AWS Bedrock model routing (Claude Haiku/Sonnet) and work
within the EKS and Terraform IaC stack managed by DevOps
Maintain guardrails configuration for every crew — escalation triggers, human
approval gates, and policy enforcement — encoded in config before any crew
enters shadow deployment
Participate in weekly DSPy evaluation cycles against gold-standard baselines to
validate crew output quality and flag drift
Instrument every crew with New Relic metrics from day one: throughput, error rate,
latency, escalation rate, and cost per task — observability is a deployment prerequisite, not an afterthought
Actively diagnose and resolve production failure modes: memory drift across multi
step workflows, hallucination under low-confidence RAG retrieval, context
degradation in long-running state machines, and prompt injection via untrusted
integration inputs
Use post-deployment observability data to identify improvement candidates and
raise them in RAID — closing the feedback loop the Observability Agent depends on
Contribute to the continuous improvement cycle: every crew in production is a
measurement and improvement loop, not a delivery milestone
Work within the 10-day delivery lifecycle — Request → Discovery → Design →
Development → QA → CI/CD → Monitoring → Continuous Improvement — delivering
to standard at each stage
Collaborate with the PM during Discovery to assess process automation feasibility
using FUDV scoring — frequency, uniformity, digitisation, volume — and push back
credibly where AI reliability or data quality is not there yet
Contribute to Thursday technical reviews and Friday retrospectives with substantive
input — not status updates but engineering judgment
Use the build agent platform as a personal productivity multiplier — flag platform
gaps via RAID rather than working around them silently
FOR THIS EXCITING MISSION YOU ARE EQUIPPED WITH…
3–5 years of experience in AI/ML development with 1+ years in agentic AI or
advanced LLM applications shipped to a production environment — not prototype
or hackathon experience
Hands-on experience with DSPy typed signatures and assertion guards — not just
LangChain familiarity
Practical exposure to at least one agent framework or platform such as LangGraph,
Google ADK, Amazon Bedrock AgentCore, LangChain, CrewAI, or equivalent; the
role values transferable agentic engineering patterns over any single required
framework
Experience building and committing N8N workflow automation in a production
codebase
Demonstrated ability to design specific, testable escalation boundaries in a live
operational AI system
Can show their work — a GitHub profile, a shipped system, or a concrete
before/after on a workflow they automated carries more weight than academic
credentials
Strong Python programming skills with AI/ML libraries and practical agentic
engineering patterns; able to work across frameworks when needed, with exposure
to at least one of LangGraph, Google ADK, Amazon Bedrock AgentCore, LangChain,
CrewAI, or equivalent
Production experience with AWS Bedrock or equivalent cloud-based LLM routing
and model management
Knowledge of vector databases, embedding systems, and retrieval-augmented
generation — including retrieval quality assessment and hallucination mitigation
Understanding of MLOps and AIOps practices: CI/CD for AI systems, evaluation
harnesses, and gold-standard baseline testing
Hands-on experience with New Relic or equivalent observability tooling for
production AI systems — metric design, dashboard instrumentation, and anomaly
diagnosis
Familiarity with containerisation, EKS, and Terraform IaC sufficient to work within a
DevOps-managed infrastructure without creating integration delays
Test-driven development for non-deterministic systems — 100% unit test coverage
before merge is a non-negotiable standard in this team
Experience with agile delivery in a timeboxed sprint model — able to take a
structured process description from design to shadow deployment within a 10-day
lifecycle
Strong code documentation discipline — every module peer-handoff ready, every
runbook complete during build, every decision traceable in RAID
Ability to work within an architecture set by a Tech Lead — executing with full
ownership and quality pride within defined guardrails, escalating cleanly when
constraints need revisiting
Writes to be understood, not to be impressive — RAID entries a director can triage, runbooks a department SME can follow, code a peer can extend without asking the
author
Calm under non-determinism — diagnoses production failures methodically using
observability data rather than thrashing or going silent
Dog-food mentality — uses the build agent platform to accelerate their own work
and actively contributes to improving it
Mission-driven — understands that the goal is workforce augmentation, not
automation for its own sake, and builds every crew with that principle at its centre
Domain experience in at least one of: Finance, Controlling, HR, Operations, Customer Service, or Customer Experience — brings escalation boundary instinct
that no intake card can fully replicate
Experience with enterprise system integrations: Zammad, Genesys, UiPath, or
equivalent CRM/CX/RPA platforms
Familiarity with ChromaDB/Milvus/pgVector or equivalent vector store for RAG
pipeline development
Experience contributing to a shared pattern library or internal engineering
knowledge base in a multi-engineer AI team