About the Role
The individual will lead the design, development, and deployment of complex, autonomous agentic AI solutions within production environments, working across multi-agent frameworks and orchestrating large language models. The position is for scalable applied AI, with a focus on delivering robust and fair agentic systems that deliver organizational value, autonomy, and operational efficiency.
Responsibilities
- Architect and deploy scalable multi-agent AI systems with advanced tool integration.
- Build and optimize AI pipelines using frameworks like LangGraph, ADK, and LangChain.
- Integrate agentic AI into enterprise workflows to automate complex business processes.
- Review and execute code written by AI agents to ensure correctness and maintainability.
- Review agent configurations, prompts, and tool wrappers to prevent behavioral or performance regressions.
- Audit agent-generated ML workflows to catch flaws like target leakage.
- Launch, monitor, and debug distributed training jobs and auto-repair loops.
- Design evaluation suites and end-to-end scenarios for non-deterministic agents.
- Debug agent workflows and align systems with regulatory and ethical standards.
- Ensure security, transparency, fairness, and reliability across AI systems.
- Mentor engineers and promote modern LLM practices across teams.
Requirements
- 4+ years of industry experience, including 1+ years with agentic orchestration frameworks (LangChain, LangGraph, ADK).
- Proficiency in multi-agent framework design, tool calling, API orchestration, and vector databases/RAG.
- Experience with continuous monitoring, telemetry, and observability tooling for non-deterministic AI agents.
- Strong background in LLM evaluation methodologies (code-based, LLM-as-a-Judge) and end-to-end testing.
- Strong Python skills (type hints, absl, absltest) and senior-level code/prompt reviewing capabilities.
- Practical knowledge of the ML lifecycle (data preprocessing, feature engineering, model design, JAX/TensorFlow).
- Experience launching and debugging distributed training jobs across cloud platforms (AWS, GCP, Azure).
- Knowledge of MLOps practices, data privacy, model security, and AI compliance regulations.
Preferred Skills
- On-device ML (TFLite, Edge TPU, Gemini Nano, latency and power profiling).
- Experience with Hugging Face, Neo4j, or knowledge graphs.
- Background in autonomous decision-making systems, anomaly detection, and dynamic process optimization.
- Track record of publishing, presenting, or open-sourcing agentic AI innovations.
- Strong problem-solving aptitude, collaborative mindset, and excellent communication skills.