AI Python Engineer – Chatbot Backend (Dubai, on-site)
Europe, Ukraine HOT
About the role
This is an on-site position in the UAE. We are looking for an AI Engineer to design, build, deploy, and operate production conversational AI and LLM systems for complex, multi-step customer journeys.
You will work primarily in Python and own stateful agent orchestration, persistent context, tool and API integration, MCP, RAG, evaluation, observability, testing, and containerized deployment. This is a hands‑on role for an engineer who is comfortable moving beyond prototypes and supporting AI systems through the full production lifecycle.
What you will do
Stateful LLM applications and agent orchestration
- Design complex conversational workflows that perform multi‑step search, servicing, booking, payment, and related operational tasks.
- Build stateful agent graphs with primary assistants, specialist agents, sub‑agents, routing layers, and structured handoffs.
- Implement short‑ and long‑term context management using checkpointers, cache, and persistent stores so users can resume journeys across sessions.
- Apply planning‑and‑execution patterns for complex requests while keeping simple interactions fast and predictable.
- Implement tool binding, structured outputs, exception handling, retries, fallbacks, and recovery paths for production reliability.
- Support asynchronous and streaming agent execution with progress updates to user interfaces.
Tool integration, MCP and enterprise APIs
- Build and maintain MCP clients and servers for secure access to internal tools, APIs, and data sources.
- Develop production Python services that integrate with enterprise REST APIs for search, validation, servicing, cart, payment, and related workflows.
- Implement secure authentication and authorization flows using OAuth 2.0, OIDC, managed identities, and service credentials.
- Define typed tool schemas, validate inputs and outputs, and handle partial failures across multi‑step executions.
- Deliver streaming tool results through WebSockets, Server‑Sent Events, or equivalent real‑time interfaces.
RAG, data and state management
- Create focused RAG and FAQ workflows with reliable retrieval, grounding, and response traceability.
- Work with SQL, NoSQL, vector databases, caches, and persistent checkpoints across transactional and conversational workloads.
- Implement data ingestion, chunking, vector representations, semantic search, routing, and retrieval optimization where appropriate.
- Design robust state models and validation layers for agent memory, tool results, and application configuration.
Evaluation, observability and performance
- Build automated evaluation pipelines using scenario‑based tests, LLM‑as‑judge methods, regression suites, and synthetic‑user simulations.
- Instrument agent workflows for traces, structured logs, prompt versions, tool activity, errors, quality signals, latency, and cost.
- Create repeatable conversation‑review and feedback loops that support prompt improvement and model comparison.
- Develop concurrency‑safe load tests for agent endpoints and investigate bottlenecks across models, tools, databases, and streaming layers.
- Troubleshoot hallucinations, routing failures, state corruption, retrieval quality, latency, and integration defects in production.
Deployment and engineering operations
- Build typed, asynchronous, production‑grade Python services with clear packaging, configuration, and dependency management.
- Write unit, integration, contract, regression, and asynchronous tests for AI workflows and external integrations.
- Create container images and support Kubernetes‑based deployment across development, testing, and production environments.
- Contribute to CI/CD and GitOps workflows, including security scanning, smoke tests, image publishing, and controlled releases.
- Implement secrets management, health endpoints, structured logging, telemetry, graceful shutdown, and operational runbooks.
- Work closely with AI, backend, frontend, platform, security, and product specialists while owning technical deliverables end‑to‑end.
Frontend and conversational interface integration
- Define contracts for streaming responses, tool progress, errors, and structured widgets used by conversational interfaces.
- Collaborate with frontend engineers on TypeScript‑based chat components and real‑time user experiences.
- Ensure backend streaming behavior remains compatible with WebSocket, SSE, and component‑based UI patterns.
Required skills and experience
- Ability to investigate ambiguous problems, communicate clearly in English at an upper‑intermediate level or higher, and collaborate in an international onsite environment.
- Strong commercial experience with Python and production backend development, including asynchronous programming, typing, packaging, configuration, and error handling.
- Hands‑on experience building and operating LLM‑powered applications in production, not only notebooks or proof‑of‑concept demos.
- Practical experience with LangGraph, LangChain, or comparable frameworks for stateful agent orchestration.
- Deep understanding of agent architecture: state graphs, primary and specialist agents, routing, tool binding, checkpointers, streaming, memory, retries, and failure handling.
- Hands‑on experience with MCP or comparable tool‑use protocols, including secure connections to APIs, tools, and data sources.
- Experience building asynchronous APIs with FastAPI, Starlette, or comparable Python frameworks, including WebSockets or SSE.
- Experience integrating REST APIs and enterprise systems, including OAuth/OIDC authentication and authorization flows.
- Practical knowledge of Redis or equivalent state/cache technologies, SQL and NoSQL databases, and vector search.
- Strong understanding of RAG architecture, vector representations, retrieval optimization, grounding, and source traceability.
- Experience with Docker, Kubernetes, cloud‑native deployment, environment overlays, and GitOps‑based delivery.
- Experience with CI/CD, automated security and smoke checks, secrets management, monitoring, and production support.
- Strong testing discipline across unit, integration, async, API, regression, and load‑testing scenarios.
- Practical experience with observability for LLM applications, including tracing, structured logging, evaluation data, latency, and cost analysis.
- Strong software‑engineering habits: clean code, code review, documentation, version control, maintainable architecture, and operational ownership.
Will be a plus
- Experience with Langfuse, OpenTelemetry, or equivalent observability and prompt‑management tooling.
- Experience designing synthetic‑user simulators, scenario catalogs, LLM‑as‑judge evaluations, and regression suites.
- Experience with prompt versioning, A/B testing, conversation review, or online feedback and training loops.
- Hands‑on experience with LoRA/QLoRA fine‑tuning, local model serving, training‑data preparation, or multi‑GPU workflows.
- Experience with TypeScript, Angular, RxJS, and widget‑driven or component‑based conversational interfaces.
- Experience with vector representations, semantic routing, intent classification, or local inference components.
- Experience with load testing and performance analysis for AI or high‑concurrency API workloads.
- Experience with conversational analytics, data exploration, or operational dashboards.
- Knowledge of travel, transportation, booking, servicing, payment, route, fare, or loyalty workflows.
Work model and relocation
The role is full‑time and onsite in Dubai, UAE, for an initial one‑year assignment with possible extension. The selected candidate will be employed through the designated UAE company. Planned support includes the UAE employment visa, employee medical insurance, an initial flight to Dubai, a return flight at the end of the assignment, an approved broker fee, apartment‑search assistance and local arrival support (subject to the final written offer).