About the Role
We are looking for a hands‑on Lead Engineer AI/ML & Backend Systems to build and scale AI‑powered products and the backend platforms behind them.
This is not a research‑only role. The ideal candidate should combine strong AI/ML and Generative AI expertise with solid backend engineering, distributed systems, event‑driven architecture, observability, and production ownership.
You will work across AI systems, APIs, data pipelines, Go/Python services, messaging platforms, and production monitoring while providing technical direction and mentoring engineers.
Key Responsibilities
- Lead the design, development, and production deployment of AI/ML and Generative AI solutions.
- Build LLM‑powered applications using prompt engineering, RAG, embeddings, structured outputs, tool/function calling, and agentic workflows where appropriate.
- Design and build scalable backend services, APIs, microservices, and asynchronous processing systems.
- Develop production‑grade services using Python and contribute to or work with existing Go‑based services.
- Design and operate event‑driven systems using Kafka, RabbitMQ, Pub/Sub, or equivalent platforms.
- Apply distributed‑system patterns such as idempotency, retries, timeouts, ordering, backpressure, failure recovery, and eventual consistency.
- Build and manage data pipelines and work with relational databases, caches, search systems, and vector stores as required.
- Define systematic evaluation of AI outputs for accuracy, relevance, groundedness, consistency, and regression.
- Own production systems end to end: design, development, deployment, monitoring, optimization, and incident resolution.
- Build strong observability using metrics, logs, distributed tracing, dashboards, and alerts.
- Track and improve system availability, latency, throughput, error rates, AI quality, queue health, and infrastructure efficiency.
- Lead architecture reviews, code reviews, technical decisions, and engineering standards.
- Mentor engineers and collaborate with Product, Data, DevOps/SRE, and other engineering teams.
Must Have
- 4-7 years of relevant engineering experience with strong hands‑on software/backend development experience.
- Strong proficiency in Python and experience building production‑grade services.
- Hands‑on experience developing and deploying AI/ML and LLM‑based applications.
- Strong understanding of RAG, embeddings, prompt/context engineering, AI evaluation, and modern LLM application patterns.
- Working knowledge of Go (Golang) with the ability to understand, debug, and contribute to Go services.
- Strong understanding of backend architecture, APIs, microservices, concurrency, caching, and database design.
- Strong understanding of distributed systems and event‑driven architecture.
- Hands‑on experience with Kafka, RabbitMQ, Pub/Sub, or similar messaging/streaming platforms.
- Strong proficiency in SQL and relational databases.
- Strong production‑debugging and root‑cause‑analysis skills.
- Strong hands‑on experience with observability and monitoring, including metrics, logs, tracing, dashboards, and alerting.
- Experience with tools such as Prometheus, Grafana, OpenTelemetry, ELK/Kibana, Datadog, New Relic, or equivalent.
- Ability to design systems for scalability, reliability, performance, maintainability, and cost efficiency.
- Experience mentoring engineers and driving technical design and code quality.
Good to Have
- Strong hands‑on development experience with Go.
- Experience with LangGraph, LangChain, LlamaIndex, or equivalent AI orchestration frameworks.
- Experience with vector databases/search such as pgvector, FAISS, Qdrant, Pinecone, or equivalent.
- Exposure to MLOps/LLMOps, model/prompt versioning, AI regression testing, and guardrails.
- Experience with Redis, Elasticsearch/OpenSearch, Docker, Kubernetes, and cloud platforms.
- Experience building high‑scale, high‑throughput distributed systems.
Experience & Education
- 57 years of relevant experience across backend engineering, AI/ML, or applied AI.
- B.E./B.Tech/M.E./M.Tech in Computer Science, IT, AI/ML, Data Science, or a related discipline, or equivalent practical experience.
What Success Looks Like
You should be equally comfortable discussing model quality and AI evaluation as you are discussing API latency, Kafka consumers, retries, idempotency, distributed tracing, database performance, and production incidents.
Location: Noida Sector 135 (WFO)