The client is looking at hiring a AI Lead Engineer / Technical Lead - AI & Generative AI for their team at Bangalore.
Relevant Experience: 7+ years in AI/ML, including at least 3+ years of hands-on experience building and deploying production-grade AI/ML systems.
Role Overview:
Client is looking for a highly experienced AI Lead Engineer to provide technical leadership for the design, development, deployment, and scaling of enterprise-grade AI/ML and Generative AI solutions.
The ideal candidate should combine strong hands-on engineering capabilities with AI architecture and technical leadership skills. This role requires deep expertise across traditional ML, Deep Learning, LLMs, RAG, Agentic AI, MLOps, cloud platforms, and production AI engineering.
The individual will be responsible for translating complex business requirements into scalable AI solutions, making critical architecture and model-selection decisions, leading experimentation, driving solutions from POC to production, and mentoring AI/ML engineers.
This is a hands-on technical leadership position. Candidates whose experience is primarily project/program management without significant AI engineering and architecture ownership may not be suitable.
Key Responsibilities:
1. AI/ML & Deep Learning Engineering
- Design, develop and optimize advanced Machine Learning and Deep Learning solutions.
- Build production-grade AI applications using Python and modern software engineering practices.
- Apply strong knowledge of data structures, algorithms, software architecture, design patterns and clean architecture.
- Evaluate and select appropriate ML/DL algorithms and model architectures.
- Work extensively with Transformers, NLP, LLMs and Deep Learning architectures.
- Lead advanced model evaluation, error analysis and systematic model improvement.
- Establish robust ML experimentation, benchmarking and statistical evaluation frameworks.
- Design and implement enterprise-grade Generative AI and LLM-based solutions.
- Implement RAG architectures, embeddings, semantic search and vector databases.
- Fine-tune and customize LLMs using techniques such as LoRA, PEFT and QLoRA.
- Design prompt engineering and LLM orchestration frameworks.
- Build Agentic AI architectures involving agents, tools and AI workflows.
- Establish LLM evaluation methodologies covering quality, accuracy, reliability and performance.
- Implement approaches for hallucination mitigation, guardrails and responsible model behavior.
- Evaluate foundation models and determine appropriate model selection based on business and technical requirements.
3. AI Architecture & Production Engineering
- Architect scalable, secure and reliable end-to-end AI/ML platforms.
- Design AI services using microservices, REST APIs and FastAPI/Flask.
- Build distributed and asynchronous AI processing architectures.
- Design scalable model-serving and inference architectures.
- Optimize GPU/CPU utilization, inference latency and throughput.
- Apply optimization techniques including quantization, distillation, batching and caching.
- Design AI systems for high availability, fault tolerance and scalability.
- Lead architecture and technical design reviews.
- Build and manage production AI/ML deployment pipelines.
- Work with Docker, Kubernetes and CI/CD pipelines.
- Implement ML lifecycle management using MLflow, Kubeflow or equivalent platforms.
- Design AI/ML solutions leveraging AWS SageMaker, Amazon Bedrock and other AWS AI services.
- Establish model versioning, deployment and rollback strategies.
- Implement model monitoring, drift detection and automated/controlled retraining strategies.
- Optimize cloud infrastructure for performance, scalability and cost.
- Work with advanced SQL and large-scale datasets.
- Design scalable data-processing pipelines supporting AI/ML workloads.
- Implement feature stores and appropriate data/model version-control mechanisms.
- Establish data-quality and governance practices for production AI.
- Build robust experiment tracking and reproducibility processes.
Technical Leadership Responsibilities:
The AI Lead Engineer will act as a technical leader for the AI engineering team and will be expected to:
- Define the technical roadmap and AI engineering strategy.
- Lead AI architecture and design reviews.
- Mentor AI/ML Engineers and Senior AI Engineers.
- Establish engineering, coding, experimentation and model-evaluation standards.
- Review technical designs, code, experiments and production architecture.
- Collaborate closely with Product, Engineering, Operations and Business teams.
- Translate business problems into technically feasible AI solutions.
- Drive AI initiatives from ideation/POC through production deployment and scaling.
- Make informed technical trade-offs across accuracy, latency, cost, scalability and maintainability.
- Create and maintain high-quality technical documentation.
- Communicate complex AI concepts and technical decisions effectively to technical and non-technical stakeholders.
Mandatory Skills:
1. AI/ML Technical Leadership
- Strong hands-on experience building AI/ML solutions.
- Advanced ML/DL algorithm and architecture expertise.
- Model experimentation, evaluation and optimization.
- Strong Python and production software engineering capabilities.
- Experience technically leading AI/ML Engineers.
- Generative AI
- Transformers
- NLP
- RAG
- Embeddings
- Semantic search
- LLM orchestration
- Agentic AI
- Model evaluation and guardrails
3. Production AI & MLOps
Strong experience taking AI/ML solutions from development into production, including:
- Docker
- CI/CD
- MLflow/Kubeflow or equivalent
- Versioning and rollback
- GPU/CPU optimization
- Scaling and production observability
4. Solution Architecture & Business Translation
Ability to:
- Understand complex business problems.
- Translate requirements into AI architecture.
- Decide between traditional ML, DL, RAG, LLM fine-tuning, prompting and hybrid approaches.
- Design scalable enterprise AI solutions.
- Make architecture decisions considering accuracy, scalability, latency, maintainability and cost.
5. Technical Leadership & Mentoring
Demonstrated experience in:
- Technical mentoring
- Coding and engineering standards
- Building technical capability within AI teams
- Cross-functional stakeholder collaboration
- Driving AI POCs into production
Preferred Skills:
- AWS AI/ML ecosystem experience.
- AWS SageMaker and Amazon Bedrock.
- Advanced GPU/inference optimization.
- LoRA, PEFT and QLoRA.
- Model quantization and distillation.
- Cloud infrastructure and AI cost optimization.
- Experience building enterprise-wide AI/GenAI platforms.
Minimum Qualifications:
- 7+ years of experience in AI/ML, Data Science, Machine Learning Engineering or closely related AI engineering roles.
- At least 3+ years of strong hands-on experience building and deploying production AI/ML systems.
- Proven technical ownership of complex AI/ML or Generative AI solutions.
- Demonstrated experience taking AI solutions from experimentation/POC to enterprise production.
- Strong technical leadership and stakeholder-management capabilities.
- AI/ML Technical Leadership - Mandatory
- Production AI & MLOps - Mandatory
- Solution Architecture & Business Translation - Mandatory
- Technical Leadership & Mentoring - Mandatory
Skill Sets :
AI/ML & Deep Learning: Advanced Python and production-grade software engineering
Software architecture, design patterns & clean architecture
Data structures & algorithms
Advanced ML/DL algorithms and model architecture selection
Transformers, NLP, LLMs and deep learning architectures
Advanced model evaluation, error analysis and model improvement
ML experimentation, benchmarking and statistical analysis
Generative AI / LLM: LLM fine-tuning and customization (LoRA, PEFT, QLoRA)
RAG architecture, embeddings, vector databases and semantic search
Prompt engineering and LLM orchestration
Agentic AI architectures and tool use
LLM evaluation, hallucination mitigation and guardrails
Model selection and foundation-model evaluation
AI Architecture & Engineering: Design scalable end-to-end AI/ML platforms
Distributed processing and asynchronous architectures
Model serving, inference optimization and GPU utilization
Quantization, distillation, batching and caching
High availability, scalability and fault-tolerant AI systems
MLOps & Cloud: Docker, Kubernetes and CI/CD
MLflow/Kubeflow or equivalent
AWS SageMaker, Bedrock and other AWS AI services
Model versioning, deployment and rollback strategies
Model monitoring, drift detection and retraining
Cloud cost and infrastructure optimization
Data & Platform: Advanced SQL
Large-scale data processing and data pipelines
Feature stores and data/version management
Data quality, data governance and experiment tracking
Leadership: Technical roadmap and AI strategy
Architecture/design reviews
Technical mentorship and team development
Cross-functional collaboration with Product, Engineering, Operations and Business
Translating business requirements into AI solutions
Driving POCs from experimentation through production
Technical documentation and stakeholder communication