Engineering Manager

Darwinbox Digital Solutions Pvt. Ltd.

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Darwinbox Digital Solutions Pvt. Ltd. is seeking an Engineering Manager for AI Platform in Bengaluru on-site. You will lead a high‑performing team responsible for building scalable AI/ML platforms, MLOps pipelines, and GenAI capabilities, partnering with product and data science to deliver AI-powered solutions.

You will drive architectural decisions, oversee reliability, and mentor engineers while contributing hands-on when needed to ensure platform excellence and cost efficiency.

Qualifications

  • Extensive software engineering and engineering management experience.

Responsibilities

  • Lead and mentor a team of software engineers for AI platform capabilities.
  • Define technical direction, architecture, and execution plans for the team.
  • Drive hiring, onboarding, and performance management while fostering a culture of excellence.
  • Build scalable AI/ML lifecycle infrastructure including model development, deployment, and monitoring.
  • Collaborate with Product, Data Science, DevOps, Security, and Infrastructure teams.

Skills

Programming: Python/Java/Go/Scala
Distributed systems
Microservices
LLMs/Generative AI
Prompt management

Education

Bachelor's or Master's in CS/IT/Engineering/AI/ML

Tools

Docker
Kubernetes
CI/CD
Infrastructure as Code
Cloud Platforms (AWS/GCP/Azure)

Job description

Job Title: Engineering Manager - AI Platform
About the Role

We are looking for an experienced and technically strong Engineering Manager - AI Platform to lead the engineering team responsible for building scalable AI/ML platforms and infrastructure.

In this role, you will be responsible for the technical direction, architecture, delivery, reliability, and continuous improvement of AI platform capabilities that enable engineering, data science, and product teams to build and deploy AI-powered solutions at scale.

You will work closely with Product, Data Science, Data Engineering, DevOps, Security, Infrastructure, and Business teams to build reliable, secure, scalable, and cost-efficient AI/ML platforms.

The ideal candidate has strong experience in software engineering and engineering management , along with hands‑on experience in AI/ML platforms, MLOps, cloud infrastructure, distributed systems, and Generative AI technologies . The candidate should be comfortable making architectural decisions, mentoring engineers, managing delivery, and working hands‑on when required.

What You'll Do
Engineering Leadership & People Management

Lead, mentor, and manage a team of software engineers working on AI/ML platform capabilities.

Define engineering goals, technical direction, and execution plans for the team.

Drive hiring, onboarding, performance management, career development, and succession planning.

Establish a culture of engineering excellence, ownership, collaboration, innovation, and continuous learning.

Conduct regular one‑on‑ones, provide technical and career guidance, and support engineers in their professional development.

Build a high‑performing engineering team capable of delivering reliable and scalable AI platform solutions.

Identify skill gaps and create development plans to strengthen technical capabilities within the team.

Define and build scalable platforms supporting the complete AI/ML lifecycle.

Design infrastructure for model development, experimentation, training, evaluation, deployment, monitoring, and lifecycle management.

Build reusable APIs, SDKs, services, and platform components that enable product and data science teams to consume AI capabilities.

Develop and maintain scalable model serving and inference infrastructure.

Support machine learning workloads across batch and real‑time inference environments.

Build systems for model versioning, experimentation, deployment, rollback, and monitoring.

Improve the reliability, scalability, performance, and cost efficiency of AI workloads.

MLOps & LLMOps

Establish and improve MLOps capabilities for model development and production deployment.

Build CI/CD pipelines for machine learning models and AI applications.

Implement automated model testing, validation, deployment, monitoring, and rollback mechanisms.

Develop LLMOps capabilities supporting Generative AI applications.

Build infrastructure for LLM application development, evaluation, deployment, and monitoring.

Support prompt management, model evaluation, model versioning, and inference monitoring.

Establish processes for tracking model performance, latency, availability, usage, and cost.

Generative AI & LLM Platform

Build scalable infrastructure supporting LLMs, Generative AI, RAG, vector search, embeddings, and AI agents .

Design and develop reusable platform components for GenAI applications.

Integrate and manage multiple foundation models and model providers.

Build frameworks for prompt management, experimentation, evaluation, and observability.

Implement mechanisms to monitor token usage, inference latency, model quality, and infrastructure costs.

Establish appropriate security controls, guardrails, and governance mechanisms for GenAI applications.

Work with Product and Data Science teams to accelerate the development and deployment of AI‑powered products.

Define technical architecture and design scalable distributed systems for AI/ML workloads.

Lead architecture reviews and make technology and design decisions.

Design highly available, fault‑tolerant, secure, and observable systems.

Identify architectural bottlenecks and drive modernization initiatives.

Ensure platform components follow engineering standards for scalability, maintainability, security, and performance.

Evaluate new technologies, frameworks, and infrastructure solutions relevant to AI/ML engineering.

Balance short‑term delivery requirements with long‑term platform scalability and maintainability.

Design and operate AI/ML infrastructure across AWS, GCP, Azure, or equivalent cloud environments.

Work closely with DevOps and Infrastructure teams on compute, storage, networking, security, and deployment architecture.

Optimize cloud infrastructure and AI workloads for performance and cost.

Support GPU‑based workloads and scalable model inference infrastructure.

Implement infrastructure‑as‑code and automated deployment practices.

Ensure appropriate monitoring, logging, alerting, and observability across AI platform components.

Platform Reliability & Observability

Define and maintain reliability, availability, scalability, and performance standards for the AI platform.

Establish monitoring and observability for models, services, infrastructure, and AI applications.

Define and track SLIs, SLOs, SLAs, and operational metrics.

Lead incident management and root‑cause analysis for platform issues.

Implement proactive measures to reduce system failures and improve platform resilience.

Drive capacity planning and performance optimization.

Security, Governance & Responsible AI

Work closely with Security and Compliance teams to implement appropriate security controls for AI platforms.

Ensure secure handling of sensitive and regulated data used by AI/ML systems.

Implement access controls, authentication, authorization, secrets management, and data protection mechanisms.

Support model governance, auditability, lineage, and explainability requirements.

Establish appropriate controls around LLM usage, data exposure, prompt injection, model misuse, and sensitive information leakage.

Ensure AI platform architecture aligns with organizational security and compliance requirements.

Cross‑Functional Collaboration

Partner with Product, Data Science, Data Engineering, Security, DevOps, Infrastructure, and Business teams.

Translate business requirements into scalable technical solutions.

Work with stakeholders to define priorities, roadmap, milestones, and delivery plans.

Communicate technical architecture, risks, dependencies, and delivery status to senior leadership.

Identify opportunities where AI/ML can improve customer experience, operational efficiency, risk management, and business outcomes.

Establish best practices around coding standards, testing, code reviews, CI/CD, observability, and documentation.

Drive automation across development, testing, deployment, and operational processes.

Conduct architecture and code reviews to ensure engineering quality.

Improve development velocity and engineering productivity through reusable platform components and tooling.

Promote a culture of continuous improvement and technical innovation.

What We're Looking For
Educational Background

Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Artificial Intelligence, Machine Learning, or a related technical field.

Relevant certifications in cloud, AI/ML, or software architecture are an added advantage.

Professional Experience

10+ years of experience in software engineering, platform engineering, or related technical roles.

3+ years of experience managing engineering teams or leading large‑scale technical initiatives.

Strong experience building and operating distributed systems and cloud‑native platforms.

Hands‑on experience with AI/ML platforms, MLOps, ML infrastructure, or AI engineering.

Experience managing and mentoring software engineering teams.

Strong experience working with Product, Data Science, DevOps, Security, and Infrastructure teams.

Technical Expertise

Strong programming experience in Python, Java, Go, Scala, or equivalent languages .

Strong understanding of software architecture, distributed systems, microservices, APIs, and event‑driven architectures.

Experience with cloud platforms such as AWS, GCP, or Azure .

Strong knowledge of Docker, Kubernetes, CI/CD, Infrastructure as Code, and observability .

Experience with ML frameworks such as PyTorch, TensorFlow, or equivalent .

Experience with ML lifecycle and MLOps platforms such as MLflow, Kubeflow, Airflow, SageMaker, Vertex AI, or equivalent .

Experience designing and operating scalable model serving and inference systems.

Strong understanding of databases, data pipelines, distributed computing, and cloud infrastructure.

Experience with LLMs, Generative AI, RAG, embeddings, vector databases, prompt engineering, or AI agents.

Strong understanding of system design, scalability, reliability, security, and performance optimization.

Nice to Have
AI/ML & Generative AI

Experience building an AI/ML platform from the ground up.

Experience with LLMOps and production‑grade Generative AI systems.

Experience with RAG architectures, vector databases, embeddings, and semantic search.

Experience with AI agents and agent orchestration frameworks.

Experience with model evaluation and AI quality measurement frameworks.

Experience optimizing GPU utilization and model inference performance.

Strong experience with AWS, GCP, or Azure AI/ML services.

Experience with Kubernetes‑based ML infrastructure.

Experience with GPU clusters and distributed model training.

Experience with Terraform or other Infrastructure‑as‑Code technologies.

Experience with Kafka or other event‑streaming platforms.

Experience working in FinTech, banking, lending, payments, or other regulated industries.

Understanding of data privacy, information security, and regulatory requirements.

Experience building AI systems handling sensitive financial or customer data.

Familiarity with responsible AI, model governance, and AI risk management.

Experience managing teams of 8–20+ engineers.

Experience hiring and building engineering teams.

Experience working with senior engineering and business leadership.

Strong stakeholder management and communication skills.

Tools & Technologies

Programming: Python, Java, Go, Scala

Generative AI: LLMs, RAG, Vector Databases, Embeddings, AI Agents, Prompt Management

Data & Streaming: Kafka, Spark, Airflow, SQL

Observability: Prometheus, Grafana, OpenTelemetry, CloudWatch, Datadog, or equivalent

CI/CD: GitHub Actions, Jenkins, GitLab CI, ArgoCD, or equivalent

Security: IAM, Secrets Management, Encryption, API Security, Data Protection

Position: Engineering Manager - AI Platform

Employment Type: Full‑Time

Experience Level: Senior Management / Engineering Leadership (10+ Years)

Work Model: On‑site

Job Snapshot

Updated Date

22-09-2026

Job ID

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Amgen SA • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Senior Machine Learning Platform Engineer
Senior Machine Learning Platform Engineer

Amgen SA • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Manager AI
Manager AI

Xpheno • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Manager AI-ML
Manager AI-ML

Ecolab Global Services • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Lead Engineer
Lead Engineer

Tutorcloud Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Senior AIML Engineer
Senior AIML Engineer

MNC Group • Bengaluru

On-site
INR 4,000,000 - 6,500,000
AI Engineer
AI Engineer

Jobvite, Inc. • Chennai District

On-site
INR 4,000,000 - 7,000,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 4,000,000 - 6,000,000
AI Engineer
AI Engineer

Saama Technologies • Pune District

On-site
INR 1,800,000 - 2,800,000