Cloud AI Platform Architect

EY

Hyderabad, Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EY seeks a seasoned AI Platform & Infrastructure leader to architect and govern enterprise AI platforms across hybrid and multi-cloud environments. You will drive platform design for model hosting, inference endpoints, vector stores and API integration, while ensuring security, observability and cost efficiency.

Responsibilities include leading cloud automation, CI/CD/CT pipelines, containerization with Kubernetes, and MLOps governance to productionize GenAI use cases across BFSI, manufacturing,

Qualifications

  • B.Tech/B.E. in Computer Science/IT/AI/ML mandatory; M.Tech/MS in Cloud Computing, AI, Data Engineering or Machine Learning preferred.
  • Hands-on experience building enterprise AI platforms, model hosting, inference systems and vector databases.
  • Deep expertise in AWS, Azure and GCP with focus on AI and infrastructure services like SageMaker, Bedrock, Azure ML, Vertex AI and Kubernetes platforms.

Responsibilities

  • Lead the design and implementation of enterprise AI platforms across hybrid and multi-cloud environments.
  • Architect scalable cloud-native AI foundations using major cloud services and enterprise Kubernetes platforms.
  • Build secure and resilient infrastructure for AI model development, training, deployment and runtime operations.
  • Design reusable platform services for model hosting, inference endpoints, vector databases and API integration.
  • Establish platform observability with logging, monitoring and cost optimization for AI systems.

Skills

AI Platform
Cloud Architecture
MLOps
Kubernetes
Terraform
CI/CD
Security & Governance
Python

Education

B.Tech/B.E. in CS/IT/AI/ML
M.Tech/MS in Cloud Computing/AI/Data Eng/ML

Tools

AWS SageMaker
Azure ML
Vertex AI
EKS
AKS
GKE
Terraform
Bicep/ARM/CloudFormation

Job description

Key Responsibilities:

  • AI Platform & Infrastructure Engineering and Architecture Leadership
    • Lead the design and implementation of enterprise AI platforms that support AI/ML, GenAI, Agentic AI and RAG workloads across hybrid and multi-cloud environments.
    • Architect scalable cloud-native AI foundations using AWS SageMaker and Bedrock, Azure ML and Azure OpenAI, GCP Vertex AI and enterprise Kubernetes platforms such as EKS, AKS and GKE.
    • Build secure and resilient infrastructure for AI model development, training, deployment and runtime operations including compute, GPU environments, storage, networking, secrets management and access control.
    • Design reusable platform services for model hosting, inference endpoints, vector databases, prompt orchestration, agent runtime support and enterprise API integration.
    • Establish platform observability with centralized logging, monitoring, tracing, telemetry, performance diagnostics and cost optimization for AI systems.
    • Enable secure AI platform controls with policy enforcement, access governance, auditability and support for Responsible AI, compliance and risk requirements.
    • Drive standardization of AI platform architecture through reusable patterns, landing zones, environment templates and enterprise engineering best practices.
  • Client & Stakeholder Management
    • Serve as primary technical advisor to CxOs, account leadership teams and enterprise engineering stakeholders on AI platform strategy, cloud modernization and deployment architecture.
    • Conduct technical workshops demonstrating platform blueprints, deployment models, MLOps capabilities, observability approaches and AI operational readiness.
    • Bridge platform engineering with business strategy by translating complex infrastructure capabilities into scalable business outcomes and delivery roadmaps.
    • Build trusted relationships through hands-on PoC delivery, architecture discussions and strategic guidance on AI platform adoption
  • Cloud, Automation & Deployment Engineering Leadership
    • Lead automation of AI infrastructure provisioning using Infrastructure as Code tools such as Terraform, Bicep, ARM templates and CloudFormation.
    • Design and implement CI/CD and CT pipelines for AI platforms, model deployment services, APIs, microservices and environment lifecycle management.
    • Drive containerization and orchestration patterns using Docker and Kubernetes including workload isolation, autoscaling, high availability and release automation.
    • Build deployment frameworks for AI and GenAI applications using REST services, FastAPI, model serving frameworks, serverless patterns and integration middleware.
    • Automate configuration management, secrets handling, policy validation, artifact promotion and environment provisioning across dev, test and production landscapes.
    • Implement release engineering practices including blue-green deployment, canary rollout, rollback mechanisms and platform upgrade planning.
    • Ensure platform reliability, scalability and supportability through automated testing, operational runbooks, resilience engineering and incident readiness.
  • MLOps, Data Pipelines & Operationalization Leadership
    • Lead the setup and scaling of MLOps and LLMOps capabilities including experiment tracking, model versioning, model registry, artifact management, pipeline orchestration and deployment governance.
    • Establish standardized operationalization pipelines for model training, validation, packaging, deployment, monitoring and retraining across enterprise AI use cases.
    • Drive integration of AI platforms with batch and real-time data pipelines using Airflow, Prefect, Spark, Kafka, Databricks and cloud-native data services.
    • Design and govern reliable data flows for AI workloads including ingestion, transformation, feature engineering, feature store integration, vector store enablement and metadata lineage.
    • Implement monitoring for models and services covering drift detection, latency, throughput, failure diagnostics and service-level performance.
    • Collaborate closely with data science, AI engineering and application teams to productionize prototypes, notebooks and GenAI use cases into enterprise-grade deployments.
    • Drive operational best practices around reproducibility, traceability, auditability and support for regulated or risk-sensitive environments.
  • Asset and Accelerator Development
    • Lead creation of EY IP including AI platform reference architectures, reusable IaC modules, deployment blueprints, MLOps templates and observability accelerators.
    • Develop reusable frameworks for secure AI environment setup, platform monitoring dashboards, model deployment factories and RAG infrastructure enablement.
    • Package technical accelerators as EY market offerings for rapid client deployment across industries and cloud ecosystems.
  • Business Development & GTM Initiatives
    • Lead technical solutioning for AI platform transformation RFPs including architecture blueprints, deployment approaches, cloud landing patterns and live PoC demonstrations.
    • Develop industry-specific GTM strategies that combine EY assets with hyperscaler AI platform services and enterprise automation capabilities.
    • Collaborate with sales teams to position EY as a trusted partner for AI platform engineering, cloud AI operationalization and secure enterprise AI deployment.
  • Program Governance & Delivery Excellence
    • Oversee end-to-end AI platform transformation programs from infrastructure discovery through production deployment, operational handover and value realization.
    • Manage cross-functional delivery teams including cloud engineers, platform engineers, ML engineers, DevOps specialists and data engineers across multiple workstreams.
    • Ensure technical excellence, risk mitigation and KPI achievement in complex enterprise AI platform and deployment programs.
  • Practice Development & Team Leadership
    • Mentor consultants and seniors in AI platform engineering, cloud architecture, automation practices and MLOps operationalization across the broader EY GDS AI practice.
    • Lead technical recruitment and capability building for cloud AI platform engineering excellence.
    • Drive innovation PoCs showcasing enterprise AI platforms, scalable GenAI deployment models, advanced observability and sovereign AI infrastructure capabilities.

Required Skills & Qualifications:

Education:

  • B.Tech/B.E. (Computer Science / IT / AI / ML mandatory); M.Tech / MS in Cloud Computing, AI, Data Engineering or Machine Learning (preferred).

Core Technical Expertise (Hands-on Implementation Required):

  • AI Platform & Infrastructure Engineering: Strong hands-on experience in designing and managing enterprise AI platforms, model hosting environments, inference systems, vector database infrastructure, API-based AI services and secure runtime environments.
  • Cloud Platforms: Deep expertise in AWS, Azure and GCP with focus on AI and infrastructure services such as SageMaker, Bedrock, Azure ML, Azure OpenAI, Vertex AI, AKS, EKS, GKE, IAM, networking, storage and monitoring.
  • Automation & Deployment Engineering: Strong knowledge of Terraform, Bicep, ARM, CloudFormation, CI/CD pipelines, containerization, Kubernetes, deployment automation, microservices architecture and release engineering.
  • MLOps / LLMOps: Experience with MLflow, Kubeflow, Azure ML pipelines, Vertex AI pipelines, model registry, experiment tracking, model serving, deployment governance and monitoring.
  • Data Engineering & Operationalization: Understanding of ETL and ELT pipelines, Airflow, Prefect, Spark, Kafka, Databricks, feature stores, streaming and batch processing and production data pipelines for AI workloads.
  • Programming & APIs: Proficiency in Python, shell scripting, YAML, JSON, REST APIs, FastAPI and automation scripting for platform and cloud operations.
  • Security & Governance: Familiarity with platform security, secrets management, policy controls, auditability, observability and support for Responsible AI and enterprise governance requirements.

AI and Data Science Certifications (Good to have):

  • Microsoft Certified: Azure AI Engineer Associate / Azure DevOps Engineer / Azure Solutions Architect
  • AWS Certified Machine Learning Specialty / AWS DevOps Engineer / AWS Solutions Architect
  • Google Professional Machine Learning Engineer / Professional Cloud DevOps Engineer
  • Additional: Kubernetes certifications (CKA / CKAD), Terraform Associate, Databricks certifications, MLOps or cloud platform engineering certifications

Consulting & Leadership Experience:

  • 10-13 years in cloud platform engineering, AI platform engineering, DevOps, MLOps or AI consulting (Big 4 / tech preferred).
  • Proven track record leading enterprise AI platform implementations across BFSI, manufacturing, healthcare or public sector.
  • Experience with sovereign AI, regulated industry deployments and multi-cloud architecture programs.

Soft Skills:

  • Exceptional technical storytelling for CxO audiences.
  • Proven ability to influence senior stakeholders through working prototypes, architecture deep dives and platform transformation roadmaps.
  • Leadership of diverse technical teams with clear delivery focus.

Portfolio Requirements:

Hands-on experience in building production-grade cloud AI platforms, automated deployment frameworks, MLOps pipelines and enterprise AI infrastructure is mandatory for interview stage


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech S And T-Cloud AI Platform Architect-ISR-Manager-GDSF02
Tech S And T-Cloud AI Platform Architect-ISR-Manager-GDSF02

EY • Chennai District

On-site
INR 3,500,000 - 7,000,000
Tech S and T-Cloud AI Platform Architect-ISR-Manager-GDSF02
Tech S and T-Cloud AI Platform Architect-ISR-Manager-GDSF02

EY • Ernakulam

On-site
INR 6,000,000 - 9,000,000
Tech S And T-Cloud AI Platform Engineer-ISR-Senior-GDSF02
Tech S And T-Cloud AI Platform Engineer-ISR-Senior-GDSF02

EY • Mumbai

On-site
INR 3,500,000 - 5,500,000
Cloud AI Engineer
Cloud AI Engineer

EY • Chennai District, Bengaluru

On-site
INR 3,000,000 - 5,500,000
Cloud AI Platform Engineering Manager
Cloud AI Platform Engineering Manager

Ernst & Young LLP ( EY India ) • Dadri

On-site
INR 2,500,000 - 4,500,000
Principal AI Solutions Architect
Principal AI Solutions Architect

Metatron Hr Solutions Coimbatore • Chennai District

On-site
INR 3,600,000 - 6,000,000
Platform Engineer
Platform Engineer

Vriba Solutions • Bengaluru

On-site
INR 2,400,000 - 4,200,000
AI Practice Head
AI Practice Head

iProgrammer Solutions • Pune District

On-site
INR 2,500,000 - 4,500,000
AI Technical Lead
AI Technical Lead

Generac Power Systems • Pune District

On-site
INR 2,800,000 - 4,200,000
Agentic Data Delivery Lead
Agentic Data Delivery Lead

EXL • Maharashtra

On-site
INR 4,500,000 - 7,500,000