AI Platform Engineer / Architect

HCLTech

Dadri

On-site

INR 4,500,000 - 7,500,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HCLTech is seeking an AI PRE Engineer (L4) to lead platform reliability and production readiness across AI workloads. You will drive standards for data, models, and security, and govern release pipelines with CI/CD and GitOps practices.

The role emphasizes cross-functional leadership, architecture across multi-cloud, and strong focus on observability, FinOps for AI, and scalable deployment of GPU-enabled clusters.

Qualifications

  • Bachelor's degree in Computer Science / Engineering is mandatory.
  • Master's preferred (Distributed Systems / Cloud / AI Systems).
  • 12-18 years overall IT / platform engineering experience.
  • 7-10 years in AI / platform / cloud engineering at scale.
  • 5+ years in architecture / PRE / SRE strategy roles.

Responsibilities

  • Define and govern enterprise-wide production readiness frameworks across platform, data, model, application, and security layers.
  • Establish PRE standards, policies, and maturity models.
  • Define and drive release certification governance across all AI workloads.
  • Own production readiness lifecycle from design to deployment to scale.
  • Define SLO/SLI/SLA frameworks and reliability engineering models.
  • Align SLOs with customer experience, business KPIs, and cost targets.
  • Publish enterprise AI reference architectures and standardize platform patterns.
  • Lead incident management and continuous reliability improvements.

Skills

Platform reliability
SRE strategy
Cross-functional leadership
Technical advisory

Education

Bachelor's degree in Computer Science / Engineering
Master's preferred (Distributed Systems / Cloud / AI Systems)

Tools

Kubernetes
LangChain
OpenAI

Job description

Location - Noida, Bengaluru, Chennai, Hyderabad and Pune.

Role

An AI PRE Engineer (L4) acts as a Platform Reliability Architect and Production Readiness Leader responsible for defining, governing, and scaling enterprise-grade AI/ML platforms. This role ensures AI systems are production-ready, resilient, observable, secure, compliant, and cost-efficient at scale, while driving standardization, automation, and platform-wide governance across AI engineering, SRE, DevOps, and MLOps disciplines. The L4 PRE acts as a strategic bridge between platform engineering, AI teams, and business leadership, ensuring reliability and scalability of AI solutions.

Responsibilities
Architecture & Production Readiness Governance
  • Define and govern enterprise-wide production readiness frameworks across: Platform, data, model, application, and security layers
  • Establish organization-wide PRE standards, policies, and maturity models
  • Define and drive release certification governance (PRR frameworks) across all AI workloads
  • Own production readiness lifecycle from design -> deployment -> scale
SLO/SLI Strategy & Reliability Engineering
  • Define and govern: SLO/SLI/SLA frameworks for latency, availability, quality, safety, drift
  • Error budget strategy at platform and business levels
  • Establish reliability engineering models for AI systems
  • Align SLOs with: Customer experience, business KPIs, and cost targets
Reference Architecture & Platform Design
  • Define and publish: Enterprise AI reference architectures for: LLM applications, RAG pipelines, vector stores, agent frameworks
  • Batch, real-time, and streaming inference systems
  • Standardize: Platform architecture patterns across multi-cloud and hybrid environments
  • Drive platform abstraction and reusable AI services
  • Design and architect the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.
Deployment & Release Engineering Governance
  • Prompt/model versioning, rollback, and release controls
  • Govern: Release pipelines, CI/CD, GitOps, and policy-as-code frameworks
  • Establish release risk assessment frameworks for AI workloads
Observability & AI Ops Strategy
  • Define enterprise observability strategy covering: Metrics, logs, traces, and AI-specific telemetry: Token usage, latency, model drift, GPU utilization, quality metrics
  • Standardize: Observability patterns for AI systems across platform
  • Drive adoption of: AIOps and predictive reliability models
Capacity Engineering & FinOps for AI
  • Own: Capacity planning and forecasting models for: Token usage
  • GPU/CPU scaling
  • Vector databases and caching layers
  • Define: Cost optimization strategies (FinOps for AI platforms)
  • Govern: Resource allocation, autoscaling policies, and utilization efficiency
Resilience Engineering & System Design
  • Define: Enterprise resilience patterns, including: Circuit breakers, fallbacks, retries, timeouts
  • Semantic caching and prompt optimization
  • Architect: High-availability AI systems with: Multi-region / multi-zone failover
  • Establish: DR/BCP strategies with defined RTO/RPO targets
AI Security, Risk & Compliance Governance
  • Define and govern: AI security frameworks: Secrets management
  • Private networking
  • Egress controlsModel and prompt security
  • Establish: Red-team validation and AI safety testing standards
  • Own: Compliance frameworks mapping (ISO 27001, SOC 2, GDPR/DPDP, HIPAA)
  • Ensure: Data privacy, auditability, and governance across AI platforms
Platform Engineering & Developer Enablement
  • Define: Enterprise delivery frameworks: CI/CD pipelines
  • Drive: Platform standardization and self-service enablement
  • Build: Reusable automation frameworks and PRE accelerators
  • Lead: Enablement programs (trainings, brown-bags, playbooks)
Cross-Functional Leadership & Stakeholder Engagement
  • Partner with: AI/ML teams, platform teams, DevOps, and business stakeholders
  • Act as: Technical advisor to leadership and customers
  • Validate: Architecture, NFRs (scale, latency, cost, safety) across programs
  • Drive: Enterprise-wide adoption of PRE practices
Incident Management & Reliability Operations
  • Own: Enterprise incident response frameworks for AI systems
  • Lead: Critical incident command and escalations
  • Govern: RCA, postmortems, and systemic reliability improvements
  • Establish: Continuous improvement and
Qualifications & Experience
  • Bachelor's degree in Computer Science / Engineering (mandatory)
  • Master's preferred (Distributed Systems / Cloud / AI Systems)
  • 12-18 years overall IT / platform engineering experience
  • 7-10 years in AI / platform / cloud engineering at scale
  • 5+ years in architecture / PRE / SRE strategy roles
Core Experience:
  • Enterprise-scale:
  • Production readiness frameworks, PRR governance
  • SLO/SLI/SLA design and reliability engineering
  • Deep expertise in:
  • Kubernetes platforms (EKS/GKE/AKS) at scale
  • GPU-aware AI infrastructure
  • Strong experience in:
  • MLOps / LLMOps / AI production systems
  • Deployment strategies and release governance
  • AI stack:
  • Frameworks like LangChain, OpenAI
  • Programming:
Certifications Required
  • NVIDIA Certified Professional - AI Infrastructure & Operations
  • NVIDIA DLI - Advanced AI Infrastructure / GPU platforms
  • Kubernetes Certifications (CKA / CKS - mandatory)
  • Cloud Architect Certifications (AWS/Azure/GCP - Professional level preferred)
  • DevOps / CI-CD certifications (preferred)
  • Linux Certifications (RHCE / LFCS - advanced preferred)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

HCLTech • Bengaluru

On-site
INR 3,000,000 - 6,000,000
AI Platform Architect
AI Platform Architect

HCLTech • Dadri

On-site
INR 4,200,000 - 7,000,000
Cloud AI Platform Architect
Cloud AI Platform Architect

EY • Hyderabad, Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
Senior Engineer _AI Platform
Senior Engineer _AI Platform

Net Connect • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Staff Consulting Architect - AI Infrastructure
Senior Staff Consulting Architect - AI Infrastructure

Nutanix • Pune District

Hybrid
INR 4,000,000 - 7,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence • Bengaluru

On-site
INR 400,000 - 700,000