Platform Engineer

HCLTech

Bengaluru

On-site

INR 3,000,000 - 6,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HCLTech in Bengaluru seeks an AI PRE Engineer to ensure AI/ML platforms are production-ready, highly reliable, observable, secure, and cost-efficient. This role bridges AI engineering, SRE, DevOps, and MLOps, defining standards across platform, data, model, application, and security layers.

The candidate will manage NVIDIA Enterprise AI Software, deploy GPU-enabled Kubernetes clusters, establish SLO/SLI frameworks, and drive PRRs, rollout safety, and cost-aware operations in a multi-tenant

Qualifications

  • Bachelor’s degree in computer science, engineering, or information technology.
  • 8-12 years of IT or platform engineering experience.
  • 4-6 years implementing or managing enterprise platforms (AI, data, or cloud).
  • 3-5 years in architecture or platform strategy roles supporting multiple teams.

Responsibilities

  • Define and maintain production readiness standards across platform, data, model, application, and security layers.
  • Manage the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.
  • Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift; implement error budget policies.
  • Implement and configure the LLM inferencing services using kServe, NIM, AI Gateways.
  • Lead incident response for AI workloads; perform post-incident reviews and drive systemic fixes.

Skills

Python
Go
Java
CI/CD
Terraform
Docker
Kubernetes
AWS/GCP/Azure
Prometheus
Observability

Education

Bachelor’s degree in CS/Engineering/IT
Master’s in systems architecture/cloud/AI

Tools

Terraform
Jenkins
GitHub Actions

Job description

AI PRE Engineer (Platform Reliability / Production Readiness Engineer)
The Role

An AI PRE Engineer ensures AI/ML platforms are production-ready, highly reliable, observable, secure, and cost-efficient, bridging AI engineering, SRE, DevOps, and MLOps disciplines.

Responsibilities:
  • Define and maintain production readiness standards across platform, data, model, application, and security layers.
  • Manage the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.
  • Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift implement error budget policies.
  • Implement and configure the LLM inferencing services using kServe, NIM, AI Gateways
  • Curate deployment blueprints (canary/shadow, blue-green, A/B) for models and prompts with rollback guidance.
  • Standardize observability patterns for prompts, embeddings, latency, cost, quality, and safety telemetry.
  • Own capacity engineering (token/concurrency budgets, GPU/CPU sizing, vector scaling, cache hierarchies).
  • Define resilience patterns (timeouts, circuit breakers, fallbacks, idempotent retries, semantic/prompt caching).
  • Set AI security baselines (secrets, private networking, egress controls) and mandate red-team & safety evaluations.
  • Maintain compliance mappings (e.g., ISO 27001, SOC 2, GDPR/DPDP, HIPAA where applicable).
  • Provide CI/CD pipelines, SDKs, Helm/Terraform templates, and policy‑as‑code for consistent delivery.
  • Author PRR checklists, runbooks/playbooks, and DR/BCP blueprints (RTO/RPO, multi-region/site failover). Drive enablement (trainings, brown-bags) and maintain knowledge repositories and decision records.
  • Partner with solution teams to validate architecture and non‑functional requirements (scale, latency, cost, safety).
  • Conduct Production Readiness Reviews (PRRs) and certify releases across performance, security, privacy, and compliance.
  • Implement observability (tracing, metrics, logs), dashboards, and SLO burn and cost anomaly alerting.
  • Experience with different IDE such as Jupiter Notebook, Visual Studio Code, PyCharm, etc.
  • Familiar with AI related libraries like LangChain, PandasAI, OpenAI
  • Execute safe releases (canary/shadow/blue green), prompt/model versioning, feature flags, and rollback plans.
  • Lead incident response for AI workloads; perform post‑incident reviews and drive systemic fixes.
  • Govern token/cost budgets, autoscaling thresholds, and vector store performance for FinOps efficiency.
Qualifications & Experience
  • Bachelor’s degree in computer science, Engineering, or Information Technology
  • Master’s degree in systems architecture, Cloud Computing, or AI-related disciplines is preferred
  • 8-12 years of overall IT or platform engineering experience
  • 4-6 years implementing or managing enterprise platforms (AI, data, or cloud platforms)
  • 3-5 years in architecture or platform strategy roles supporting multiple teams or business units
  • Production readiness reviews, SLO/SLI/SLA design, incident management, RCA/postmortems, on-call support, and capacity planning for AI/ML platforms
  • Hands‑on experience with AWS/GCP/Azure, GPU‑aware infrastructure, Infrastructure as Code (Terraform), Docker, Kubernetes (EKS/GKE/AKS), and managing large‑scale, multi‑tenant clusters
  • Deploying ML/LLM workloads to production, model lifecycle management, RAG pipelines, safe rollouts (canary/shadow), rollback strategies, and managing inference scalability and latency
  • Metrics, logging, tracing, and alerting using Prometheus/Grafana/OpenTelemetry or cloud-native tools; monitoring AI‑specific signals such as model drift, latency, token usage, and GPU utilization
  • Strong coding (Python/Go/Java), CI/CD pipelines (GitHub Actions, Jenkins), GitOps, automated reliability tooling, security best practices (secrets management, access control, AI guardrails)
Certifications Required:
  • NVIDIA Certified Professional: AI Infrastructure & Operations
  • NVIDIA DLI - Deploying AI with Kubernetes & GPUs
  • NVIDIA DLI - Building AI Infrastructure with NVIDIA Technologies
  • Docker Certified Associate
  • Red Hat Certified System Administrator (RHCSA)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer / Architect
AI Platform Engineer / Architect

HCLTech • Dadri

On-site
INR 4,500,000 - 7,500,000
AI Platform Architect
AI Platform Architect

HCLTech • Dadri

On-site
INR 4,200,000 - 7,000,000
AI SRE/ AI Site Reliability Engineer
AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Devops Engineer
Devops Engineer

Airtel • Gurugram District

On-site
INR 800,000 - 1,400,000
AI Engineer
AI Engineer

STCO India • Hyderabad

On-site
INR 800,000 - 1,200,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Engineer _AI Platform
Senior Engineer _AI Platform

Net Connect • Bengaluru

On-site
INR 3,500,000 - 6,000,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
Systems Integration Specialist
Systems Integration Specialist

NTT DATA North America • Hyderabad

On-site
INR 2,000,000 - 3,500,000