Architect

HCLTech

Dadri

On-site

INR 3,000,000 - 6,000,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

HCLTech is seeking an experienced AI/ML architect to own enterprise GenAI architecture and lead end-to-end deployment of domain-specific SLMs. You will design, tune, and serve multi-model pipelines with robust ingestion, RAG, and governance.

This role demands strong expertise in large-scale AI platforms and secure, scalable production operations. Ideal candidates have 10+ years in AI/ML, hands-on experience with GenAI systems, and a track record of delivering enterprise-grade solutions with

Qualifications

  • 10+ years of experience in AI/ML architecture and deployment.
  • Expertise in GenAI models and enterprise-grade workflows.

Responsibilities

  • Own enterprise GenAI architecture and build domain SLMs using models.
  • Deliver multi-turn operational workflows for triage, diagnosis, remediation, and escalation.
  • Build ingestion pipelines for PDFs, Office docs, SharePoint, and APIs.
  • Implement idempotency, deduplication, security controls, and lineage in ingestion.

Skills

Model Engineering
RAG & GraphRAG
Inference & APIs
Security & Quality

Tools

PyTorch
Transformers
Hugging Face
DeepSpeed
FastAPI

Job description

Experience: 10+ years

Key Responsibilities

  • Own enterprise GenAI architecture and build domain SLMs using Gemma, Llama, Qwen, Phi, Mistral, or equivalent models.
  • Deliver multi-turn operational workflows for triage, diagnosis, approval-gated remediation, validation, ticket updates, escalation, durable state, resumability, and loop prevention.

Data Preparation & Ingestion

  • Build event-driven and scheduled ingestion for PDF, DOCX, PPTX, XLSX, SharePoint, APIs, and approved email folders or attachments using change notifications, delta sync, checkpoints, and reconciliation.
  • Implement idempotency, incremental reprocessing, OCR/parsing, deduplication, malware and PII/secret controls, metadata and ACL preservation, security-trimmed retrieval, durable queues, retries, dead-letter handling, lineage, freshness, and ingestion monitoring.

Fine-Tuning & Alignment

  • Lead SFT/instruction tuning, continued pretraining, domain adaptation, LoRA, QLoRA, DoRA, DPO/RLHF, knowledge distillation, quantization-aware optimization, and model compression.
  • Engineer training datasets with cleaning, labeling, synthetic data, JSONL/chat formats, versioning, quality gates, and train/validation/test splits.

RAG & Grounded AI

  • Architect RAG with chunking, embeddings, hybrid dense/sparse search, metadata filtering, reranking, context assembly, citations, and grounded-ness evaluation using Qdrant or equivalent.
  • Build GraphRAG with Neo4j/AuraDB, Amazon Neptune, etc.; apply ontologies, entity resolution, graph traversal, subgraph retrieval, Cypher/open Cypher or Gremlin, and vector-graph hybrid retrieval for multi-hop reasoning and provenance.

Inference Engineering & Model Serving

  • Deliver low-latency serving with vLLM, TensorRT-LLM, SGLang, TGI, or Triton using dynamic batching, paged attention, KV/prefix caching, speculative decoding, streaming, routing, autoscaling, and INT4/INT8/FP8, AWQ, or GPTQ quantization.
  • Develop intelligent model routing across SLMs, specialist models, and frontier LLMs using rules, semantic or complexity classification, model cascades, and learned routing; select the lowest-cost model that meets task-fit, quality, latency, context, privacy, safety, region, and availability requirements, with budgets, token metering, caching, circuit breakers, fallback, and continuous evaluation of routing accuracy, escalation rate, cost per successful task, and quality or latency regression.
  • Declare and test TTFT, inter-token and p50/p95/p99 end-to-end latency, tokens/second, requests/second, concurrency, queue time, error rate, GPU/KV-cache utilization, context and token limits, cost, and performance under sustained, burst, failover, and degraded modes.

API, Platform & Production Operations

  • Build OpenAI-compatible APIs with structured outputs, function/tool calling, orchestration, API gateway, authentication, RBAC/ACL, rate limits, quotas, secrets, and audit logging.
  • Deploy through Docker, Kubernetes, Helm, Terraform, and CI/CD across Azure/AWS/GCP with multi-node/multi-zone HA, load balancing, autoscaling, automated failover, graceful degradation, blue-green/rolling release, rollback, cross-region DR, backups, point-in-time recovery, and tested RTO/RPO.

Evaluation, Security & Governance

  • Define release gates for task success, Exact Match/F1, retrieval precision/recall, grounded-ness, citation coverage, hallucination/refusal, schema validity, safety, latency, throughput, availability, RTO/RPO, and cost.
  • Implement layered guardrails across input, retrieval, dialogue, routing, tool execution, and output: prompt-injection/jailbreak defense, PII/secret masking, grounding checks, tool allowlists, least privilege, HITL/dual approval, blast-radius and retry limits, kill switches, and deterministic fallback.
  • Establish OpenTelemetry-compatible model/agent tracing across sessions and turns for model/prompt versions, retrieval, graph paths, tool calls, guardrail decisions, approvals, tokens, cost, errors, and component latency; provide redacted dashboards, SLO alerts, drift/anomaly monitoring, continuous evaluation, trace replay, audit lineage, and rollback signals.

REQUIRED SKILLS:

  • Model Engineering: Python, PyTorch, Transformers, Hugging Face, TRL/PEFT, SFT, LoRA/QLoRA/DoRA, DPO/RLHF, BF16/FP16, DDP/FSDP/ DeepSpeed ZeRO, MLflow / W&B.
  • RAG & GraphRAG: Chunking, embeddings, hybrid search, reranking, Qdrant, Neo4j/AuraDB, Amazon Neptune, etc.; ontology, entity resolution, Cypher/Gremlin, vector-graph retrieval, grounding and citations.
  • Inference & APIs: vLLM, TensorRT-LLM, SGLang, TGI/Triton, batching, KV/prefix cache, speculative decoding, quantization, FastAPI / OpenAI-compatible APIs, structured output and tool calling; intelligent model routing using rules, semantic/complexity classifiers, cascades, cost-quality-latency policies, FinOps budgets, metering, fallback, and routing observability.
  • Security & Quality: Layered guardrails, prompt-injection defense, PII/secrets, RBAC/ACL, HITL, auditability, OpenTelemetry tracing, evaluation, drift monitoring, latency/throughput/cost SLOs.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Analytics Engineer-Senior AI / RAG Platform Engineer
Analytics Engineer-Senior AI / RAG Platform Engineer

Trigyn Technologies Limited. • Delhi

On-site
INR 2,000,000 - 3,500,000
Lead AI Engineer
Lead AI Engineer

Keka Technologies Private Limited • Nagar

On-site
INR 1,500,000 - 2,100,000
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
AI Architect
AI Architect

WorkSpan Inc. • Bengaluru

On-site
INR 3,500,000 - 5,500,000
AI Engineer
AI Engineer

Keka Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
AI/ML Technology Architect - DaAI
AI/ML Technology Architect - DaAI

Infosys • Bengaluru

On-site
INR 2,500,000 - 4,800,000
Backend AI Engineer
Backend AI Engineer

Cirruslabs • Hyderabad, Bengaluru

On-site
INR 4,000,000 - 8,000,000
Interesting Job Opportunity: AI Engineer - LLM/RAG
Interesting Job Opportunity: AI Engineer - LLM/RAG

Prismforce • Maharashtra

On-site
INR 1,400,000 - 2,200,000
AI Engineer
AI Engineer

BT Group • Bengaluru

Hybrid
INR 2,500,000 - 4,500,000
Software Engineer III Chennai, India · On-site
Software Engineer III Chennai, India · On-site

Arcadia Power, Inc. • Chennai District

On-site
INR 2,000,000 - 4,000,000
Stock options
Hybrid work in Chennai
Medical insurance (self + 5 family)
+4