AI Infrastructure and Platform Architect

Ignatiuz Inc.

Indore District

Vor Ort

INR 3.500.000 - 6.000.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine Bewerbung wie gemacht für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Ignatiuz Inc. in Indore (India) seeks an experienced AI Infrastructure and Platform Architect to design, optimize, and manage scalable AI infrastructure across on‑premises, cloud, and hybrid environments.

The role focuses on GPU‑based systems, AI platforms, and production support to ensure performance, reliability, and cost efficiency. The ideal candidate will collaborate with AI/ML developers, DevOps, data engineers, and solution architects to deliver robust AI solutions and scalable platforms.

Qualifikationen

  • Bachelor's or Master's degree in Computer Science, AI, IT, Engineering, or related field.
  • Strong experience in AI infrastructure, GPU environments, or production AI workloads.
  • Proven ability to design, deploy, and maintain AI platforms across cloud and on‑prem environments.

Aufgaben

  • Design and manage AI infrastructure for model training, fine‑tuning, inference, and LLM workloads.
  • Define CPU, GPU, RAM, storage, networking, cooling, and power requirements.
  • Review hardware and platform performance and recommend upgrades.
  • Perform capacity planning to support future workloads.
  • Build and maintain AI platforms using Linux, Docker, Kubernetes, and cloud services.
  • Configure NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI acceleration tech.
  • Monitor system health, GPU utilization, memory, storage, and network throughput.
  • Diagnose infrastructure failures, bottlenecks, and outages.
  • Implement monitoring, backup, security, and disaster recovery best practices.
  • Prepare architecture documents and runbooks; support production deployment.

Kenntnisse

AI infrastructure
GPU computing
NVIDIA CUDA cuDNN TensorRT
PyTorch TF HuggingFace
Cloud & on-prem platforms
Capacity planning
Monitoring & troubleshooting
High-performance storage & networking
MLOps & IaC
Prometheus/Grafana/NVIDIA DCGM/OpenTel

Ausbildung

Bachelor's or Master's in CS/AI/IT/Engineering or related field

Jobbeschreibung

AI Infrastructure and Platform Architect

Ignatiuz is a digital transformation and intelligent workplace consulting company with offices in the US (PA) and India (Indore). Focused on accelerating digital performance through innovation and automation . Our team has worked with a variety of Fortune 500 clients and has a track record of delivering reliable, high-quality solutions that drive business success.


Job Description

Position Summary

We are seeking an experienced AI Infrastructure andPlatform Architect to design, optimize, and manage scalable AIinfrastructure and platforms across on-premises, cloud, and hybridenvironments.


The ideal candidate should have strong experience withGPU-based systems, AI/ML platforms, infrastructure architecture, performanceoptimization, capacity planning, and production support. The role will workclosely with AI/ML developers, DevOps engineers, data engineers, and solutionarchitects to improve the performance, reliability, scalability, and costefficiency of AI solutions.


Key Responsibilities


  • Designand manage AI infrastructure for model training, fine-tuning, inference,computer vision, Generative AI, and LLM workloads.

  • DefineCPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.

  • Reviewexisting hardware and platform performance and recommend upgrades oroptimizations.

  • Performcapacity planning to support future workloads and minimize frequenthardware changes.

  • Buildand maintain AI platforms using Linux, Docker, Kubernetes, GPUorchestration, and cloud services.

  • Configureand manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AIacceleration technologies.

  • Monitorsystem health, GPU utilization, memory usage, storage performance, andnetwork throughput.

  • Diagnoseinfrastructure failures, system crashes, performance bottlenecks, andplatform outages.

  • Implementmonitoring, alerting, backup, disaster recovery, security, and operationalbest practices.

  • Preparearchitecture documents, hardware specifications, technicalrecommendations, and operational runbooks.

  • Supportproduction deployment, troubleshooting, and continuous platformimprovement.


AI Solution Optimization

The candidate should also be capable of:



  • Reviewingthe end-to-end AI solution and identifying performance, architecture, andinfrastructure gaps.

  • Recommendingimprovements to scalability, reliability, maintainability, and costefficiency.

  • SupportingAI/ML developers with model training and experimentation environments.

  • Helpingreduce training time through GPU optimization, distributed training,resource tuning, and efficient data pipelines.

  • Providingguidance on model accuracy, evaluation, hyperparameter tuning, andexperimentation practices.

  • Improvingmodel-serving and inference performance.

  • Mentoringexisting team members on AI infrastructure and production-readiness bestpractices.


Required Skills


  • AIinfrastructure and GPU-based computing

  • NVIDIAGPU architecture, CUDA, cuDNN, NCCL, and TensorRT

  • PyTorch,TensorFlow, Hugging Face, or similar frameworks

  • Cloudand on-premises AI platforms

  • Infrastructuresizing and capacity planning

  • Performancemonitoring and troubleshooting

  • High-performancestorage and networking

  • MLOps,CI/CD, automation, and Infrastructure as Code

  • Monitoringtools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry


Qualifications


  • Bachelor'sor Master's degree in Computer Science, Artificial Intelligence,Information Technology, Engineering, or a related field.

  • Strongoverall experience in infrastructure, cloud, platform engineering,architecture, or AI systems.

  • Aminimum of 3 years of direct, hands-on experience specifically workingwith AI infrastructure, machine learning platforms, GPU environments, orproduction AI workloads.

  • Strongproblem-solving, troubleshooting, communication, and technicaldocumentation skills.


The candidate's total professional experience may besignificantly higher. However, at least three years should involve genuine,relevant, hands-on work with AI systems and platforms.


Added Advantage

Preference will be given to candidates who have experiencewith:



  • LargeLanguage Models and Generative AI

  • RAGand agentic AI systems

  • Computervision and edge AI

  • Model-servingplatforms

  • AIperformance benchmarking

  • FinOpsand infrastructure cost optimization

  • HighPerformance Computing environments


Experience Validation


  • AIinfrastructure or platforms they designed or managed

  • GPUand hardware-sizing decisions

  • Modeltraining or inference environments supported

  • Performanceissues diagnosed and resolved

  • Improvementsachieved in training time, utilization, reliability, or cost

  • ProductionAI workloads they deployed or maintained


General DevOps, cloud, or system administration experiencewithout direct AI or machine learning exposure will not be sufficient for thisposition.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Architect
AI Architect

Larsen & Toubro • Chennai District

Vor Ort
INR 4.000.000 - 7.000.000
Lead Engineer – Ai Platform
Lead Engineer – Ai Platform

TutorCloud • Bengaluru

Vor Ort
INR 4.000.000 - 7.000.000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture in India • Mumbai

Vor Ort
INR 3.000.000 - 5.400.000
AI Engineer
AI Engineer

Saama Technologies Inc • Indien

Vor Ort
INR 4.200.000 - 6.000.000
Lead AI Engineer
Lead AI Engineer

Bridge Global • Ernakulam

Vor Ort
INR 3.000.000 - 5.200.000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions Corp. • Indore District

Vor Ort
INR 2.400.000 - 4.200.000
AI Architect
AI Architect

HSM Edifice Construction Services • Nagpur District

Vor Ort
INR 1.500.000 - 2.200.000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions • Indore District

Vor Ort
INR 2.800.000 - 4.200.000
Cloud Solution Architect – Enterprise AI Infrastructure & FinOps
Cloud Solution Architect – Enterprise AI Infrastructure & FinOps

Zydus Group • Ahmedabad District

Vor Ort
INR 4.000.000 - 7.000.000
Data Platform Architect - AI/ML Infrastructure on AWS
Data Platform Architect - AI/ML Infrastructure on AWS

Accenture India Private Limited • Bengaluru

Hybrid
INR 2.500.000 - 5.500.000