AI Infrastructure and Platform Architect
Ignatiuz is a digital transformation and intelligent workplace consulting company with offices in the US (PA) and India (Indore). Focused on accelerating digital performance through innovation and automation . Our team has worked with a variety of Fortune 500 clients and has a track record of delivering reliable, high-quality solutions that drive business success.
Job Description
Position Summary
We are seeking an experienced AI Infrastructure andPlatform Architect to design, optimize, and manage scalable AIinfrastructure and platforms across on-premises, cloud, and hybridenvironments.
The ideal candidate should have strong experience withGPU-based systems, AI/ML platforms, infrastructure architecture, performanceoptimization, capacity planning, and production support. The role will workclosely with AI/ML developers, DevOps engineers, data engineers, and solutionarchitects to improve the performance, reliability, scalability, and costefficiency of AI solutions.
Key Responsibilities
- Designand manage AI infrastructure for model training, fine-tuning, inference,computer vision, Generative AI, and LLM workloads.
- DefineCPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
- Reviewexisting hardware and platform performance and recommend upgrades oroptimizations.
- Performcapacity planning to support future workloads and minimize frequenthardware changes.
- Buildand maintain AI platforms using Linux, Docker, Kubernetes, GPUorchestration, and cloud services.
- Configureand manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AIacceleration technologies.
- Monitorsystem health, GPU utilization, memory usage, storage performance, andnetwork throughput.
- Diagnoseinfrastructure failures, system crashes, performance bottlenecks, andplatform outages.
- Implementmonitoring, alerting, backup, disaster recovery, security, and operationalbest practices.
- Preparearchitecture documents, hardware specifications, technicalrecommendations, and operational runbooks.
- Supportproduction deployment, troubleshooting, and continuous platformimprovement.
AI Solution Optimization
The candidate should also be capable of:
- Reviewingthe end-to-end AI solution and identifying performance, architecture, andinfrastructure gaps.
- Recommendingimprovements to scalability, reliability, maintainability, and costefficiency.
- SupportingAI/ML developers with model training and experimentation environments.
- Helpingreduce training time through GPU optimization, distributed training,resource tuning, and efficient data pipelines.
- Providingguidance on model accuracy, evaluation, hyperparameter tuning, andexperimentation practices.
- Improvingmodel-serving and inference performance.
- Mentoringexisting team members on AI infrastructure and production-readiness bestpractices.
Required Skills
- AIinfrastructure and GPU-based computing
- NVIDIAGPU architecture, CUDA, cuDNN, NCCL, and TensorRT
- PyTorch,TensorFlow, Hugging Face, or similar frameworks
- Cloudand on-premises AI platforms
- Infrastructuresizing and capacity planning
- Performancemonitoring and troubleshooting
- High-performancestorage and networking
- MLOps,CI/CD, automation, and Infrastructure as Code
- Monitoringtools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry
Qualifications
- Bachelor'sor Master's degree in Computer Science, Artificial Intelligence,Information Technology, Engineering, or a related field.
- Strongoverall experience in infrastructure, cloud, platform engineering,architecture, or AI systems.
- Aminimum of 3 years of direct, hands-on experience specifically workingwith AI infrastructure, machine learning platforms, GPU environments, orproduction AI workloads.
- Strongproblem-solving, troubleshooting, communication, and technicaldocumentation skills.
The candidate's total professional experience may besignificantly higher. However, at least three years should involve genuine,relevant, hands-on work with AI systems and platforms.
Added Advantage
Preference will be given to candidates who have experiencewith:
- LargeLanguage Models and Generative AI
- RAGand agentic AI systems
- Computervision and edge AI
- Model-servingplatforms
- AIperformance benchmarking
- FinOpsand infrastructure cost optimization
- HighPerformance Computing environments
Experience Validation
- AIinfrastructure or platforms they designed or managed
- GPUand hardware-sizing decisions
- Modeltraining or inference environments supported
- Performanceissues diagnosed and resolved
- Improvementsachieved in training time, utilization, reliability, or cost
- ProductionAI workloads they deployed or maintained
General DevOps, cloud, or system administration experiencewithout direct AI or machine learning exposure will not be sufficient for thisposition.