- Architect and deliver AI/ML/LLM-based solutions embedded into infrastructure operations, including ITOps/AIOps, self-healing systems, predictive capacity, and automated RCA
- Design and build agentic AI workflows for infrastructure automation, including ticketing, monitoring, remediation, and provisioning
- Evaluate, fine-tune, and integrate open-source and commercial LLMs into enterprise infrastructure tooling
- Build RAG pipelines, vector databases, and knowledge-grounding systems over infrastructure documentation, runbooks, and CMDB data
- Write production-grade Python code and Bash/PowerShell scripts
- Integrate AI solutions with Azure, AWS, GCP, ServiceNow, Datadog, Splunk, Prometheus/Grafana, and CI/CD pipelines
- Define and enforce AI governance guardrails for data privacy, model security, hallucination control, and cost/token management
- Partner with infrastructure leadership to identify high-ROI AI use cases and build the roadmap
- Mentor infrastructure engineers on AI-adjacent skills and act as internal AI Center of Excellence lead for the infrastructure vertical
- Own the POC-to-pilot-to-production lifecycle for AI initiatives, including MLOps/LLMOps practices
Requirements
- 10–12 years in Infrastructure/Cloud engineering or architecture, with the last 3–4 years focused on applied AI/ML/LLM work
- Strong hands-on coding ability in Python; scripting in Bash/PowerShell
- API integration and SDK usage, including OpenAI, Anthropic, LangChain, and LlamaIndex
- Understanding of LLM fundamentals: prompting, fine-tuning, embeddings, RAG, context windows, and tokens/cost
- Practical experience building agentic AI systems using LangGraph, AutoGen, CrewAI, or custom orchestration with tool-calling/function-calling
- Experience with vector databases such as Pinecone, Weaviate, FAISS, and Azure AI Search
- Deep infrastructure background in cloud architecture, networking, virtualization, ITSM, monitoring/observability, and automation using Ansible/Terraform
- Experience with MLOps/LLMOps, including model deployment, monitoring, versioning, and cost governance
- Strong architecture and solutioning skills
- Certifications in AWS/Azure/GCP AI or Solutions Architect are good-to-have
- Exposure to enterprise AI governance/responsible AI frameworks is good-to-have
- Experience presenting to CXO/leadership on AI strategy and business cases is good-to-have
- Prior experience in a Big 4/GDS/large enterprise infrastructure environment is good-to-have
- Open-source contributions or published POCs in AI/agentic systems are good-to-have
- Ability to translate ambiguous infrastructure pain points into AI-solvable use cases
- Strong stakeholder management
- Comfortable with hands-on coding/architecture and strategic roadmap/governance work
Core Competencies
Demonstrates expertise in architecting and delivering AI/ML/LLM-based solutions for infrastructure operations, with strong capabilities in Python coding, AI governance, and MLOps practices. Proven ability to mentor teams and translate infrastructure challenges into AI-driven solutions.
Highest-signal resume keywords
- AI/ML/LLM Solution Architecture
- Python Coding
- MLOps/LLMOps Practices
- API Integration
- Infrastructure Automation
ATS Optimization Keywords
Hard Skills
- Python
- Bash Scripting
- PowerShell Scripting
- API Integration
- LLM Fundamentals
- Agentic AI Systems
- Vector Databases
- Cloud Architecture
- MLOps
- Automation with Ansible/Terraform
Soft Skills
- Stakeholder Management
- Mentoring
- Strategic Roadmap Development
Certifications & Qualifications
- AWS AI Solutions Architect
- Azure AI Solutions Architect
- GCP AI Solutions Architect
Industry Keywords
- Infrastructure Operations
- ITOps
- AIOps
- Predictive Capacity
- Self-Healing Systems
- AI Governance
- Enterprise AI
- Big 4
- GDS
- ITSM
Tools & Technologies
- Azure
- AWS
- GCP
- ServiceNow
- Datadog
- Splunk
- Prometheus
- Grafana
- LangChain
- LlamaIndex