Senior AI DevOps Engineer - Azure/AWS & ...
NVS/SADE--P/1875561
Posted On 30 Sep 2026
End Date 30 Sep 2027
Required Experience 8 - 12 years
Basic Section
Grade --
Employment Type Full Time
Employee Category --
Organisational
Group Company NewVison
Company Name New Vision Softcom & Consultancy Pvt. Ltd
Function Business Units (BU)
Organization Unit DevOps
Region APAC
Country India
Base office location Pune
Working model Hybrid
Weekly Off Pune office Standard
State Maharashtra
Skills
Skill
Graduation/Equivalent Course Post Graduate
CERTIFICATION
AWS Certified Machine Learning – Specialty (MLS-C01) Microsoft Certified: Azure AI Engineer – Associate Exam AI-100
Working Language
Job Description
Role Summary
We are seeking a Senior AI DevOps Engineer to design, automate, secure, and operate scalable platforms for artificial intelligence and machine learning workloads.
This hands-on role combines cloud engineering, DevOps, Kubernetes, infrastructure as code, MLOps, and LLMOps. The engineer will support the full AI lifecycle—from experimentation and model packaging through deployment, monitoring, optimization, retraining, and retirement.
The primary cloud environment is Microsoft Azure, with additional experience in AWS, including Amazon Bedrock, highly valued. The successful candidate will partner with data scientists, ML engineers, software engineers, security teams, and business stakeholders to deliver reliable, production-grade AI solutions.
Key Responsibilities
- Design and operate secure, scalable, and highly available platforms for machine learning, generative AI, and LLM workloads.
- Develop standardized deployment patterns for traditional ML models, deep learning models, LLMs, RAG applications, and AI agents.
- Support real-time, batch, asynchronous, and event-driven model-serving architectures.
- Deploy and manage model endpoints, embedding services, prompt-processing services, retrieval pipelines, and inference APIs.
- Optimize AI platforms for performance, availability, scalability, resilience, and cost.
- Partner with data science and ML engineering teams to productionize experimental models.
- Establish repeatable processes for model packaging, validation, deployment, rollback, versioning, and retirement.
CI/CD and DevSecOps Automation
- Design and maintain CI/CD pipelines using GitHub Actions and Azure DevOps.
- Automate deployments for application code, infrastructure, container images, model artifacts, and ML workflows.
- Implement automated testing for code, infrastructure, data pipelines, model packages, security controls, and deployment configurations.
- Establish promotion workflows across development, test, staging, and production environments.
- Automate model validation, deployment-readiness checks, approval gates, rollback, canary releases, and controlled production promotion.
- Integrate source control, artifact repositories, container registries, model registries, cloud services, and observability platforms.
- Apply GitOps, pull-request-based deployment, reusable workflow templates, and environment protection practices.
Containers and Kubernetes
- Design, deploy, and operate production-grade Kubernetes and Azure Container Instances environments for AI and ML workloads.
- Build secure, efficient Docker images for model-serving APIs, data-processing jobs, feature services, and LLM applications.
- Manage Kubernetes deployments, services, ingress, configuration, secrets, storage, namespaces, resource quotas, and autoscaling.
- Implement workload scheduling for CPU-, GPU-, and memory-intensive applications.
- Use Helm, Kustomize, Kubernetes operators, and platform integrations to support model serving, distributed training, experiment tracking, and batch inference.
- Troubleshoot production issues involving containers, networking, storage, scheduling, performance, and resource utilization.
- Establish standards for cluster security, upgrades, capacity planning, backup, and disaster recovery.
Cloud Engineering and Infrastructure as Code
- Develop reusable Terraform modules and Azure Bicep templates for multi-environment cloud infrastructure.
- Build secure cloud foundations using identity and access management, private networking, encryption, secrets management, policy enforcement, and centralized logging.
- Automate infrastructure provisioning, validation, drift detection, controlled changes, and rollback.
- Integrate infrastructure deployment with CI/CD pipelines and approval workflows.
- Monitor and optimize cloud costs, including compute, storage, data transfer, GPU capacity, and AI inference usage.
- Support high availability, business continuity, disaster recovery, and regional resiliency requirements.
MLOps and Model Lifecycle Management
- Implement MLOps practices covering experimentation, training, validation, registry management, deployment, monitoring, retraining, and retirement.
- Manage model registries and experiment-tracking platforms such as MLflow, Kubeflow, or comparable tools.
- Govern model versions, metadata, lineage, ownership, approval status, and deployment history.
- Establish reproducibility across code, data, model artifacts, dependencies, configurations, and runtime environments.
- Implement model rollback, challenger models, A/B testing, canary deployments, shadow deployments, and automated model-refresh workflows.
- Monitor model, data, and concept drift and support remediation through retraining or controlled model updates.
LLMOps and Generative AI Operations
- Deploy and operate hosted and self-managed LLMs in secure, production-grade environments.
- Build deployment workflows for prompt templates, embedding models, rerankers, retrieval services, and AI application components.
- Support RAG, semantic search, agentic workflows, tool-using applications, and vector database integrations.
- Manage model configuration, context windows, token limits, concurrency, batching, caching, rate limits, quotas, and RPM/TPM usage.
- Optimize model selection, latency, throughput, response quality, token consumption, and overall GenAI cost.
- Implement controls for sensitive-data exposure, prompt injection, jailbreaks, unauthorized access, unsafe responses, and data leakage.
- Apply GenAI SDKs and orchestration frameworks such as OpenAI/Azure AI SDKs, Google ADK, LangChain, and LangGraph.
Observability, Reliability, and Incident Management
- Implement observability across AI platforms, infrastructure, applications, models, inference endpoints, retrieval systems, and vector databases.
- Monitor latency, throughput, concurrency, errors, timeouts, token usage, GPU/CPU utilization, memory, queue depth, availability, model quality, and drift.
- Use tools such as Prometheus, Grafana, OpenTelemetry, Azure Monitor, CloudWatch, Langfuse, LangSmith, or comparable platforms.
- Create dashboards, alerts, health checks, runbooks, service-level objectives, and escalation procedures.
- Implement distributed tracing and correlation across AI application components.
- Lead incident response, root-cause analysis, remediation, and post-incident improvement activities.
- Design reliability patterns such as graceful degradation, fallback models, retry logic, circuit breakers, and disaster recovery.
Security, Governance, and Responsible AI
- Embed security controls throughout the AI, cloud, infrastructure, and software delivery lifecycle.
- Implement least-privilege access, identity federation, secrets management, network segmentation, encryption, and secure service-to-service communication.
- Integrate source-code, dependency, container-image, infrastructure, and model-artifact scanning into CI/CD.
- Apply policy-as-code and automated compliance checks.
- Maintain auditability for infrastructure changes, model releases, deployments, access events, and incidents.
- Partner with security, privacy, risk, and compliance teams to support responsible AI operations.
- Promote traceability, human oversight, access control, safety evaluations, explainability, and controlled model changes.
Technical Leadership and Collaboration
- Define reference architectures, platform standards, reusable components, and technical roadmaps.
- Lead architecture reviews, technical design sessions, code reviews, and production-readiness assessments.
- Mentor engineers and promote effective DevOps, SRE, cloud, and MLOps practices.
- Translate business and AI product requirements into secure, scalable, supportable solutions.
- Communicate technical risks, dependencies, delivery status, and trade-offs to engineering and business leadership.
- Create technical documentation, operating procedures, runbooks, and knowledge-transfer materials.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Science, or a related technical field; equivalent experience may be considered.
- Seven or more years of experience in DevOps, cloud engineering, platform engineering, software engineering, data engineering, MLOps, or a related discipline.
- At least three years of hands‑on experience supporting production AI, machine learning, or MLOps platforms.
- Strong experience with GitHub Actions or Azure DevOps, Docker, Kubernetes, and Terraform.
- Practical experience with Azure, AWS, or another major cloud platform.
- Experience deploying and operating ML models, LLMs, RAG solutions, or AI-enabled applications in production.
- Experience with MLflow, Kubeflow, or a comparable MLOps platform.
- Experience monitoring application and model performance, including inference latency, data quality, and model or concept drift.
- Proficiency in Python and familiarity with REST APIs, SDKs, JSON, YAML, and configuration-driven automation.
- Strong communication, problem‑solving, technical leadership, and cross-functional collaboration skills.
Preferred Qualifications
- Experience operating AI workloads across multiple cloud platforms.
- Experience with Azure Bicep, policy as code, and secure cloud landing zones.
- Experience with Pinecone, Milvus, or other vector databases.
- Experience with KServe, Seldon, NVIDIA Triton, BentoML, TorchServe, or comparable model-serving technologies.
- Experience with GPU scheduling, CUDA‑enabled workloads, quantization, batching, and inference optimization.
- Experience with LLMOps, RAG architectures, prompt lifecycle management, AI agents, and LLM evaluation.
- Familiarity with Azure OpenAI, Azure AI Foundry, Amazon Bedrock, Amazon SageMaker, or Google Vertex AI.
- Experience with responsible AI, model risk management, privacy, or regulatory controls.
- Relevant certifications such as CKA, Terraform Associate, AWS Certified Machine Learning Engineer, Azure AI Engineer Associate, Google Professional Machine Learning Engineer, or comparable credentials.
Representative Technology Environment
- Infrastructure: Terraform, Bicep, policy as code
- GenAI: LangChain, LangGraph, RAG, embeddings, vector databases, prompt management
- Security: IAM, Key Vault, Trivy, Snyk, SonarQube, Checkov, GitHub Advanced Security