Senior AI DevOps Engineer - Azure/AWS & Python

NewVision Software

Pune District

Hybrid

INR 2,600,000 - 4,200,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

New Vision Softcom & Consultancy Pvt. Ltd in Pune, India, is seeking a Senior AI DevOps Engineer to design, automate, and operate scalable AI/ML platforms. The role emphasizes cloud engineering, DevOps, Kubernetes, IaC, MLOps, and LLMOps in Azure with AWS experience.

You will partner with data scientists, ML engineers, and security teams to deliver production-grade AI solutions, manage model endpoints, and implement robust CI/CD and security practices.

Qualifications

  • Bachelor’s degree in CS/Engineering or related field; equivalent experience considered.
  • Seven+ years in DevOps, cloud engineering, or MLOps related roles.
  • Hands-on with GitHub Actions or Azure DevOps, Docker, Kubernetes, Terraform.
  • Experience deploying ML models, LLMs, or AI-enabled apps in production.

Responsibilities

  • Design secure, scalable AI/ML platforms and endpoints.
  • Develop deployment patterns for ML/LLM workloads.
  • Support real-time and batch model serving architectures.
  • Automate CI/CD, testing, and promotion across environments.
  • Monitor, optimize, and secure AI platforms and costs.

Skills

Azure
AWS
Kubernetes
Terraform
CI/CD
MLOps
LLMOps
GitHub Actions
Azure DevOps
Python

Education

Bachelor’s degree in CS/Engineering/IT/Data Science
Post Graduate

Tools

Terraform
Azure Bicep
Docker
Kubernetes
MLflow
Kubeflow
LangChain
LangGraph
Pinecone

Job description

Senior AI DevOps Engineer - Azure/AWS & ...

NVS/SADE--P/1875561

  • Product Engineering
  • Pune

Posted On 30 Sep 2026

End Date 30 Sep 2027

Required Experience 8 - 12 years

Basic Section

Grade --

Employment Type Full Time

Employee Category --

Organisational

Group Company NewVison

Company Name New Vision Softcom & Consultancy Pvt. Ltd

Function Business Units (BU)

Organization Unit DevOps

Region APAC

Country India

Base office location Pune

Working model Hybrid

Weekly Off Pune office Standard

State Maharashtra

Skills
Skill

Graduation/Equivalent Course Post Graduate

CERTIFICATION

AWS Certified Machine Learning – Specialty (MLS-C01) Microsoft Certified: Azure AI Engineer – Associate Exam AI-100

Working Language
Job Description

Role Summary

We are seeking a Senior AI DevOps Engineer to design, automate, secure, and operate scalable platforms for artificial intelligence and machine learning workloads.

This hands-on role combines cloud engineering, DevOps, Kubernetes, infrastructure as code, MLOps, and LLMOps. The engineer will support the full AI lifecycle—from experimentation and model packaging through deployment, monitoring, optimization, retraining, and retirement.

The primary cloud environment is Microsoft Azure, with additional experience in AWS, including Amazon Bedrock, highly valued. The successful candidate will partner with data scientists, ML engineers, software engineers, security teams, and business stakeholders to deliver reliable, production-grade AI solutions.

Key Responsibilities
  • Design and operate secure, scalable, and highly available platforms for machine learning, generative AI, and LLM workloads.
  • Develop standardized deployment patterns for traditional ML models, deep learning models, LLMs, RAG applications, and AI agents.
  • Support real-time, batch, asynchronous, and event-driven model-serving architectures.
  • Deploy and manage model endpoints, embedding services, prompt-processing services, retrieval pipelines, and inference APIs.
  • Optimize AI platforms for performance, availability, scalability, resilience, and cost.
  • Partner with data science and ML engineering teams to productionize experimental models.
  • Establish repeatable processes for model packaging, validation, deployment, rollback, versioning, and retirement.
CI/CD and DevSecOps Automation
  • Design and maintain CI/CD pipelines using GitHub Actions and Azure DevOps.
  • Automate deployments for application code, infrastructure, container images, model artifacts, and ML workflows.
  • Implement automated testing for code, infrastructure, data pipelines, model packages, security controls, and deployment configurations.
  • Establish promotion workflows across development, test, staging, and production environments.
  • Automate model validation, deployment-readiness checks, approval gates, rollback, canary releases, and controlled production promotion.
  • Integrate source control, artifact repositories, container registries, model registries, cloud services, and observability platforms.
  • Apply GitOps, pull-request-based deployment, reusable workflow templates, and environment protection practices.
Containers and Kubernetes
  • Design, deploy, and operate production-grade Kubernetes and Azure Container Instances environments for AI and ML workloads.
  • Build secure, efficient Docker images for model-serving APIs, data-processing jobs, feature services, and LLM applications.
  • Manage Kubernetes deployments, services, ingress, configuration, secrets, storage, namespaces, resource quotas, and autoscaling.
  • Implement workload scheduling for CPU-, GPU-, and memory-intensive applications.
  • Use Helm, Kustomize, Kubernetes operators, and platform integrations to support model serving, distributed training, experiment tracking, and batch inference.
  • Troubleshoot production issues involving containers, networking, storage, scheduling, performance, and resource utilization.
  • Establish standards for cluster security, upgrades, capacity planning, backup, and disaster recovery.
Cloud Engineering and Infrastructure as Code
  • Develop reusable Terraform modules and Azure Bicep templates for multi-environment cloud infrastructure.
  • Build secure cloud foundations using identity and access management, private networking, encryption, secrets management, policy enforcement, and centralized logging.
  • Automate infrastructure provisioning, validation, drift detection, controlled changes, and rollback.
  • Integrate infrastructure deployment with CI/CD pipelines and approval workflows.
  • Monitor and optimize cloud costs, including compute, storage, data transfer, GPU capacity, and AI inference usage.
  • Support high availability, business continuity, disaster recovery, and regional resiliency requirements.
MLOps and Model Lifecycle Management
  • Implement MLOps practices covering experimentation, training, validation, registry management, deployment, monitoring, retraining, and retirement.
  • Manage model registries and experiment-tracking platforms such as MLflow, Kubeflow, or comparable tools.
  • Govern model versions, metadata, lineage, ownership, approval status, and deployment history.
  • Establish reproducibility across code, data, model artifacts, dependencies, configurations, and runtime environments.
  • Implement model rollback, challenger models, A/B testing, canary deployments, shadow deployments, and automated model-refresh workflows.
  • Monitor model, data, and concept drift and support remediation through retraining or controlled model updates.
LLMOps and Generative AI Operations
  • Deploy and operate hosted and self-managed LLMs in secure, production-grade environments.
  • Build deployment workflows for prompt templates, embedding models, rerankers, retrieval services, and AI application components.
  • Support RAG, semantic search, agentic workflows, tool-using applications, and vector database integrations.
  • Manage model configuration, context windows, token limits, concurrency, batching, caching, rate limits, quotas, and RPM/TPM usage.
  • Optimize model selection, latency, throughput, response quality, token consumption, and overall GenAI cost.
  • Implement controls for sensitive-data exposure, prompt injection, jailbreaks, unauthorized access, unsafe responses, and data leakage.
  • Apply GenAI SDKs and orchestration frameworks such as OpenAI/Azure AI SDKs, Google ADK, LangChain, and LangGraph.
Observability, Reliability, and Incident Management
  • Implement observability across AI platforms, infrastructure, applications, models, inference endpoints, retrieval systems, and vector databases.
  • Monitor latency, throughput, concurrency, errors, timeouts, token usage, GPU/CPU utilization, memory, queue depth, availability, model quality, and drift.
  • Use tools such as Prometheus, Grafana, OpenTelemetry, Azure Monitor, CloudWatch, Langfuse, LangSmith, or comparable platforms.
  • Create dashboards, alerts, health checks, runbooks, service-level objectives, and escalation procedures.
  • Implement distributed tracing and correlation across AI application components.
  • Lead incident response, root-cause analysis, remediation, and post-incident improvement activities.
  • Design reliability patterns such as graceful degradation, fallback models, retry logic, circuit breakers, and disaster recovery.
Security, Governance, and Responsible AI
  • Embed security controls throughout the AI, cloud, infrastructure, and software delivery lifecycle.
  • Implement least-privilege access, identity federation, secrets management, network segmentation, encryption, and secure service-to-service communication.
  • Integrate source-code, dependency, container-image, infrastructure, and model-artifact scanning into CI/CD.
  • Apply policy-as-code and automated compliance checks.
  • Maintain auditability for infrastructure changes, model releases, deployments, access events, and incidents.
  • Partner with security, privacy, risk, and compliance teams to support responsible AI operations.
  • Promote traceability, human oversight, access control, safety evaluations, explainability, and controlled model changes.
Technical Leadership and Collaboration
  • Define reference architectures, platform standards, reusable components, and technical roadmaps.
  • Lead architecture reviews, technical design sessions, code reviews, and production-readiness assessments.
  • Mentor engineers and promote effective DevOps, SRE, cloud, and MLOps practices.
  • Translate business and AI product requirements into secure, scalable, supportable solutions.
  • Communicate technical risks, dependencies, delivery status, and trade-offs to engineering and business leadership.
  • Create technical documentation, operating procedures, runbooks, and knowledge-transfer materials.
Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Science, or a related technical field; equivalent experience may be considered.
  • Seven or more years of experience in DevOps, cloud engineering, platform engineering, software engineering, data engineering, MLOps, or a related discipline.
  • At least three years of hands‑on experience supporting production AI, machine learning, or MLOps platforms.
  • Strong experience with GitHub Actions or Azure DevOps, Docker, Kubernetes, and Terraform.
  • Practical experience with Azure, AWS, or another major cloud platform.
  • Experience deploying and operating ML models, LLMs, RAG solutions, or AI-enabled applications in production.
  • Experience with MLflow, Kubeflow, or a comparable MLOps platform.
  • Experience monitoring application and model performance, including inference latency, data quality, and model or concept drift.
  • Proficiency in Python and familiarity with REST APIs, SDKs, JSON, YAML, and configuration-driven automation.
  • Strong communication, problem‑solving, technical leadership, and cross-functional collaboration skills.
Preferred Qualifications
  • Experience operating AI workloads across multiple cloud platforms.
  • Experience with Azure Bicep, policy as code, and secure cloud landing zones.
  • Experience with Pinecone, Milvus, or other vector databases.
  • Experience with KServe, Seldon, NVIDIA Triton, BentoML, TorchServe, or comparable model-serving technologies.
  • Experience with GPU scheduling, CUDA‑enabled workloads, quantization, batching, and inference optimization.
  • Experience with LLMOps, RAG architectures, prompt lifecycle management, AI agents, and LLM evaluation.
  • Familiarity with Azure OpenAI, Azure AI Foundry, Amazon Bedrock, Amazon SageMaker, or Google Vertex AI.
  • Experience with responsible AI, model risk management, privacy, or regulatory controls.
  • Relevant certifications such as CKA, Terraform Associate, AWS Certified Machine Learning Engineer, Azure AI Engineer Associate, Google Professional Machine Learning Engineer, or comparable credentials.
Representative Technology Environment
  • Infrastructure: Terraform, Bicep, policy as code
  • GenAI: LangChain, LangGraph, RAG, embeddings, vector databases, prompt management
  • Security: IAM, Key Vault, Trivy, Snyk, SonarQube, Checkov, GitHub Advanced Security
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GenAI Engineer - Database
GenAI Engineer - Database

Nielseniq India • Pune District, Mumbai, Gurugram District

On-site
INR 3,000,000 - 4,500,000
Manager AI
Manager AI

Xpheno • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior AI Engineer
Senior AI Engineer

Cummins • Pune District

On-site
INR 4,000,000 - 7,000,000
Manager AI-ML
Manager AI-ML

Ecolab Global Services • Bengaluru

On-site
INR 2,500,000 - 3,500,000
AIML
AIML

Qentelli LLC • Hyderabad

On-site
INR 1,800,000 - 3,600,000
Cloud Solution Architect – Enterprise AI Infrastructure & FinOps
Cloud Solution Architect – Enterprise AI Infrastructure & FinOps

Zydus Group • Ahmedabad District

On-site
INR 4,000,000 - 7,000,000
AI Engineer
AI Engineer

Jobvite, Inc. • Chennai District

On-site
INR 4,000,000 - 7,000,000
AI/ML Architect
AI/ML Architect

Indihire Consultants • Bengaluru

On-site
INR 3,500,000 - 7,000,000
AI Engineer
AI Engineer

Saama Technologies • Pune District

On-site
INR 1,800,000 - 2,800,000
Lead AI Engineer
Lead AI Engineer

Keka Technologies Private Limited • Nagar

On-site
INR 1,500,000 - 2,100,000