DevOps Engineer

KnowledgeWorks Global Ltd.

Mumbai

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

KnowledgeWorks Global Ltd. in Mumbai is seeking an experienced Observability and DevOps leader to design and implement an end-to-end observability and alert stack for web and AI/ML services, covering Prometheus, Grafana, ELK/OpenSearch, and Datadog.

You will own scaling infrastructure across on-prem and cloud, apply IaC with Terraform and Ansible, lead CI/CD with Jenkins, GitHub Actions or ArgoCD, containerize with Docker and Kubernetes, and embed security checks into pipelines.

Qualifications

  • 5-10 years of hands-on DevOps/SRE/Infrastructure engineering experience, with 2-3 years in senior/lead capacity.
  • Strong Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket).
  • Cloud platforms exposure (AWS, Azure, and/or GCP) with on-premise/hybrid infra experience.
  • Expert-level Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi).
  • Kubernetes and Docker experience, including multi-cluster management and CI/CD end-to-end.

Responsibilities

  • Design and implement end-to-end observability & alert stack for web apps and AI/ML services.
  • Develop tooling like Prometheus, Grafana, ELK/OpenSearch, Datadog; enable AI observability signals.
  • Lead infrastructure scaling across on-prem and cloud with capacity planning and cost optimization.
  • Embed security checks (SAST/DAST, secrets scanning) into CI/CD pipelines (DevSecOps).
  • Provision and manage GPU infra for model deployment; optimize multi-tenant GPU usage.
  • Drive CI/CD pipelines (Jenkins, GitHub Actions, ArgoCD/Flux) and containerization ( Docker/Kubernetes ).
  • Collaborate with architects and data scientists to translate designs into infra and delivery plans.

Skills

DevOps/SRE leadership
Cross-functional collaboration
SRE best practices
Security and compliance awareness
Networking fundamentals
Automation scripting
Cloud architecture

Tools

Prometheus
Grafana
ELK/OpenSearch
Datadog
Terraform
Ansible
Jenkins
GitHub Actions
ArgoCD
Flux
Docker
Kubernetes
Helm
SAST/DAST tools
SonarQube
Snyk
Checkmarx
OWASP DependencyCheck
Trivy
Grype
Clair
Cosign/Sigstore
HashiCorp Vault
AWS Secrets Manager
Azure Key Vault
NVIDIA CUDA
Python
Bash
Go

Job description

Observability Strategy (Web & AI Applications)
  • Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services.
  • Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent
  • Build AI-specific observability: model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals.
Infrastructure Scaling & Reliability:
  • Design and manage infrastructure capable of scaling across on-premise data centers and public cloud
  • Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking.
  • Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable.
  • Lead disaster recovery, backup, and high-availability strategy for critical systems.
DevOps & CI/CD Delivery :
  • Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans.
  • Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or ArgoCD/Flux for GitOps) across multiple projects and teams.
  • Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management.
  • Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps).
  • Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model
  • Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM,
  • Optimize GPU utilization, cost, and throughput across multi-tenant workloads; manage CUDA/driver/toolkit
  • Collaborate with data science/ML engineering teams on MLOps pipelines - model versioning, experiment
  • 5-10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 2-3 years in a senior or lead capacity.
  • Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket).
  • Strong background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure.
  • Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi).
  • Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management.
  • Hands-on experience building and maintaining CI/CD pipelines end to end.
  • Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production.
  • Proficiency in scripting/automation languages: Python, Bash, and/or Go.
  • Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems.
  • Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans.
  • Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines.
  • Container and image security: vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore).
  • Familiarity with Secrets management and credential hygiene using tools such such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault
  • Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Devops Engineer
Devops Engineer

Larsen & Toubro • Chennai

On-site
INR 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

NAVVYASA CONSULTING PRIVATE LIMITED • Gurugram District

On-site
INR 800,000 - 1,200,000
DevOps Engineer
DevOps Engineer

Auric AI Labs • Bengaluru

On-site
INR 900,000 - 1,500,000
DevOps Engineer
DevOps Engineer

algoleap • Hyderabad

On-site
INR 1,800,000 - 2,400,000
AI DevOps Engineer — Mid/Senior Level
AI DevOps Engineer — Mid/Senior Level

Stackular • Hyderabad

On-site
INR 1,200,000 - 2,400,000
DevOps II
DevOps II

Capital Factory • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Architect - DevOps and ML Op's
Senior Architect - DevOps and ML Op's

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 6,000,000
AI/ML Engineer - AI Observability (ML Ops Engineer)
AI/ML Engineer - AI Observability (ML Ops Engineer)

Vconstruct • Pune District, Nagpur District

On-site
INR 2,400,000 - 4,200,000
DevOps Specialist
DevOps Specialist

Opstree • India

On-site
INR 1,800,000 - 3,400,000
DevOps Engineer
DevOps Engineer

Comviva • Maharashtra

On-site
INR 1,500,000 - 2,400,000