Get more replies from employers
Send a job-specific resume in minutes.
GyanSys Inc. in Bengaluru seeks a hands-on AI Deployment Engineer responsible for ML engineering, model deployment, governance and observability across cloud, on-premises, hybrid and air-gapped environments.
You will own the complete lifecycle of DL models, LLMs and SLMs, build MLOps/LLMOps pipelines, optimize performance and cost, and implement governance, monitoring and secure access with Keycloak.
Location: Bangalore (Working from Office / Hybrid)
We are seeking a hands-on AI Deployment Engineer specializing in ML Engineering, Model Deployment, Model Governance, and Model Observability. The engineer will own the complete lifecycle of Deep Learning models, LLMs, and SLMs across cloud, on-premises, hybrid, and air-gapped environments.
Build and manage MLOps and LLMOps pipelines.
Deploy, host, and scale Deep Learning models, LLMs, and SLMs and Inference optimisation
Manage end-to-end model lifecycle including versioning, deployment, rollout, rollback, and retirement.
Host models on Databricks, Kubernetes, OpenShift, and GPU-based infrastructure.
Implement model governance, lineage, approval workflows, and compliance controls.
Build model monitoring, observability, tracing, logging, and drift detection capabilities.
Optimize model performance, latency, throughput, GPU utilization, and cost.
Support cloud, on-premises, hybrid, and air-gapped environments.
Strong GPU knowledge including NVIDIA GPUs, CUDA, multi-GPU deployments, and inference optimization.
Experience in Model Registry, Model Governance, Model Monitoring, Drift Detection, and AI Observability.
Experience with Vector Databases (Pinecone, Chroma, FAISS, Milvus, Azure AI Search).
REST APIs, WebSockets, Streaming HTTP.
Experience with MLflow, OpenTelemetry, LangFuse, Splunk, and Grafana/ELK.
Experience across Cloud, On-Premises, Hybrid, and Air-Gapped environments.
Experience with Auth setup like Keycloak