Job Title: Senior MLOps + DevOps Engineer (On-Prem AI Platform)
Role Overview:
We are looking for a Senior MLOps + DevOps Engineer (8+ years) to architect, build, and scale AI/ML platforms in an on-prem enterprise environment.
This role requires end-to-end ownership of ML systems, infrastructure, CI/CD, and production reliability, enabling scalable deployment of machine learning and GenAI solutions.
Key Responsibilities:
1. Platform Architecture & Ownership
- - Design and own end-to-end ML platform architecture (data training deployment monitoring)
- - Define and enforce best practices for scalable and secure ML systems
- - Standardize MLOps + DevOps frameworks and processes
2. Model Deployment & Serving
- - Deploy and manage ML/LLM models on GPU-based on-prem infrastructure
- - Optimize inference performance (latency, throughput, batching)
- - Implement model versioning, A/B testing, and rollback strategies
3. CI/CD & Automation
- - Design and implement CI/CD pipelines for ML models, APIs, and data workflows
- - Enable automated testing, deployment, and release management
4. Infrastructure & Containerization
- - Manage Linux-based (RHEL preferred) on-prem infrastructure
- - Containerize applications using Docker
- - Deploy and orchestrate workloads using Kubernetes / OpenShift
- - Operate within restricted or air-gapped environments
5. Data & System Integration
- - Build pipelines integrating structured databases and high-volume logs/streaming data
- - Support batch and real-time inference architectures
6. Monitoring, Observability & Reliability
- - Implement end-to-end observability (model + infra)
- - Use tools like Prometheus, Grafana, ELK stack
- - Ensure high availability, SLA adherence, and incident response
7. GenAI & Advanced ML Systems
- - Deploy RAG pipelines and vector databases
- - Manage LLM serving frameworks
- - Work with agent orchestration frameworks
8. Leadership & Collaboration
- - Mentor engineers on MLOps and DevOps best practices
- - Collaborate with cross-functional teams
- - Drive design reviews and production readiness
Required Skills:
- - Strong Python and scripting (Bash)
- - Deep understanding of ML lifecycle and productionization
- - Experience deploying ML/LLM systems in production
- - Linux, Docker, Kubernetes/OpenShift
- - CI/CD tools (Jenkins/GitLab CI)
- - SQL and data pipeline experience
Good to Have:
- - GPU optimization knowledge
- - MLflow / Kubeflow
- - Terraform / Ansible
- - Experience in on-prem or restricted environments
Experience:
- 8+ years in MLOps / DevOps / Platform Engineering
- Proven experience scaling production ML systems
Ideal Candidate:
A hands-on platform architect who can operate across ML systems and infrastructure, driving automation, scalability, and reliability.