We are seeking an experienced, proactive, and self-driven Senior DevOps Engineer (5–7 Years Experience) to manage end-to-end cloud and on-premise infrastructure, automation pipelines, and deployment frameworks for complex AI/ML, real-time voice, and data-intensive applications.In this role, you will bridge the gap between AI engineering, core software development, and infrastructure operations. You will take full ownership of designing, building, and maintaining robust CI/CD pipelines, container orchestration environments, data/messaging clusters (Kafka, ELK, MongoDB), and specialized GPU/CPU hardware configurations.
- Experience Required: 5 to 7 Years
- Reporting & Collaboration: Direct collaboration with Engineering Leads, AI/ML Engineers, and Project Managers/Scrum Leads.
Key Responsibilities
1. Cloud & Hardware Infrastructure Management (GPU & CPU)
- Resource Optimization: Oversee end-to-end hardware resource management, from GPU/CPU procurement estimations based on traffic engineering to server topology design and optimization.
- AI/ML Workload Support: Configure, tune, and maintain compute environments required to run Large Language Models (LLMs), Computer Vision (CV) solutions, and Real-Time Voice Bots using NVIDIA GPU, CUDA, and inference server platforms.
- Cloud Operations: Architect, manage, and scale cloud infrastructure across major platforms (AWS, Azure, or GCP, OCI).
2. Infrastructure as Code (IaC) & Container Orchestration
- Automation: Implement and maintain Infrastructure as Code using Terraform, Ansible, or equivalent automation frameworks.
- Containerization: Architect multi-container environments using Docker and manage production-grade cluster orchestration with Kubernetes (including Helm charts).
- Network & Security: Manage cloud networking, load balancing, DNS, SSL/TLS, reverse proxies (Nginx), IAM, access controls, and security configurations.
3. Data Platforms & Messaging Infrastructure
- Streaming & Search: Set up, configure, and maintain high-throughput streaming and logging environments (Kafka, ELK/OpenSearch), managing topics, partitioning, and multi-node cluster setups.
- Database Management: Manage and scale database clusters, specifically MongoDB, ensuring high availability, backups, and disaster recovery.
4. CI/CD & Enterprise Integrations
- Pipeline Automation: Design, implement, and standardise automated CI/CD pipelines (via GitHub Actions, GitLab CI/CD, Jenkins, or Azure DevOps) to replace manual deployment processes.
- Integrations: Assist in establishing API/system integrations with client enterprise platforms, such as ERP systems and accounting software (e.g., Tally).
5. Delivery, Reliability & Operational Leadership
- Ownership & Sprint Tracking: Drive daily priority management, task allocation, and tracking for infrastructure tasks; bridge technical teams and project managers/scrum leads.
- Monitoring & Alerting: Set up end-to-end monitoring and logging (Prometheus, Grafana, ELK) to ensure optimum system availability, performance, and log tracing.
- Disaster Recovery: Establish and enforce best practices for backup, failover, system availability, and disaster recovery.
Required Technical Skills & Qualifications
- Experience: 5–7 years of hands-on experience in DevOps, Cloud, and Infrastructure Engineering.
- OS & Scripting: Strong expertise in Linux system administration, networking, and Shell/Python scripting.
- Containers & Orchestration: Advanced experience with Docker and Kubernetes management.
- CI/CD: Expertise in setting up enterprise-level continuous integration and deployment pipelines.
- Data & Messaging Systems: Direct experience in managing Kafka clusters, ELK Stack, PostgreSQL and MongoDB databases.
- Hardware & Traffic Engineering: Proven track record in capacity planning, traffic management, and hardware estimation for data-heavy workloads.
Preferred & AI-Specific Skills
- GPU Infrastructure: Direct experience managing NVIDIA GPU environments, CUDA setups, vector databases, and model serving/inference environments for LLM or AI/ML workloads.
- Cloud & IaC: Deep knowledge of at least one major cloud platform (AWS, Azure, GCP) and IaC tools (Terraform, Ansible).
- Enterprise Integrations: Experience connecting client-side ERP systems or third-party enterprise platforms.
- Agile Leadership: Strong sprint management capability and experience operating in fast-paced or startup engineering environments.
What We Look For
- Ownership Mindset: Takes full responsibility for infrastructure stability, performance, and deployments across dev, staging, and production environments.
- Problem-Solving Skills: Strong capability to troubleshoot complex deployment, networking, or hardware-level issues independently.
- Communication & Collaboration: Clear visibility into sprint tasks, effective cross-team communication, and documentation skills.