Job Purpose
The Linux & Kubernetes Administrator is responsible for the administration, maintenance, optimization, automation, and security of enterprise Linux platforms and containerized environments supporting mission-critical business applications. The role focuses on ensuring high availability, scalability, operational efficiency, and security compliance across cloud, on-premises and hybrid infrastructure environments leveraging Linux, Kubernetes, Docker, OpenShift, and cloud-native technologies
Roles & Responsibilities
Linux Infrastructure Administration
- Install, configure, administer, and support Red Hat Enterprise Linux (RHEL), Ubuntu, SUSE Linux, and other enterprise Linux distributions.
- Manage system lifecycle activities, including provisioning, patching, upgrades, hardening, and decommissioning.
- Monitor system health, availability, resource utilization, and performance metrics.
- Manage user administration, access controls, and privilege management.
Container Platform Administration
- Deploy, administer, and support container platforms including Kubernetes, OpenShift, Docker Enterprise, Rancher, and container runtime environments.
- Manage Kubernetes clusters across development, testing, and production environments.
- Configure and maintain Namespaces, Pods, StatefulSets, Deployments, Services, Ingress Controllers, and Persistent Volumes.
- Ensure container platform availability, scalability, and performance.
Cloud-Native Platform Management
- Support hybrid and multi-cloud containerized workloads running on L & T Vyoma, private cloud & Sovereign cloud environments.
- Implement cloud-native operational best practices for containerized applications.
- Manage container registries, image repositories, and artifact management solutions.
- Support microservices-based application deployments.
- Support AI/ML, GPU, and high-performance computing infrastructure.
- Support GPU/ AI environment
Platform Reliability & Performance Management
- Identify and resolve Linux and container platform performance issues.
- Conduct root cause analysis (RCA) for production incidents.
- Perform proactive health checks, capacity planning, and resource optimization.
- Implement monitoring and alerting for infrastructure and container workloads.
Security & Compliance
- Implement Linux OS hardening as per CIS, STIG, and organizational standards.
- Manage vulnerability remediation and security patching.
- Configure security policies for Kubernetes and OpenShift environments.
- Implement container image scanning, compliance monitoring, and runtime protection.
- Ensure compliance with ISO 27001, PCI-DSS, SOC2, and organizational governance requirements.
Automation & DevOps Enablement
- Develop automation scripts using Shell Scripting, Python, Ansible, and PowerShell where applicable.
- Implement Infrastructure as Code (IaC) using Terraform, Ansible, and GitOps methodologies.
- Support CI/CD pipelines integrated with Kubernetes and container platforms.
- Drive operational automation and service reliability improvements.
Backup, Recovery & Business Continuity
- Implement backup and recovery procedures for Linux and containerized platforms.
- Support disaster recovery planning and testing.
- Ensure platform resilience and adherence to RPO/RTO commitments.
- Participate in business continuity and disaster recovery exercises.
Monitoring & Observability
- Administer enterprise monitoring solutions such as Prometheus, Grafana, ELK, Splunk, Dynatrace, AppDynamics, and Zabbix.
- Configure dashboards, alerts, and operational reporting.
- Analyze trends and recommend performance improvements.
- Support observability initiatives across cloud-native platforms.
Stakeholder & Service Management
- Participate in Incident, Problem, Change, and Release Management activities.
- Collaborate with Application, DevOps, Cloud Engineering, Security, and Infrastructure teams.
- Support audits, compliance reviews, and customer governance meetings.
- Mentor junior administrators and provide technical leadership.
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Relevant Experience
- Minimum 8-12 years of experience in Linux System Administration.
- Minimum 5+ years of hands-on experience supporting Kubernetes/OpenShift platforms in production environments.
- Experience managing large-scale enterprise and cloud-native environments.
- Strong exposure to container orchestration technologies.
- Experience supporting mission-critical 24x7 managed services operations.
- Experience in DevOps, CI/CD, Infrastructure as Code, and automation frameworks.
- Understanding of SRE principles and platform reliability engineering practices.
- Scripting knowledge in Python, Bash, or PowerShell.