About the Role
We are seeking an experienced and highly motivated Senior DevOps, System Administrator to design, implement, and manage secure, scalable, and highly available cloud-native infrastructure. The candidate should possess deep expertise in Linux, OpenShift, Kubernetes, CI/CD, containerization, cloud-native technologies, automation, and monitoring micro‑service‑based applications.
You will play a critical role in building and maintaining enterprise‑grade DevOps platforms, optimizing application delivery pipelines, improving operational efficiency, and ensuring infrastructure reliability, security, and compliance.
Key Responsibilities
- Platform Engineering
- Design, deploy, and manage enterprise‑scale OpenShift clusters across Development, UAT, DR, and Production environments.
- Ensure high availability, scalability, resiliency, disaster recovery, and performance of container platforms.
- Implement Kubernetes/OpenShift best practices including deployments, stateful sets, daemon sets, cron jobs, RBAC, network policies, ingress controllers, storage classes, persistent volumes, and operators.
- Manage container lifecycle, image repositories, and runtime security.
- DevOps & CI/CD
- Design, implement, and maintain end‑to‑end CI/CD pipelines.
- Automate build, test, security scanning, deployment, and rollback processes.
- Integrate CI/CD tools such as GitLab, ArgoCD (GitOps), Tekton, Maven, Gradle, Nexus Repository, SonarQube, JUnit, Trivy, and OWASP Dependency Check.
- Implement GitOps practices using ArgoCD.
- Manage branching strategies, pull requests, code reviews, release management, and version control.
- Containerization & Cloud‑Native
- Build and optimize Docker images following security and performance best practices.
- Deploy and manage large‑scale microservice architectures.
- Optimize application startup, container resource allocation, JVM tuning, and memory utilization.
- Implement readiness, liveness, and startup probes.
- Configure Horizontal Pod Autoscaler (HPA).
- Infrastructure & Automation
- Strong expertise in Linux administration (RHEL, Ubuntu).
- Develop Bash/Shell/Python scripts for automation.
- Automate operational tasks including backup, log rotation, certificate renewal, platform maintenance, and deployment automation.
- Manage Infrastructure as Code (IaC) using Helm Charts.
- Middleware & Web Servers
- Configure and troubleshoot NGINX, Apache HTTP Server, and IBM HTTP Server (IHS).
- Configure reverse proxy, load balancing, SSL/TLS, HTTP/2, WebSocket, CORS, compression, and security headers.
- Optimize web server performance and troubleshoot gateway, timeout, buffering, and proxy‑related issues.
- Database & Caching
- Configure and manage Redis Standalone and Redis Cluster.
- Monitor database connectivity and connection pools.
- Optimize HikariCP and JDBC connection pooling.
- Messaging & Streaming
- Deploy and administer Apache Kafka.
- Troubleshoot messaging bottlenecks and optimize throughput.
- Monitoring, Logging & Observability
- Design centralized monitoring and logging platforms.
- Configure and manage Prometheus, Grafana, Alertmanager, ELK Stack, and Loki.
- Analyze application, infrastructure, and Kubernetes/Openshift logs.
- Create dashboards, alerts, and SLA monitoring.
- Security & Compliance
- Implement DevSecOps best practices.
- Integrate ACS/ACM, SonarQube, and OWASP Dependency Check.
- Manage secrets, certificates, RBAC, network policies, security contexts, and pod security standards.
- Ensure compliance with enterprise standards including ISO 27001 and GRC policies.
- Incident & Operations Management
- Troubleshoot production issues involving OpenShift, Redis, Kubernetes, Grafana, Prometheus, networking, DNS, load balancers, storage, JVM, and databases.
- Perform root cause analysis (RCA).
- Manage incidents, problem records, and change requests according to SLA guidelines.
- Plan and execute Disaster Recovery (DR) and failover activities.
Preferred Qualifications
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- 4+ years of experience in DevOps and Platform Engineering.
- Hands‑on experience with OpenShift Operators and Operator Lifecycle Manager (OLM).
- Strong understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, SSL/TLS, load balancing, and firewalls.
- Experience with enterprise‑scale production environments supporting mission‑critical applications.
Preferred Certifications
- RedHat Certified Specialist in OpenShift Administration
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- Kubernetes Security Specialist (CKS)
- RedHat Certified Engineer (RHCE)
Soft Skills
- Strong analytical and troubleshooting skills.
- Excellent communication and documentation abilities.
- Ability to work independently and collaboratively in cross‑functional teams.
- Strong problem‑solving and decision‑making skills.
- Ability to manage multiple priorities in a fast‑paced environment.
- Commitment to continuous learning and adoption of emerging technologies.
Experience and Education
- 4 years of work experience; Bachelor of Engineering/Technology, MCA, or equivalent.