A technology company located in Karnataka, Bengaluru is seeking an experienced professional to focus on system reliability and Kubernetes operations. The role entails designing scalable architectures, automating deployments, managing cloud services, and implementing monitoring solutions. The ideal candidate will collaborate with various teams, ensuring high availability and security of Kubernetes environments. Key responsibilities include incident response and capacity planning to accommodate future workloads, along with maintaining comprehensive documentation.
Responsibilities
Collaborate with software development teams to ensure reliability is prioritized.
Design scalable architectures for critical applications.
Manage high availability of Kubernetes clusters.
Drive automation to streamline deployment in cloud environments.
Implement CI/CD pipelines for Kubernetes applications.
Leverage cloud services to optimize Kubernetes solutions.
Implement monitoring solutions for applications and infrastructure.
Respond to incidents minimizing downtime.
Conduct security audits for Kubernetes environments.
Maintain documentation for processes and troubleshooting.
Collaborate with security to apply best practices and run audits.
Document configurations and troubleshooting steps.
Tools
Kubernetes
AWS
GCP
Azure
Terraform
Ansible
Job description
Responsibilities
System Reliability:
a. Collaborate with software development teams to ensure reliability is a key consideration throughout the software development life cycle.
b. Design and implement scalable and resilient architectures for mission-critical applications.
Kubernetes Operations:
a. Design, implement, and manage Kubernetes clusters, ensuring high availability, fault tolerance, and scalability.
b. Perform upgrades, patch management, and security enhancements for Kubernetes infrastructure.
Automation and Infrastructure as Code (IaC):
a. Drive automation efforts to streamline deployment, scaling, and management of applications on Kubernetes and/or cloud environments.
b. Implement CI/CD pipelines for deploying and updating Kubernetes applications.
c. Develop and maintain Infrastructure as Code scripts (e.g., Terraform, Ansible) for provisioning and managing cloud and container resources.
Cloud Integration:
a. Leverage cloud services (AWS, GCP, Azure) to optimize Kubernetes infrastructure and seamlessly integrate with other cloud-native solutions.
b. Implement best practices for deploying and managing Kubernetes on cloud platforms.
Monitoring and Alerting:
a. Implement effective monitoring and alerting solutions for Kubernetes clusters, applications, and underlying infrastructure.
b. Proactively identify and address performance bottlenecks and reliability issues.
Incident Response:
a. Respond to and resolve incidents related to Kubernetes infrastructure and applications, ensuring minimal downtime and impact on users.
b. Conduct post-incident reviews and implement improvements to prevent future issues.
Capacity Planning:
a. Perform capacity planning to ensure the Kubernetes infrastructure can accommodate current and future workloads in the cloud.
Security:
a. Collaborate with the security team to implement and maintain security best practices for Kubernetes environments in the cloud.
b. Conduct regular security audits and vulnerability assessments.
Collaboration and Documentation:
a. Work closely with development, operations, and other teams to ensure a collaborative approach to infrastructure and application reliability.
b. Maintain clear and comprehensive documentation for processes, configurations, and troubleshooting steps.