Overview
About the Role:
The role will involve working across infrastructure, deployment, security, performance, and AI-related environments, while ensuring systems remain reliable, scalable, secure, and cost-effective.
Key Responsibilities:
- Manage and maintain AWS infrastructure, services, and cloud environments.
- Monitor and optimise server and cloud costs to improve infrastructure efficiency.
- Perform server debugging, troubleshooting, and performance optimisation.
- Conduct load testing and performance analysis to identify bottlenecks and improve system reliability.
- Set up and maintain Ansible automation pipelines and Jenkins CI/CD pipelines.
- Manage and work with different types of servers and hosting environments based on project requirements.
- Implement and manage user roles, access controls, and security policies across infrastructure and services.
- Maintain and strengthen server and infrastructure security, including access management and security best practices.
- Set up and optimise Ollama and AI-based environments for internal/project requirements.
- Use AI tools and effective prompting techniques to improve development, troubleshooting, automation, and operational workflows.
- Set up, configure, and manage infrastructure services such as Redis, Solr, OpenSearch, and similar technologies.
- Manage and maintain Supabase environments, including configuration, access, and operational requirements.
- Monitor infrastructure health, identify potential issues, and proactively resolve performance or availability problems.
- Collaborate with development and technical teams to ensure smooth application deployment and infrastructure operations.
- Continuously evaluate tools and processes to improve automation, scalability, reliability, security, and operational efficiency.
Requirement and Qualification:
- 2+ years of hands-on experience as a DevOps Engineer or in a similar infrastructure/cloud role.
- Strong hands-on experience with AWS and cloud infrastructure management.
- Experience with server administration, troubleshooting, debugging, and performance optimisation.
- Good understanding of CI/CD practices and experience working with Jenkins and Ansible.
- Experience in load testing, performance monitoring, and identifying infrastructure bottlenecks.
- Working knowledge of different types of servers and hosting environments.
- Understanding of user roles, permissions, access policies, and infrastructure security.
- Strong understanding of server and cloud security best practices.
- Comfortable using AI tools and prompting techniques for technical problem-solving, automation, and day-to-day work.
- Strong troubleshooting, analytical, and problem-solving skills.
- Ability to work independently as well as collaborate effectively with development and technical teams.