Site Reliability Engineer (SRE)
We are seeking an experienced Site Reliability Engineer (SRE) to ensure the availability, performance, scalability, and reliability of customer‑facing platforms. The SRE will work closely with DevOps, DBA, Development, and Security teams to provision infrastructure, deploy applications, automate workflows, and maintain operational excellence. This role has a direct impact on system stability, customer satisfaction, and overall platform performance.
Key Responsibilities & Deliverables
- Manage, monitor, and maintain highly available Windows and Linux environments
- Ensure system scalability by analyzing performance metrics and trends
- Handle routine service requests while identifying and implementing automation opportunities
- Design, implement, and maintain infrastructure as code (IaC) using Terraform, ARM Templates, and AWS CloudFormation
- Manage data backups and implement disaster recovery strategies
- Design and deploy CI/CD pipelines using GitHub Actions, Jenkins, Octopus, Ansible, and Azure DevOps
- Enforce security best practices throughout the software development lifecycle (SDLC)
- Follow and promote ITIL best practices and standards
- Act as a subject matter expert for emerging cloud technologies with a focus on AWS
Technical Skills & Proficiencies
- Strong hands‑on experience with AWS
- Experience administering Windows and Linux servers
- CI/CD experience with GitHub Actions, Jenkins, and Octopus
- Infrastructure automation using Ansible, Terraform, or similar tools
- Experience with Azure and AWS cloud services
- Experience with observability and monitoring tools such as New Relic, Application Insights, AppDynamics, or Datadog
- Hands‑on experience with Docker and Kubernetes
- Scripting skills in Bash, PowerShell, or Python
- SQL Server database maintenance and administration (preferred)
- Strong understanding of networking concepts including VNET, subnets, private link, and VNET peering
- Azure AD, OAuth, certificates
- AKS, App Services, ASE
- Load balancer, application gateway, firewall
- API Management
- Azure SQL and databases
- Experience analyzing application logs, IIS logs, system logs, security logs, and AWS CloudTrail events
Experience Requirements
- 5+ years of experience in SRE, DevOps, or system administration
- Proven expertise in supporting high availability Windows and Linux environments
- Strong experience with the WISA stack (Windows, IIS, SQL Server, ASP.NET)
- 3+ years of experience working with cloud platforms (AWS, Azure & GCP)
- 1+ year of experience working with container technologies such as Docker and Kubernetes
- Experience working in agile methodologies such as Scrum, Kanban, or Lean
Education
- Bachelor’s Degree or Diploma in Computer Science, Information Systems, or equivalent practical experience
Skills
- devops, gcp, azure, database, cloud, aws