Senior Site Reliability Engineer

Sails Software Solutions

Visakhapatnam

On-site

INR 2,800,000 - 4,000,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sails Software Solutions in Visakhapatnam, India seeks a Senior Site Reliability Engineer to architect and maintain robust cloud infrastructure. You will drive a DevOps culture, automate deployments, and ensure reliability for microservices using AWS, Kubernetes, and ECS.

The role emphasizes IaC, monitoring, incident response, and collaboration with development teams to optimize performance and security.

Qualifications

  • 6+ years in SRE/DevOps/Cloud Eng.
  • AWS services knowledge: EC2, S3, RDS, IAM, VPC, Lambda, CloudWatch.
  • Kubernetes proficiency and container orchestration.
  • Experience with ECS Fargate/EC2.
  • IaC with Terraform/CloudFormation/Pulumi.
  • Scripting: Python, Bash, or Go.
  • Networking, load balancing, DNS, firewall in cloud.
  • Microservices, API gateways, and service meshes.

Responsibilities

  • Design and build scalable, secure cloud infrastructure using AWS and Kubernetes.
  • Lead IaC efforts with Terraform or CloudFormation.
  • Develop security and cost-optimization practices.
  • Ensure observability with monitoring/logging (Grafana, Datadog).
  • Own SRE lifecycle including incident management and postmortems.
  • Collaborate with teams to deploy and operate microservices.

Skills

Site Reliability Engineering
DevOps
Cloud Engineering
AWS
Kubernetes
ECS/Fargate
IaC
Python
Bash
Go
Networking
Service Meshes
API Gateways

Tools

Terraform
CloudFormation
Pulumi
Grafana
Datadog
Jenkins
Helm
SonarQube
Istio
Linkerd

Job description

SRE-AWS/GCP Technical Skills


  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Engineering.

  • Expertise in AWS services such as EC2, S3, RDS, IAM, VPC, Lambda, CloudWatch, etc.

  • Strong knowledge of Kubernetes and container orchestration best practices.

  • Experience managing services on Amazon ECS (Fargate or EC2).

  • Proficient in infrastructure-as-code tools like Terraform, CloudFormation, or Pulumi.

  • Skilled in scripting languages such as Python, Bash, or Go.

  • Solid grasp of networking, load balancing, DNS, and firewall rules in cloud environments.

  • Deep understanding of microservices architectures, API gateways, and service meshes.


Soft Skills


  • Proven leadership and cross-functional collaboration skills.

  • Strong problem-solving and incident-resolution mindset.

  • Clear communication, documentation, and stakeholder reporting abilities.

  • Passion for continuous improvement and automation.


Preferred Qualifications


  • AWS certifications such as AWS Certified DevOps Engineer, Solutions Architect – Professional, or equivalent.

  • Familiarity with service meshes like Istio or Linkerd.

  • Experience with serverless architectures and event-driven systems.

  • Knowledge of regulatory compliance (SOC2, ISO 27001, GDPR) in cloud environments.


Skills – AWS Cloud, CICD, EC2, Kubernete, Grafana, Datadog, Python

SRE- AWS Job Summary We are looking for an experienced and driven Senior Site Reliability Engineer (SRE) to architect, implement, and maintain robust cloud infrastructure. This role demands a deep understanding of AWS, Kubernetes, ECS, and the ability to build scalable, secure, and highly available infrastructure from scratch. The ideal candidate will be a strong advocate for DevOps principles, automation, and reliability, and will possess the skills to support and optimize complex microservices-based architectures.


Key Responsibilities


  • Infrastructure Design & Implementation

  • Design and build highly scalable, fault-tolerant, and secure cloud infrastructure using AWS, Kubernetes, and ECS.

  • Lead efforts in infrastructure as code (IaC) using tools like Terraform or CloudFormation.

  • Develop and enforce best practices for infrastructure provisioning, security, and cost optimization.


System Reliability & Performance


  • Ensure availability, performance, scalability, and security of production systems.

  • Implement observability strategies including monitoring, logging, and alerting using tools such as Prometheus, Grafana, ELK, or Datadog.

  • Analyse system performance metrics and proactively identify potential issues and bottlenecks.


DevOps & Automation


  • Build and maintain CI/CD pipelines to streamline code deployments across environments.

  • Drive automation in infrastructure provisioning, configuration management, and operational tasks.

  • Ensure repeatable and reliable deployments using containers and orchestration tools like Kubernetes and ECS.


Service Management


  • Own the SRE lifecycle, including incident management, postmortems, root cause analysis, and runbook creation.

  • Collaborate closely with development and QA teams to ensure seamless microservices integration, deployment, and lifecycle management.

  • Maintain service-level objectives (SLOs), service-level agreements (SLAs), and error budgets.


Security & Compliance


  • Implement and enforce cloud security best practices for networking, identity and access management, and data protection.

  • Support audits, compliance assessments, and vulnerability remediation.

  • Monitor for security anomalies and work with security teams to respond to threats.


Technical Skills


  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Engineering.

  • Expertise in AWS services such as EC2, S3, RDS, IAM, VPC, Lambda, CloudWatch, etc.

  • Strong knowledge of Kubernetes and container orchestration best practices.

  • Experience managing services on Amazon ECS (Fargate or EC2).

  • Proficient in infrastructure-as-code tools like Terraform, CloudFormation, or Pulumi.

  • Skilled in scripting languages such as Python, Bash, or Go.

  • Solid grasp of networking, load balancing, DNS, and firewall rules in cloud environments.

  • Deep understanding of microservices architectures, API gateways, and service meshes.


Preferred Qualifications


  • AWS certifications such as AWS Certified DevOps Engineer, Solutions Architect – Professional, or equivalent.

  • Familiarity with service meshes like Istio or Linkerd.

  • Experience with serverless architectures and event-driven systems.

  • Knowledge of regulatory compliance (SOC2, ISO 27001, GDPR) in cloud environments.


Skills – AWS Cloud, CICD, EC2, Kubernete, Grafana, Datadog, Python

Key Responsibilities: Cloud Platform: GCP


  • Infrastructure Automation: Design, implement, and manage infrastructure as code using Terraform to provision and manage GCP resources.

  • Container Orchestration: Deploy and manage Kubernetes clusters, ensuring efficient operation of containerized applications.

  • Continuous Integration/Continuous Deployment (CI/CD): Develop and maintain CI/CD pipelines using Jenkins to automate application build, test, and deployment processes.

  • Containerization: Collaborate with development teams to containerize applications using Docker and manage deployments with Helm Charts.

  • Code Quality Assurance: Integrate and manage SonarQube to ensure code quality and security standards are met.

  • Monitoring and Logging: Implement and manage monitoring solutions using Datadog to ensure system health, performance, and security.

  • Collaboration: Work closely with cross-functional teams, including developers, QA, and operations, to streamline processes and improve productivity.


Requirements


  • Experience: 5+ years in DevOps or cloud engineering roles, with at least 3 years of relevant experience in the specified technologies.

  • Technical Proficiency:

    • Hands-on experience with GCP services and architecture.

    • Proficiency in Terraform for infrastructure as code implementations.

    • Strong understanding and experience with Kubernetes and Docker.

    • Experience in setting up and managing CI/CD pipelines using Jenkins.

    • Familiarity with Helm Charts for application deployment.

    • Experience with SonarQube for code quality analysis.

    • Proficiency in monitoring and logging tools, particularly Datadog.



  • Scripting Skills: Proficiency in scripting languages such as Bash or Python is an added advantage.

  • Strong problem-solving abilities and analytical thinking.

  • Excellent communication skills, both verbal and written.

  • Ability to work collaboratively in a team environment.

  • Strong organizational and time management skills.


Skills – Terraform, Kubernetes, Cluster, Docker, GCP, SonarQube Experience Level Senior Level
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE-AWS/GCP
SRE-AWS/GCP

Sailssoftware • Visakhapatnam

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems India Pvt Ltd • Chennai District

On-site
INR 1,800,000 - 3,200,000
Senior SRE Engineer
Senior SRE Engineer

V2 Solutions • Chennai District

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

WorkSpan • Bengaluru

On-site
INR 2,200,000 - 3,600,000