Site Reliability Engineer

DeepIQ

Hyderabad

On-site

INR 1,200,000 - 2,100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DeepIQ is seeking a hands-on Site Reliability Engineer to improve the reliability, security, scalability, and operational efficiency of our cloud platform. The role focuses on AWS-hosted services running on Amazon EKS, with GitOps using Argo CD and infrastructure as code with Terraform.

You will work with engineering teams to operate production environments, enhance observability, automate deployments, respond to incidents, and strengthen security.

Qualifications

  • 2–3 years of SRE, DevOps, cloud infra, or related role.
  • Hands-on AWS experience and Linux administration.
  • Experience with Kubernetes (Amazon EKS) and Terraform.
  • Familiarity with GitOps (Argo CD) and GitHub workflows.
  • Strong troubleshooting, documentation, and communication skills.

Responsibilities

  • Operate and improve production and non-production environments in AWS.
  • Manage Kubernetes workloads and EKS-hosted apps.
  • Automate infrastructure and deployments using Terraform and Argo CD.
  • Monitor reliability with Prometheus, Grafana, and related tooling.
  • Participate in on-call rotation and incident response.
  • Support security, cost optimization, and disaster recovery.

Skills

AWS
Kubernetes
GitOps
Terraform
CI/CD
Automation

Tools

Argo CD
Terraform
GitHub

Job description

DeepIQ is looking for a hands‑on Site Reliability Engineer (SRE) to help improve the reliability, security, scalability, and operational efficiency of our cloud platform.

Our applications are primarily hosted on AWS and deployed on Amazon EKS. We use Argo CD for GitOps deployments, Terraform for infrastructure as code, GitHub for source control and CI/CD workflows, and Linear for project and incident tracking.

In this role, you will work closely with engineering teams to operate our production environments, improve observability, automate infrastructure and deployments, respond to incidents, and strengthen the security of our cloud platform.

ResponsibilitiesCloud Infrastructure and Kubernetes
About the Role

DeepIQ is looking for a hands‑on Site Reliability Engineer (SRE) to help improve the reliability, security, scalability, and operational efficiency of our cloud platform.

Our applications are primarily hosted on AWS and deployed on Amazon EKS. We use Argo CD for GitOps deployments, Terraform for infrastructure as code, GitHub for source control and CI/CD workflows, and Linear for project and incident tracking.

In this role, you will work closely with engineering teams to operate our production environments, improve observability, automate infrastructure and deployments, respond to incidents, and strengthen the security of our cloud platform.

Cloud Infrastructure and Kubernetes
  • Operate, maintain, and improve production and non-production environments running in AWS.
  • Manage applications and supporting services deployed on Amazon EKS.
  • Troubleshoot Kubernetes workloads, networking, storage, resource utilization, and application deployment issues.
  • Manage Kubernetes resources such as deployments, services, ingress controllers, ConfigMaps, secrets, autoscaling policies, and Helm charts.
  • Improve the scalability, availability, and cost efficiency of the platform.
  • Support AWS services such as EC2, IAM, VPC, S3, RDS, Route 53, CloudFront, load balancers, and CloudWatch.
Infrastructure as Code and GitOps
  • Build and maintain reusable Terraform modules for AWS and Kubernetes infrastructure.
  • Manage application deployments using Argo CD and GitOps practices.
  • Maintain clear separation between development, staging, and production environments.
  • Review infrastructure changes through pull requests and automated validation.
  • Identify manual operational processes and replace them with reliable automation.
  • Help maintain GitHub-based CI/CD workflows for application builds, testing, security checks, and deployments.
Monitoring, Logging, and Observability
  • Build and maintain monitoring, alerting, logging, and dashboarding across applications and infrastructure.
  • Monitor platform health, application performance, Kubernetes workloads, and AWS services.
  • Define actionable alerts that minimize noise and identify real production issues.
  • Centralize and improve application, infrastructure, audit, and security logs.
  • Help implement and maintain tools such as Amazon CloudWatch, Prometheus, Grafana, OpenTelemetry, OpenSearch, or similar observability platforms.
  • Work with developers to improve application metrics, structured logging, distributed tracing, and health checks.
  • Define and track service-level indicators, service-level objectives, availability, latency, and error rates.
Reliability and Incident Management
  • Participate in production support and an on‑call rotation.
  • Investigate production incidents and restore services quickly and safely.
  • Perform root‑cause analysis and document incident findings, corrective actions, and preventive measures.
  • Create and maintain operational runbooks and troubleshooting documentation.
  • Improve backup, disaster recovery, high availability, and business continuity procedures.
  • Conduct capacity planning, resilience testing, and failure scenario reviews.
  • Track operational improvements, incidents, and follow‑up work in Linear.
Security and Compliance
  • Apply security best practices across AWS, Kubernetes, CI/CD pipelines, and application deployments.
  • Review and improve AWS IAM roles, policies, service accounts, and least‑privilege access.
  • Help manage secrets securely using AWS Secrets Manager, Kubernetes secrets, or similar solutions.
  • Implement container image scanning, dependency scanning, infrastructure scanning, and vulnerability management.
  • Support Kubernetes security practices such as RBAC, network policies, workload identity, pod security controls, and secure container configurations.
  • Monitor security events and assist with incident investigation and remediation.
  • Support patching, certificate management, access reviews, audit logging, and security compliance activities.
Collaboration and Continuous Improvement
  • Work closely with software engineers to improve application reliability and production readiness.
  • Participate in architecture, infrastructure, and deployment design reviews.
  • Help engineering teams understand operational risks and reliability requirements.
  • Promote automation, observability, security, and infrastructure‑as‑code best practices.
  • Document platform architecture, operational procedures, and technical decisions.
  • Identify opportunities to reduce cloud costs without negatively affecting reliability or performance.
Required Qualifications
  • 2–3 years of experience in site reliability engineering, DevOps, cloud infrastructure, platform engineering, or a related role.
  • Hands‑on experience working with AWS.
  • Practical experience operating Kubernetes environments, preferably Amazon EKS.
  • Experience troubleshooting Kubernetes applications, networking, resource constraints, and deployment failures.
  • Experience with Terraform or another infrastructure‑as‑code tool.
  • Experience with Git‑based development and deployment workflows.
  • Familiarity with GitOps and continuous delivery tools such as Argo CD.
  • Experience implementing or supporting monitoring, logging, dashboards, and alerts.
  • Working knowledge of Linux administration, networking, DNS, TLS, HTTP, and load balancing.
  • Ability to write automation scripts using Python, Bash, or a similar language.
  • Understanding of cloud security, IAM, secrets management, and least‑privilege access.
  • Strong troubleshooting, documentation, and communication skills.
Preferred Qualifications
  • Experience with Prometheus, Grafana, Amazon CloudWatch, OpenTelemetry, OpenSearch, or the Elastic Stack.
  • Experience managing Helm charts and Kubernetes ingress controllers.
  • Experience with GitHub Actions or another CI/CD platform.
  • Familiarity with service meshes, Kubernetes autoscaling, and cluster autoscaling.
  • Experience with container and infrastructure security tools such as Trivy, Checkov, tfsec, Snyk, or similar platforms.
  • Familiarity with AWS security services such as GuardDuty, Security Hub, Inspector, CloudTrail, AWS Config, and IAM Access Analyzer.
  • Experience supporting multi‑tenant SaaS applications.
  • Familiarity with SOC 2, ISO 27001, or other security and compliance frameworks.
  • Exposure to FinOps, AWS cost monitoring, and cloud resource optimization.
  • AWS or Kubernetes certifications are helpful but not required.
What Success Looks Like
  • Develop a strong understanding of DeepIQ’s AWS, EKS, and application architecture.
  • Improve visibility into application and infrastructure health.
  • Reduce noisy alerts and improve incident detection.
  • Automate recurring infrastructure and operational tasks.
  • Strengthen cloud and Kubernetes security controls.
  • Improve deployment reliability through Terraform, GitHub, and Argo CD.
  • Create useful operational documentation and incident‑response runbooks.
  • Help reduce production incidents, recovery time, and unnecessary cloud costs.
What We Are Looking For

We are looking for someone who is curious, dependable, and comfortable taking ownership of operational problems. You should enjoy troubleshooting complex systems, automating repetitive work, and collaborating with developers to make applications easier and safer to operate.

  • You do not need to know every tool listed in this description. However, you should have a strong foundation in AWS, Kubernetes, infrastructure automation, and production troubleshooting, along with the willingness to learn and take ownership.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
AWS DevOps Engineer
AWS DevOps Engineer

DataBeat.io Media • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Devops Engineer
Devops Engineer

SAIGroup • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Competitive compensation
Equity
Benefits
Site Reliability Engineer
Site Reliability Engineer

Insight Global • Bengaluru

On-site
INR 2,500,000 - 5,000,000
AWS DevOps Engineer
AWS DevOps Engineer

Keka Technologies Private Limited • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Site Reliability Engineer — Multi-Cloud Infrastructure
Site Reliability Engineer — Multi-Cloud Infrastructure

Skit.ai • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Cybage Software • Pune District

On-site
INR 2,500,000 - 4,000,000