Senior DevOps Engineer

Goodfit

Gurgaon

On-site

INR 2,600,000 - 4,200,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Goodfit is seeking a Senior DevOps /SRE Engineer to design, build, and operate scalable cloud infrastructure, CI/CD pipelines, and containerized workloads using AWS, Kubernetes (EKS/ECS), and Terraform.

You will implement observability, security, and GenAI-enhanced processes, lead incident response, and collaborate with product and security teams to ensure 24x7 availability for financial applications.

Qualifications

  • AWS Certification required.
  • Hands-on with Kubernetes (EKS) and container orchestration.
  • Terraform and IaC experience.
  • CI/CD with Jenkins, Git workflows.
  • Security and compliance awareness in cloud and container environments.

Responsibilities

  • Design, build, and maintain CI/CD pipelines in AWS.
  • Manage multi-account AWS environments and EKS/ECS clusters.
  • Implement IaC with Terraform and CI/CD tooling.
  • Maintain observability with Prometheus, Grafana, New Relic.
  • Ensure security with IAM, Secrets Manager, CIS hardening.
  • Lead incident response and disaster recovery planning.
  • Collaborate with Dev, Security, and Product teams.
  • Explore GenAI integrations in DevOps workflows.
  • Define network and API gateway configurations.
  • Establish blue-green/canary deployments.
  • Mentor junior engineers.

Skills

AWS Expertise
Kubernetes & Containers
CI/CD
Infrastructure as Code
Monitoring & Observability
API Management
Linux Administration
Scripting Bash/Python
Networking
Security & Compliance
GenAI in DevOps
Soft Skills
Agile/Scrum

Education

AWS Certification

Tools

Jenkins
Git
Helm
Terraform
Prometheus
Grafana
New Relic
Trivy/Snyk
Copilot for DevOps

Job description

Senior DevOps /SRE Engineer
Role Summary

We are looking for a highly skilled and certified Senior DevOps Engineer to join our growing infrastructure team. In this role, you will be responsible for designing, building, and maintaining robust CI/CD pipelines, containerised workloads, cloud infrastructure, and observability platforms that underpin our business-critical financial applications.

The ideal candidate brings deep AWS expertise, strong Linux and scripting skills, hands-on experience with Kubernetes orchestration (EKS/ECS), Infrastructure-as-Code, and a forward-looking mindset toward GenAI-augmented DevOps practices.

Key Responsibilities
1. Cloud Infrastructure & AWS
  • Architect, provision, and manage production-grade AWS environments across multiple accounts and regions.

  • Lead EKS (Elastic Kubernetes Service) cluster management: node group scaling, upgrades, RBAC, pod security policies, and network policies.

  • Design and manage ECS (Elastic Container Service) workloads using Fargate and EC2 launch types.

  • Configure and manage AWS services: VPC, EC2, RDS, S3, IAM, Route53, ALB/NLB, ACM, CloudWatch, SNS, SQS, Lambda, Secrets Manager, Parameter Store, WAF, Shield, and API Gateway.

  • Implement and maintain multi-account AWS organisation strategies with SSO, SCPs, and guardrails.

  • Ensure high availability, disaster recovery (DR), and business continuity for all production workloads.

2. Containerisation & Orchestration
  • Design and maintain Docker images and multi-stage Dockerfiles optimised for security and performance.

  • Define Helm charts for Kubernetes application deployments across dev, staging, and production environments.

  • Implement Kubernetes best practices: resource limits, health probes, HPA/VPA, pod disruption budgets, and rolling deployments.

  • Manage container image repositories and lifecycle policies via ECR.

3. CI/CD & DevOps Toolchain
  • Build, maintain, and improve CI/CD pipelines using Jenkins (declarative & scripted pipelines).

  • Integrate automated testing, code quality gates (SonarQube), vulnerability scanning, and artifact management into pipelines.

  • Manage Git repositories, branching strategies (GitFlow / trunk-based), and code review workflows.

  • Implement blue-green, canary, and rolling deployment strategies to support zero-downtime releases.

  • Collaborate with development teams on shift-left security and DevSecOps practices.

4. Infrastructure as Code (IaC)
  • Own end-to-end Terraform codebases: module design, state management, remote backends (S3 + DynamoDB), and workspace strategies.

  • Implement Terraform CI/CD with plan/apply automation, drift detection, and policy-as-code (OPA/Sentinel or Checkov).

  • Maintain version-controlled IaC with documented runbooks and change management processes.

5. Observability, Monitoring & Alerting
  • Design and manage a unified observability stack using Prometheus, Grafana, and New Relic.

  • Build Grafana dashboards for infrastructure KPIs, application performance, and business metrics.

  • Define alerting rules, SLOs, SLIs, and error budgets aligned with business-critical service SLAs.

  • Implement distributed tracing and log aggregation (ELK / CloudWatch Logs / Loki).

  • Conduct regular capacity planning reviews and performance tuning exercises.

6. API Gateway & Networking
  • Manage AWS API Gateway (REST & HTTP APIs): throttling, WAF integration, authorisers, usage plans, and versioning.

  • Design and maintain secure VPC architecture: subnets, NACLs, security groups, VPC peering, Transit Gateway, and PrivateLink.

  • Implement mTLS, API security policies, and rate-limiting for all external-facing services.

7. Linux & System Administration
  • Deep expertise in Linux (RHEL/CentOS/Amazon Linux/Ubuntu) administration and hardening.

  • Perform kernel tuning, filesystem management, process management, and performance profiling.

  • Implement OS-level security hardening aligned with CIS benchmarks.

  • Manage SSH, PAM, sudoers, and user access controls in production environments.

8. Scripting & Automation
  • Write production-quality Bash scripts for system automation, alerting, log rotation, patching, and health-check routines.

  • Develop Python scripts for AWS automation (boto3), data processing, operational tooling, and custom Prometheus exporters.

  • Build internal CLI tools and runbook automations to reduce MTTR and eliminate toil.

9. GenAI & Emerging Technology
  • Actively explore and integrate GenAI tools into DevOps workflows — AI-assisted code reviews, incident summarisation, runbook generation, and intelligent alerting.

  • Evaluate and prototype LLM-powered internal tools (Copilot for DevOps, ChatOps bots, AIOps capabilities).

  • Stay current on AI/ML infrastructure trends and support data/ML teams with MLOps tooling where applicable.

  • Experiment with prompt engineering for automation use cases and share findings for team knowledge sharing.

10. Security & Compliance
  • Implement secrets management using AWS Secrets Manager, HashiCorp Vault, or equivalent.

  • Conduct regular vulnerability assessments, patching cycles, and security audits of cloud and container environments.

  • Ensure compliance with NBFC regulatory requirements (data residency, audit trails, access logging).

  • Integrate SAST/DAST tools and container image scanning (Trivy/Snyk) into CI/CD pipelines.

11. Incident Management & Reliability
  • Act as the escalation point for production incidents impacting business-critical applications.

  • Lead post-incident reviews (PIRs), root cause analyses (RCAs), and corrective action implementation.

  • Define and maintain runbooks, SOPs, and disaster recovery playbooks.

  • Drive SRE practices including chaos engineering and game days for resilience validation.

12. Collaboration & Leadership
  • Mentor junior and mid-level DevOps engineers; conduct code reviews and enforce quality standards.

  • Partner with development, QA, security, and product teams to align DevOps practices with organisational goals.

  • Present infrastructure roadmaps and cost optimisation proposals to technology leadership.

  • Contribute to vendor evaluations, RFPs, and toolchain decisions.

Required Qualifications & Skills
Technical Skills — Mandatory
  • AWS (Expert Level) — AWS Certification required:

    • AWS Certified Solutions Architect – Associate or Professional

    • AWS Certified DevOps Engineer – Professional (preferred)

    • AWS Certified SysOps Administrator (added advantage)

  • Kubernetes & Containers: EKS, ECS, Docker, Helm (3+ years hands-on)

  • CI/CD: Jenkins (advanced pipeline development), Git (GitFlow, PR workflows)

  • Infrastructure as Code: Terraform (advanced — modules, workspaces, remote state)

  • Monitoring & Observability: Prometheus, Grafana, New Relic

  • API Management: AWS API Gateway (REST & HTTP)

  • Linux: Advanced — RHEL/Amazon Linux/Ubuntu, bash scripting, system hardening

  • Scripting: Bash (advanced), Python (intermediate–advanced with boto3)

  • Networking: VPC design, DNS, load balancing, TLS/SSL, firewall rules

  • Security: IAM, KMS, Secrets Manager, RBAC, CIS hardening, vulnerability scanning

GenAI & Modern Practices
  • Hands-on exposure or experimentation with GenAI tools (GitHub Copilot, ChatGPT API, LangChain)

  • Understanding of LLMOps or MLOps concepts is a strong plus

  • Ability to leverage AI for code generation, documentation, and operational automation

Soft Skills
  • Strong analytical and problem-solving mindset under high-pressure production scenarios

  • Excellent written and verbal communication — ability to translate complex technical concepts for non-technical stakeholders

  • Ownership mentality with a proactive approach to identifying and resolving risks

  • Team player with experience working in Agile/Scrum delivery models

Experience Requirements
  • Minimum 4 years in a dedicated DevOps / SRE / Cloud Infrastructure role

  • Minimum 3 years working on AWS production environments at scale

  • Minimum 2 years with Kubernetes (EKS preferred) in production

  • Experience in BFSI / NBFC / FinTech domain is a strong advantage

  • Experience managing business-critical, 24x7 production applications

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Engineer, DevOps
Senior Staff Engineer, DevOps

Semtech • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Staff Engineer, DevOps
Senior Staff Engineer, DevOps

Sierra Wireless • India

On-site
INR 2,500,000 - 4,000,000
AWS DevOps Engineer
AWS DevOps Engineer

Keka Technologies Private Limited • Hyderabad

On-site
INR 1,800,000 - 2,800,000
DevOps Engineer
DevOps Engineer

NeuralGarage • Bengaluru

On-site
INR 1,200,000 - 2,400,000
AWS DevOps Engineer
AWS DevOps Engineer

DataBeat.io Media • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Vouchagram India • New Delhi

On-site
INR 2,500,000 - 4,500,000
Principal DevOps Engineer
Principal DevOps Engineer

Valor PayTech • Chennai District

On-site
INR 1,500,000 - 2,500,000
AWS DevOps Engineer
AWS DevOps Engineer

DataBeat • Hyderabad

On-site
INR 1,800,000 - 3,000,000
DevOps Engineer
DevOps Engineer

Benchmark IT Solutions • Maharashtra

On-site
INR 900,000 - 1,300,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000