Positions: 3
What Is This Role About?
We are expanding our core infrastructure engineering team responsible for the reliability, security, scalability, and operational efficiency of a regulated fintech platform. This is a hands‑on production engineering role. You will work across AWS infrastructure, networking, databases, CI/CD, observability, security, disaster recovery, and automation. We are looking for engineers who can independently investigate complex problems, make sound technical decisions, automate repetitive work, and take end-to-end ownership of production systems. This is not a ticket‑driven DevOps role.
What You'll Own
- Architect and operate multi‑account AWS environments using AWS Organizations, Control Tower, IAM, and IAM Identity Center.
- Design, build, operate, and optimize production workloads across ECS, EC2, ALB/NLB, RDS PostgreSQL, Lambda, S3, EFS, and related AWS services.
- Design and troubleshoot AWS networking including VPCs, subnets, route tables, NAT, Transit Gateway, security groups, load balancers, DNS, VPC endpoints, and hybrid connectivity.
- Build and govern Infrastructure as Code using Terraform/Terragrunt, including reusable modules, standards, state management, and infrastructure lifecycle management.
- Own and improve GitLab CI/CD pipelines, runners, deployment strategies, environment promotion, rollback mechanisms, and release automation.
- Build reliable container platforms with strong understanding of Docker image optimization, runtime behaviour, resource allocation, scaling, and ECS operations.
- Lead troubleshooting of complex distributed‑system failures across infrastructure, network, application, container, and database layers.
- Operate Amazon RDS for PostgreSQL, including performance analysis, query tuning, indexing, VACUUM/ANALYZE, backup/recovery, replication, capacity planning, and operational troubleshooting.
- Build observability across logs, metrics, traces, dashboards, alerts, and SLOs using platforms such as CloudWatch, Grafana, and OpenTelemetry‑compatible tooling.
- Manage secrets, IAM boundaries, encryption, certificates, KMS, and security controls across infrastructure and applications.
- Design and validate disaster recovery strategies across AWS regions and accounts, including recovery automation and periodic DR exercises.
- Lead production incident investigation, root‑cause analysis, corrective actions, and prevention of recurring failures.
- Identify and drive AWS cost optimization across compute, databases, storage, networking, logging, and data transfer without compromising reliability.
- Participate in architecture reviews, capacity planning, security reviews, production‑readiness reviews, and infrastructure design decisions.
- Replace manual operational processes with reliable automation using Python, Bash, APIs, Infrastructure as Code, and modern DevOps tooling.
- Work closely with developers, QA, security, and operations teams to improve software delivery and production reliability.
- Mentor engineers, review infrastructure changes, and help establish engineering standards across the team.
Must-Have Skills
Candidates are expected to bring strong hands‑on experience in the following areas from day one:
- 8+ years of hands‑on experience in DevOps, SRE, Cloud Infrastructure, or Platform Engineering roles with direct production accountability.
- Deep AWS experience, including VPC architecture, multi‑AZ design, networking, IAM, compute, storage, databases, scaling, and production troubleshooting.
- Strong AWS networking fundamentals including routing, NAT, Transit Gateway, DNS, load balancing, private connectivity, and network troubleshooting.
- Advanced Terraform expertise, including reusable module design and authoring, remote state management, environment separation, and infrastructure lifecycle management.
- Strong container fundamentals including Docker image optimization, runtime troubleshooting, resource tuning, and production container orchestration.
- Strong Linux fundamentals and hands‑on Bash and/or Python automation.
- Strong understanding of CI/CD architecture and production deployment strategies.
- Deep debugging capability across infrastructure, application, network, container, and database layers.
- Hands‑on observability experience across logs, metrics, dashboards, alerting, and production incident investigation.
- Strong understanding of IAM, secrets management, encryption, least privilege, and cloud security fundamentals.
- Cloudflare, CDN, WAF, or reverse‑proxy experience including routing, caching, TLS/SSL, security policies, and troubleshooting.
- Ability to independently investigate unfamiliar production problems using logs, metrics, traces, system behaviour, documentation, and first principles.
Preferred Skills
- Kubernetes production operations and troubleshooting.
- GitOps and declarative infrastructure/application delivery.
- OpenTelemetry, distributed tracing, and modern observability practices.
- Platform engineering, Internal Developer Platforms, self‑service infrastructure, or developer experience initiatives.
- Apache Airflow or similar workflow orchestration platforms.
- Multi‑region AWS disaster recovery architecture and DR automation.
- Advanced PostgreSQL administration and performance tuning.
- Cloud security and compliance experience in regulated environments.
- SOC 2, ISO 27001, CERT‑In, SEBI, or similar regulatory/compliance exposure.
- IAM federation, SSO, and enterprise identity architecture.
- Infrastructure policy‑as‑code and automated compliance controls.
- Software supply‑chain security, container scanning, secrets scanning, and CI/CD security controls.
AI & Automation Expectations
AI‑assisted engineering will increasingly be part of infrastructure operations. We expect engineers to use it responsibly, not blindly. You should be able to:
- Use AI tools to accelerate investigation, scripting, documentation, automation, and operational analysis.
- Understand and validate AI‑generated code, infrastructure changes, SQL, scripts, and recommendations before using them.
- Never execute AI‑generated production changes without understanding their behaviour and potential blast radius.
- Identify opportunities where repetitive operational work can be safely automated.
- Design automation with appropriate validation, auditability, permissions, rollback, and human controls.
- Treat AI as an engineering accelerator, not a substitute for technical understanding.
How We Expect You to Operate
This is not a checkbox role. We expect engineers to:
- Take ownership of problems from detection through permanent resolution.
- Make technical decisions within their area of responsibility without waiting for step‑by‑step instructions.
- Investigate incidents using evidence rather than assumptions.
- Argue with data, not opinions.
- Understand the systems they operate rather than treating infrastructure as a collection of configuration files.
- Automate repetitive work instead of repeatedly performing manual operations.
- Consider reliability, security, performance, operability, and cost when making infrastructure decisions.
- Understand the blast radius before making production changes.
- Challenge existing architecture when there is a technically better approach.
- Document important architectural decisions, operational procedures, and failure scenarios.
- Be comfortable working across boundaries, from VPCs and DNS to CI/CD, containers, databases, observability, and security.
- Leave systems more reliable and easier to operate than they found them.
This Role Is Not a Fit If You
- Have primarily certification‑based or theoretical AWS knowledge without substantial hands‑on production experience.
- Have mostly operated pre‑built pipelines and infrastructure without understanding how they work underneath.
- Avoid production systems or are uncomfortable taking ownership during incidents.
- Depend on step‑by‑step instructions to investigate unfamiliar technical problems.
- Escalate issues without first investigating logs, metrics, configuration, dependencies, and system behaviour.
- Use AI‑generated solutions without understanding and validating what they do.
- Prefer narrowly defined responsibilities and are uncomfortable troubleshooting across infrastructure layers.
- Treat recurring manual operational work as normal instead of looking for ways to automate it.
- Focus only on making a system work without considering reliability, security, scalability, operability, and cost.