Senior Software Engineer - Sre Focused

Alpheya

Bengaluru

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A leading WealthTech startup in Bengaluru is seeking a seasoned Site Reliability Engineer to spearhead production incident response. The role demands expertise in Go and Kubernetes, with a hands-on approach to debugging and fixing issues in production. The ideal candidate will have over 7 years in reliability-focused positions, demonstrating a commitment to engineering hygiene and operational standards. This is an opportunity to work in a dynamic fintech environment, ensuring robust performance across their digital wealth management platform.

Qualifications

  • 7+ years of experience in SRE or Production Engineering focused on reliability.
  • Proven debugging skills in distributed systems and production environments.
  • Strong hands-on experience with Kubernetes and Go programming.

Responsibilities

  • Lead production incident response and communication.
  • Debug and fix issues across Go services and the broader system.
  • Troubleshoot Kubernetes workloads and deployments.

Skills

Production Engineering
Go programming
Kubernetes
Incident leadership
Observability

Tools

PostgreSQL
Grafana
Snowflake

Job description

We are a B2B WealthTech startup based in Abu Dhabi and backed by BNY Mellon (America’s oldest bank and first company to list on NYSE) and Lunate (a new $50B AUM alternative asset management firm based in Abu Dhabi, UAE). The company has raised $300M to build a state of the art wealth technology platform.

Our mission is to power and grow our clients’ Wealth franchises through differentiated experiences, financial solutions, and insights. Our digital wealth management platform- will enable banks and other financial institutions in the Middle East to grow and further penetrate affluent, HNW and UHNW investor segments.

While still leveraging the capabilities and knowledge of large organizations, our fintech is a startup with truly cross-functional and agile teams.

For more information, please visit www.alpheya.com

Role

We're building a team that owns production incident response, deep debugging, and permanent fixes across application, data, and deployment layers.

This is not a tickets-only ops role. You will write code, ship fixes safely, and harden the platform so issues don't repeat.

Note: This is a software engineering role with real production ownership. You’ll combine engineering and operations to own outcomes end-to-end: investigate incidents, ship code fixes, and prevent repeat issues through tests, observability, and hardening.

  • Lead and execute production incident response: triage, mitigation, stakeholder communication, and coordination across teams
  • Debug and fix issues across Go services (mandatory) and the broader stack (Node.js services where relevant)
  • \
  • Work across service boundaries: GraphQL/RPC, distributed tracing, dependency failures, performance bottlenecks, and safe degradation patterns
  • Troubleshoot Kubernetes workloads and deployments
  • Diagnose PostgreSQL/CNPG issues
  • Handle production bugs that span application + data pipelines (ETL/Snowflake mappings), including backfills/replays and data-quality validation
  • Build prevention: add regression tests, improve observability , and maintain runbooks/service passports
  • Drive reliability improvements: SLOs/SLIs, alert quality, release readiness checks, and operational standards across teams
  • 7+ years in SRE / Production Engineering / Platform Engineering (reliability-focused)
  • Strong Go (mandatory): ability to read, debug, and ship production fixes in Go codebases
  • Proven experience debugging distributed systems in production (latency, error rates, timeouts, retries, cascading failures)
  • Strong hands-on experience with Kubernetes in production environments
  • Experience with Helm and GitOps workflows (FluxCD preferred; ArgoCD acceptable)
  • Solid PostgreSQL troubleshooting experience (performance, incident patterns, migrations)
  • Observability experience (metrics/logging/tracing; Datadog/Grafana/Tempo/Loki experience is a plus)
  • Strong incident leadership: calm under pressure, clear communication, structured problem-solving
  • Engineering hygiene: PR discipline, reviews, testing mindset, safe rollouts/rollbacks
  • Comfortable with IAM/security fundamentals in real production systems: OAuth2/OIDC basics, RBAC/least privilege, and safe secrets handling
Good to Have
  • Node.js backend experience in production
  • Experience in FinTech / regulated environments / high-availability systems (auditability, change control, incident rigor)
  • Data reliability experience: ETL monitoring, reconciliation, Snowflake operations, schema/mapping drift handling
  • Reliability patterns common to trading/fintech platforms: correctness and data integrity mindset (idempotency, reconciliation), resilient partner integrations, and strong observability for critical user journeys
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Data & Application Reliability
Senior Site Reliability Engineer - Data & Application Reliability

Alpheya • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BayOne Solutions • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AssetMark Global Wealth • Hyderabad

On-site
INR 1,400,000 - 2,100,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

AlphaSense Oy • Pune District

On-site
INR 4,000,000 - 7,000,000
Full Stack Engineer
Full Stack Engineer

BPMLinks • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Production Support Engineer
Production Support Engineer

Luma Financial Technologies, LLC • India

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

BuildxPartners • Sadar Bazar

On-site
INR 3,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Full Stack Engineer
Full Stack Engineer

Algoworks • India

On-site
INR 1,500,000 - 2,000,000