Site Reliability Engineer (SRE) Operations

Teciem

Hinoba-an

Hybrid

PHP 990,000 - 1,518,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Teciem in Bangalore, India, invites an experienced Senior Site Reliability Engineer to own the Kondor UP production platform. You will define SLOs, lead on-call, and drive automation to reduce toil.

The role spans observability, incident response, IaC and GitOps, and working closely with InfraOps, AppOps, NetOps, and SecOps to ensure secure, scalable operations in a multi-tenant SaaS environment. Hybrid work model: in-office Bengaluru with remote options; English required; 5+ years of SRE

Qualifications

  • 5+ years in a Site Reliability Engineer or production operations role.
  • Proven experience defining and operating against SLIs, SLOs, and error budgets.
  • Strong automation skills and programming in at least one language (Python/Go/Bash).

Responsibilities

  • Own the Kondor UP production platform reliability, availability, and performance.
  • Lead on-call rotation, incident response, post-mortems, and customer communication.
  • Define observability strategy with metrics, tracing, and log aggregation; maintain tenant-aware dashboards.

Skills

SRE operations
Python/Go/Bash
Prometheus/Grafana
Kubernetes
ArgoCD/GitOps
Terraform/IaC

Education

Engineering degree (CS/IT/Math)

Tools

AWS EKS
Helm
Jaeger/OTel
Loki/OpenSearch
Terraform
Karpenter

Job description

The Work We Do
Teciem designs, builds, and delivers treasury and capital markets software solutions for financial institutions worldwide. We serve banks of every size and geography, offering the right setup for the right need. Our solutions are designed to replace multiple disconnected systems with one complete, front-to-back platform, helping customers to capture trading and business opportunities quickly, clearly and with control. We cover the entire trading lifecycle, ensuring that everything - from execution to position keeping, to risk management – runs smoothly. With decades of experience and one of the largest, most diverse client bases in the industry, we turn deep industry knowledge into software that covers most asset classes, meets complex real-world treasury and capital market's needs, and adapts as markets evolve.

About Kondor UP
Kondor UP is the SaaS edition of Kondor, Teciem's flagship treasury management platform. Operated by TCM (TeCIEM) on AWS, Kondor UP delivers treasury capabilities as a fully managed, always-on service - processing real financial value on behalf of tier-1 banks under strict regulatory requirements. The platform runs on Amazon EKS (Kubernetes), backed by Amazon RDS for SQL Server (Multi-AZ), Amazon MSK (Kafka), Amazon MQ, and a comprehensive observability stack (OpenTelemetry, Prometheus, Grafana, Loki, Jaeger). Changes are delivered through GitOps (ArgoCD) and Terraform, with blue/green deployments and feature flags ensuring zero-downtime operations. Segregation of duties, full audit trails, and always-on availability are not edge cases in this domain - they are the baseline.

Role Summary

As a Senior Site Reliability Engineer (SRE) - SRE Operations, you will be responsible for the reliability, availability, and performance of the Kondor UP production platform. This is an operations-focused engineering role: you own the production system, define and defend SLOs, lead on-call and incident response, and relentlessly drive down toil through automation and engineering. You apply software engineering discipline to operational problems - turning production failures into systemic improvements, building the observability that gives teams real-time insight, and shifting reliability practices left into the delivery lifecycle. The SRE Operations engineer works closely with the InfraOps, AppOps, NetOps, and SecOps teams, and is a senior contributor within the SRE workstream.

Key Responsibilities
  1. Service Reliability & SLO Management
    Define, instrument, and own Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets across critical Kondor UP services (deal capture, risk engine, market data feeds, API gateway). Maintain the customer-facing SLA commitments; monitor error budget burn rates (2h / 24h windows) and trigger reliability work when burn thresholds are breached. Produce monthly SLA packs and tenant-scoped reliability reports for customer delivery. Drive Production Readiness Reviews (PRR) before new features or services go live.
  2. Incident Response & On-Call
    Participate in the 24/5 on-call rotation as a senior responder; lead P1/P2 incident coordination, remediation, and customer communication. Triage and elevate across the stack (application, platform, network, data) with authority and speed. Facilitate and own blameless post-mortems: structured root-cause analysis, clear action items, tracked follow-through to prevent recurrence. Maintain and continuously improve on-call runbooks, escalation paths, and the incident management playbook. Enforce Segregation of Duties (SoD) and fully auditable change records during and after every production incident.
  3. Observability & Alerting
    Own the end-to-end observability strategy for Kondor UP: metrics (Prometheus + Grafana), distributed tracing (OpenTelemetry / Jaeger), log aggregation (Loki + Fluent Bit), synthetic monitoring, complementing AWS CloudWatch. Design and maintain alerting rules that are actionable and low-noise - minimising alert fatigue while ensuring timely detection of degradation. Build and maintain tenant-aware SLO dashboards (golden signals: latency, traffic, errors, saturation per tenant). Instrument Istio service mesh metrics for east-west traffic observability within EKS clusters.
  4. Production Operations & Platform Health
    Operate and maintain containerised workloads on Amazon EKS - pod health, Helm release management, node group lifecycle (Karpenter), HPA/KEDA tuning. Manage day-2 operations for platform data services: Amazon RDS SQL Server (Multi-AZ, PITR, failover testing), Amazon MSK (Kafka broker health, consumer lag), Amazon MQ, MemoryDB (Valkey). Enforce Kyverno admission policies and ensure production environments remain compliant with security and resource standards. Debug complex production issues across containers, microservices, service meshes, and data layers - identify root cause, fix, document, and implement preventive measures.
  5. Toil Reduction & Automation
    Identify, measure, and relentlessly eliminate operational toil through automation - scripting (Python, Bash, Go), self-service tooling, and operational runbook automation. Build and maintain internal operational tooling to enable safe, repeatable, auditable production operations. Set and enforce the team’s toil threshold policy: if toil exceeds a defined percentage of engineering time, reliability work takes priority over feature delivery.
  6. Infrastructure as Code & GitOps
    Apply GitOps principles (ArgoCD) for all application and configuration changes in production - no manual changes, every state transition is a git commit. Contribute to and review Terraform modules (IaC) for platform infrastructure - ensuring all changes are version-controlled, peer-reviewed, and auditable. Validate and gate production deployments: blue/green readiness, canary analysis, smoke tests, rollback criteria.
  7. Capacity Planning & FinOps
    Conduct capacity planning - model workload growth per tenant, set resource headroom policy, and validate auto-scaling behaviour (Karpenter, HPA) under load. Run regular load and performance tests; validate SLO compliance under projected peak traffic.
  8. Production Readiness & Resilience
    Conduct failure-mode analysis and disaster recovery testing (RDS PITR restoration, cross-region failover, AZ failure simulation). Design and run chaos experiments (LitmusChaos or equivalent) to validate resilience assumptions and improve system robustness. Enforce security guardrails in production: secrets management (AWS Secrets Manager / KMS), CVE triage for running workloads, network policy compliance. Collaborate with the SecOps team on incident response procedures involving security events in production.
Required Skills & Experience
SRE & Production Operations (mandatory)
  • 5+ years in a Site Reliability Engineer or production operations engineering role, operating a SaaS or always-on service at scale.
  • Proven hands-on experience defining and operating against SLIs, SLOs, error budgets, and customer SLAs.
  • Demonstrated incident management experience: 24/5 on-call, structured root-cause analysis, blameless post-mortems.
  • Strong software engineering fundamentals and proficiency in at least one scripting/programming language (Python, Go, or Bash) for automation and operational tooling.
AWS & Kubernetes (mandatory)
  • Hands-on experience operating production workloads on AWS: Amazon EKS, EC2, IAM, VPC, S3, RDS, Route 53, CloudWatch.
  • Solid experience with Kubernetes and Helm for orchestration of containerised workloads in production.
  • Working knowledge of Kubernetes networking (CNI, CoreDNS, NetworkPolicy) and Linux/Unix operating systems.
Observability
  • Proficiency with Prometheus and Grafana (dashboards, recording rules, alerting).
  • Experience with OpenTelemetry, distributed tracing (Jaeger or Tempo), and log aggregation (Loki or OpenSearch/ELK).
  • Ability to define and instrument meaningful SLIs from application and infrastructure telemetry.
IaC & GitOps
  • Experience with Terraform (modules, remote state, AWS provider) for infrastructure changes in production.
  • Familiarity with ArgoCD or equivalent GitOps tooling for continuous delivery and drift detection.
Security & Compliance
  • Knowledge of best practices for data encryption (KMS, TLS/mTLS), secrets management, and least-privilege IAM in production.
  • Awareness of audit and compliance requirements for financial services (SoD, change records, DORA, ISO 27001).
Nice to Have
  • Experience operating a service mesh (Istio) for traffic management, mTLS, and fine-grained observability.
  • Experience with Kyverno or OPA for admission control and compliance guardrails in Kubernetes.
  • Experience with HashiCorp Vault or AWS Secrets Manager / KMS for secrets management at scale.
  • Familiarity with FinOps tooling and cloud cost optimisation for a multi-tenant SaaS platform.
  • Experience with chaos engineering and disaster recovery practices (LitmusChaos, Gremlin, GameDay exercises).
  • Exposure to treasury or capital markets systems (Kondor, Summit, or equivalent TMS/risk platforms).
Experience
  • AWS certifications: Solutions Architect Professional, DevOps Engineer Professional, or SysOps Administrator.
  • Kubernetes certifications: CKA (Certified Kubernetes Administrator) or CKAD.
Profile Experience

5+ years SRE or production operations engineering; 3+ years on AWS in a SaaS or always-on context.

Education

Engineering degree (Computer Science, Information Technology, Mathematics, or equivalent) or equivalent experience.

Languages

English (mandatory - working language)

Location

Bangalore, India

Work model

Hybrid - Bangalore (TeCIEM India office) + remote

On-call Participation in a 24/5 on-call rotation is a core requirement of this role

Team Context & Reporting

The SRE Operations Engineer (TP3) reports to the Cloud SRE Manager and operates within the SRE workstream.

Primary interfaces:
- InfraOps - joint ownership of EKS cluster health, node lifecycle, platform baseline; escalation path for infrastructure-layer incidents.
- AppOps - first-line production support; SRE provides reliability tooling, runbooks, and post-mortem leadership.
- NetOps - network-layer triage during incidents; collaboration on SLI instrumentation for connectivity paths.
- SecOps - production security posture, CVE triage, incident response coordination, compliance evidence.
- DevOps / Platform Engineering - reliability gates in CI/CD pipelines; PRR process; shared ownership of GitOps toolchain.
- Architecture - resilience pattern review, capacity planning inputs, ADR contributions.

Diverse Minds, Shared Ambition

At Teciem, we believe that our strength comes from the diversity of our people. Different perspectives, backgrounds, and experiences fuel our innovation and help us build solutions that truly make a difference in the world of financial technology. We’re committed to creating a workplace where everyone feels respected, heard, and empowered to grow. Here, you can bring your whole self to work, contribute your unique ideas, and be part of a team driven by shared ambition. We welcome talent from all walks of life and encourage applications from individuals of all genders, races, ages, abilities, identities, and beliefs. Together, we’re shaping a culture where diversity isn’t just celebrated — it’s essential to our success.

Purpose – Why we exist

We empower financial institutions to build resilient and future-ready economies, worldwide.

Vision – What the future holds

To lead innovation in treasury and capital markets technology, building on the solid foundations of our mission -critical and industry - defining solutions.

Mission – How we get there

We place our clients' success and ambitions at our core, continuously evolving and innovating our solutions to deliver outstanding business value and real economic impact.

You help us simplify Treasury and capital markets can be intricate, but you play a key role in making them easier to navigate. Your ideas and expertise help us transform complicated processes into intuitive, streamlined solutions used by financial institutions worldwide. Every improvement you make creates clarity, efficiency, and real-world impact for our clients. You shape the future with AI, every day AI isn’t a buzzword here — it’s embedded in how we build, innovate, and deliver. Whether you’re working on smarter automation, data-driven insights, or enhanced user experiences, your contribution fuels the next generation of intelligent financial technology. You’ll be part of a team that uses AI to make our products faster, sharper, and more meaningful for the industry. You grow through collaboration. We believe the best outcomes happen when great minds come together. You’ll work alongside talented colleagues across engineering, product, design, and client-facing teams — sharing knowledge, solving problems, and learning constantly. Collaboration isn’t just how we work; it’s how we grow, innovate, and support each other.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) Operations
Senior Site Reliability Engineer (SRE) Operations

Teciem • Hinoba-an

Hybrid
PHP 2,772,000 - 4,291,000
Senior Run DevOps Engineer
Senior Run DevOps Engineer

Teciem • Hinoba-an

On-site
PHP 1,200,000 - 2,000,000
Network Operations Engineer (NetOps)
Network Operations Engineer (NetOps)

Teciem • Hinoba-an

On-site
PHP 1,200,000 - 1,900,000
Team Lead/Manager – Kondor Back Office
Team Lead/Manager – Kondor Back Office

Teciem • Hinoba-an

On-site
PHP 924,000 - 1,386,000
Senior SRE — FinTech SaaS Operations (Hybrid)
Senior SRE — FinTech SaaS Operations (Hybrid)

Teciem • Hinoba-an

Hybrid
PHP 990,000 - 1,518,000
Senior SRE Operations Engineer - Cloud & Observability
Senior SRE Operations Engineer - Cloud & Observability

Teciem • Hinoba-an

Hybrid
PHP 2,772,000 - 4,291,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Senior AI Engineer
Senior AI Engineer

Teciem • Hinoba-an

On-site
PHP 1,650,000 - 2,640,000
Accounting Operations Lead
Accounting Operations Lead

Teciem • Hinoba-an

On-site
PHP 792,000 - 1,188,000
AI Engineering Lead
AI Engineering Lead

Teciem • Hinoba-an

Hybrid
PHP 1,980,000 - 3,301,000