Lead Site Reliability Engineer

Trintech Solutions Private Limited (India)

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Trintech Solutions Private Limited (India) is seeking a Lead Site Reliability Engineer to drive reliability across a broad service portfolio. You will own the integrated reliability plan, coordinate work across teams and ensure incident learning translates into measurable improvements.

This hands-on IC role will focus on Kubernetes, OKD/OpenShift, Azure, VMware infrastructure, and Linux/Windows workloads, with Grafana, Prometheus, Datadog, Sumo Logic and PagerDuty supporting visibility and

Qualifications

  • Proven track record leading reliability improvements across production services.
  • Experience with monitoring, alerting and incident lifecycle management.
  • Hands-on with Kubernetes-based environments and hybrid infra.

Responsibilities

  • Build a prioritised reliability roadmap from incident data and customer impact.
  • Lead adoption of service indicators and error-budget decisions with stakeholders.
  • Coordinate alert improvements across multiple monitoring tools and runbooks.
  • Lead complex incident response and post-incident reviews to drive lasting improvements.
  • Develop and review automation to reduce repetitive operational work.
  • Coordinate readiness, capacity and recovery work across teams.

Skills

SRE engineering
Incident management
Cross-team coordination

Tools

Kubernetes
OKD/OpenShift
Azure
VMware
Linux
Windows
Grafana
Prometheus
Datadog
Sumo Logic
PagerDuty
Terraform
Argo CD

Job description

Description About the role At Trintech, the Lead Site Reliability Engineer leads the technical delivery of reliability improvements across a broad service portfolio. You will turn production evidence into a prioritised engineering plan, coordinate work across teams and ensure incident learning results in measurable improvement. This is a technical individual-contributor role with no direct reports. You will remain hands-on across Kubernetes and OKD/OpenShift, Azure, VMware-based infrastructure, and Linux and Windows workloads. Grafana, Prometheus, Datadog, Sumo Logic and PagerDuty support investigation, service visibility and incident response.

Your impact
  • Own the integrated reliability engineering plan across services, regions and workstreams, with clear technical priorities and accountable follow-through.
  • Strengthen the SRE team's effectiveness through practical direction, better operational tools, sustainable response practices and verified engineering outcomes.
What you will do
  • Build a prioritised reliability roadmap from customer impact, incident recurrence, service risk and operational effort, agreeing delivery capacity with the SRE manager.
  • Lead adoption of meaningful service indicators, objectives and error-budget decisions with product and service owners, establishing ownership and review practices.
  • Coordinate alert improvements across Grafana, Prometheus, Datadog, Sumo Logic and PagerDuty, addressing duplication, missing coverage, routing and diagnostic context.
  • Lead technical response to complex incidents and strengthen regional handovers, escalation readiness and blameless reviews; verify completion of consequential follow-ups.
  • Direct and contribute to automation that removes repeated operational work, with code review, testing, bounded permissions and safe recovery behaviour.
  • Coordinate readiness, capacity and recovery work across application, data, infrastructure and Platform Engineering teams, resolving dependencies and technical blockers.
  • Review high-risk changes and remain hands-on through diagnosis, prototypes and targeted implementation where direct involvement improves the outcome.
  • Report customer impact, reliability trends, repeat failures, alert quality and engineering progress; raise gaps in capacity or ownership with clear recommendations.
What you will bring
  • A record of leading reliability improvements across multiple production services and teams, with evidence of sustained outcomes after delivery.
  • Strong practical judgement across Kubernetes, Linux, application behaviour, networks, storage and hybrid infrastructure, with the ability to involve specialists effectively.
  • Hands-on depth in monitoring, PromQL, log analysis and alert design, including the ability to distinguish telemetry faults from service failures.
  • Experience leading significant incidents, improving on-call practices and driving corrective actions through to verified resolution.
  • Credible software and automation skills, including reviewing and implementing maintainable, tested operational tooling and controlled changes.
  • Ability to coordinate dependencies, mentor technical leaders and negotiate delivery priorities and reliability trade-offs without relying on line-management authority.
Nice to have
  • Experience with OKD/OpenShift, Azure, VMware, Windows, Datadog, Sumo Logic, PagerDuty, Terraform, SaltStack, Azure DevOps or Argo CD.
  • Experience supporting distributed teams and SaaS services with database, integration, batch-processing or customer-facing financial-workflow dependencies.
How you will work and lead
  • Set technical priorities and guide Senior and Principal contributors, agreeing resource commitments with their managers and retaining clear service ownership.
  • Partner with the SRE manager on on-call readiness and protected engineering capacity; provide coaching and technical input into development and recruitment.
  • Make delivery and recovery decisions within agreed authority, escalating business-risk acceptance, funding and organisation-wide architecture choices to accountable owners.
What success looks like
  • The reliability roadmap delivers measurable reductions in customer disruption, recurring incidents and avoidable operational effort.
  • Critical services have agreed owners, meaningful health measures, actionable alerts and tested recovery procedures.
  • Regional handovers, escalation and incident follow-ups work consistently, with response load and capacity constraints visible to management.
  • Teams deliver coordinated improvements with verified outcomes while engineers develop stronger independent technical judgement.

At our core, Trintechers stand committed to fostering a culture rooted in our core values – Humble, Empowered, Reliable, and Open. Together, these values guide our actions, define our identity, and inspire us to continuously strive for excellence in everything we do.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin or disability.

The Growth is Creating Great Opportunities! Our team is expanding, and we want to hire the most talented people we can. Continued success depends on it!

Thanks for your interest in working on our team! Trintech gives people time back for what matters most. Our AI Financial Close solutions enable thousands of clients worldwide to lead productivity transformation across their finance and accounting organizations — driving efficiencies, ensuring accuracy to mitigate risk, and empowering strategic decision-making. Make time count with Trintech. As the leader in AI Financial Close Management, Trintech is headquartered in Plano, Texas with offices and strategic resellers across United States, Europe, Australia, South America, Africa, and Asia Pacific. With a strong partner ecosystem, Trintech collaborates with over 100 companies to create a network of interconnected businesses. To learn more about Trintech, visit www.trintech.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Trintech • Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Software Eng Tech Lead
Software Eng Tech Lead

Trintech Solutions Private Limited (India) • Bengaluru

On-site
INR 3,800,000 - 7,000,000
Principal Platform Engineer
Principal Platform Engineer

Trintech Solutions Private Limited (India) • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Software Engineering Technical Lead, Bangalore
Software Engineering Technical Lead, Bangalore

Trintech Solutions Private Limited (India) • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior Software Test Engineer (ACCT)
Senior Software Test Engineer (ACCT)

Trintech Solutions Private Limited (India) • Dadri

On-site
INR 1,200,000 - 1,800,000
Senior Software Test Engineer (ACCT)
Senior Software Test Engineer (ACCT)

Trintech • India

On-site
INR 1,200,000 - 1,800,000
Senior Software Test Engineer (FR)
Senior Software Test Engineer (FR)

Trintech • India

On-site
INR 1,200,000 - 2,400,000
Senior Software Test Engineer (FR)
Senior Software Test Engineer (FR)

Trintech Solutions Private Limited (India) • Dadri

On-site
INR 1,400,000 - 2,000,000
Principal Software Test Engineer
Principal Software Test Engineer

Trintech Solutions Private Limited (India) • Dadri

On-site
INR 1,200,000 - 2,100,000
Senior Software Engineer (Smalltalk)
Senior Software Engineer (Smalltalk)

Trintech Solutions Private Limited (India) • Bengaluru

On-site
INR 1,500,000 - 2,300,000