Senior Software Engineer (Site Reliability Engineering)

SentiLink

Bengaluru

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Employer paid group health insurance
401(k) plan with employer match
Flexible paid time off
Regular company-wide in-person events
Home office stipend

Job summary

SentiLink is seeking a Software Engineer – SRE in Bengaluru, India, to enhance system reliability and performance. You’ll tackle production challenges, collaborate with backend teams, and improve observability through robust coding and tooling.

The ideal candidate has over 7 years of software development experience with strong coding skills in Python or Go, and a solid understanding of PostgreSQL and distributed systems. The position includes remote collaboration with US teams and comes with great perks.

Qualifications

  • 7–8+ years of software development experience in backend or SRE roles.
  • Strong coding skills in Python, Go, or similar.
  • Deep understanding of PostgreSQL, performance tuning, and query optimization.
  • Experience with distributed systems, containers, and Kubernetes in production.

Responsibilities

  • Design, write, and ship production-grade code to improve performance.
  • Identify and resolve application issues across the stack.
  • Collaborate with US-based engineering teams on incidents and design reviews.
  • Contribute to logs, metrics, dashboards, and alerts for observability.

Skills

Python
Go
PostgreSQL
Docker
Kubernetes
Datadog
CloudWatch
Prometheus
Infrastructure as Code (IaC)

Tools

Terraform

Job description

Software Engineer – SRE

Are you a passionate software engineer who thrives on solving tough production problems, writing robust code, and making systems fast, reliable, and observable? We're looking for a Software Engineer to join our SRE team, where you'll work at the intersection of software and infrastructure, improving the reliability and performance of our most critical applications. This is a software engineering role at its core. You'll spend your time writing production code, profiling services, and debugging across the stack, with the leverage of knowing your work improves every service, not just one. You'll collaborate closely with backend and product teams to fix performance bottlenecks, chase down complex issues, and build tooling that supports better observability, cost‑efficiency, and operational excellence. You'll also help improve our incident response processes, evolve our monitoring strategy, and contribute directly to high‑impact production systems.

Responsibilities
  • Design, write, and ship production‑grade code to fix bugs, improve performance, and increase reliability across multiple services.
  • Tackle complex coding challenges in live services — requiring solid understanding of algorithms, data structures, and system architecture.
  • Identify and resolve application issues by diving deep across the stack — including backend code, database interactions, and infrastructure components — with the ability to implement code‑level fixes where needed.
  • Profile APIs and services under load to identify bottlenecks and implement fixes at the code, database, or configuration level.
  • Design and evaluate infrastructure solutions — weighing tradeoffs in architecture, tooling choices, and system configuration — and clearly document rationale for decisions made.
  • Collaborate effectively with US‑based engineering teams, including availability for overlap hours in the late evening IST to support real‑time coordination on incidents, design reviews, and cross‑functional initiatives.
  • Communicate clearly and proactively — write structured updates, flag blockers early, and synthesize technical context for both engineering and non‑engineering audiences.
  • Build tools, automation, and test frameworks to simulate traffic, validate behavior under stress, and prevent regressions.
  • Improve observability by contributing to logs, metrics, dashboards, alerts, and distributed tracing.
  • Collaborate on defining and measuring SLIs, SLOs, and SLAs to align reliability goals with business outcomes.
  • Contribute to the development and maintenance of incident response runbooks and help improve operational processes to minimize downtime.
  • Participate in incident response and root cause analysis and contribute to the engineering on‑call rotation.
  • Support cost optimization efforts.
  • Research and evaluate new monitoring technologies or best practices to continuously improve system visibility and reliability.
Requirements
  • 7–8+ years of software development experience, with meaningful time in backend or SRE/infrastructure‑adjacent roles.
  • Strong coding skills in languages such as Python, Go, or similar.
  • Proven ability to write and ship production code — and debug it under real‑world conditions.
  • Deep understanding of PostgreSQL or similar relational databases, including performance tuning and query optimization.
  • Hands‑on experience with distributed systems, containers (Docker), and Kubernetes in production — including making and defending architectural decisions, not just operating existing setups.
  • Experience with observability and monitoring platforms like Datadog, CloudWatch, or Prometheus.
  • Strong written and verbal communication skills — able to write clear technical proposals, postmortems, and async updates for distributed teams. Comfortable presenting tradeoffs and decisions to senior stakeholders.
  • Comfortable working with US‑based counterparts, including regular overlap hours in the late evening IST.
  • Demonstrates a structured execution approach — breaks down ambiguous problems, tracks progress against milestones, and surfaces blockers before they become delays.
  • Knowledge of SLIs, SLOs, and SLAs, and how they align with business objectives.
  • Familiarity with cloud platforms (AWS preferred) and distributed systems.
  • Comfortable working across unfamiliar codebases and resolving issues from code to infrastructure.
  • Strong analytical, problem‑solving, and collaboration skills.
  • Prior experience in an IaC (Terraform/Terragrunt) or performance engineering role is a plus.
Perks
  • Employer paid group health insurance for you and your dependents
  • 401(k) plan with employer match (or equivalent for non US‑based roles)
  • Flexible paid time off
  • Regular company‑wide in‑person events
  • Home office stipend, and more!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE
Senior SRE

CloudRaft • India

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Premium health insurance & wellness
AI stack & GPU infrastructure
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer
Site Reliability Engineer

Ontrac Solutions • Delhi

On-site
INR 1,200,000 - 2,200,000
Lead Software Engineer - Site Reliability
Lead Software Engineer - Site Reliability

jobr.pro • Chennai District

On-site
INR 3,000,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

NOV • Ernakulam

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

MetaForgeIT • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Persistent Systems • Pune District

Hybrid
INR 1,200,000 - 2,200,000
Hybrid work
Career growth
Education sponsorship
+4
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000