Site Reliability Engineer IV

Capitolis

Sterling (VA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Candescent is seeking a Site Reliability Engineer IV based in Sterling, Virginia. The ideal candidate will have 9-12 years of experience, focusing on reliability, availability, and performance for digital banking applications. Responsibilities include supporting production applications on Kubernetes and AWS, troubleshooting issues, and improving observability. Candidates must have hands-on experience with Java applications, a strong understanding of SRE practices, and excellent collaboration skills. This role is a key part of the engineering team to enhance operational performance and reliability.

Qualifications

  • 9-12 years of experience in a relevant role.
  • Proven experience with SRE practices.
  • Strong experience troubleshooting using application logs and metrics.

Responsibilities

  • Support and operate production applications running on Kubernetes and AWS.
  • Participate in incident response and root cause analysis.
  • Drive continuous improvements in operational readiness.

Skills

Hands-on experience supporting Java applications in production
Strong understanding of JVM fundamentals
Experience operating Java applications on Kubernetes (EKS)
Ability to write automation and scripts (Python or any)
Strong collaboration and communication skills

Tools

AWS services (EKS, RDS/Aurora, S3, EFS, CloudWatch)
Monitoring tools

Job description

Candescent is the leading cloud-based digital banking solutions provider for financial institutions. We are transforming digital banking with intelligent, cloud-powered solutions that connect account opening, digital banking, and branch experiences for financial institutions. Our advanced technology and developer tools enable seamless, differentiated customer journeys that elevate trust, service, and innovation. Success here requires flexibility in a fast-paced environment, a client-first mindset, and a commitment to delivering consistent, reliable results as part of a performance-driven, values-led team. With team members around the world, Candescent is an equal opportunity employer.

Position: Site Reliability Engineer IV

Experience: 9-12 Years

Location: Bangalore (Ecospace)

Candescent Site Reliability Engineering (SRE) mission is to proactively ensure the reliability, availability and performance of our Digital First banking applications. As a member of the SRE team, you will focus on building and operating highly reliable application platforms by applying SRE principles such as automation, observability, resilience and continuous improvement.

You will partner closely with application and platform teams to define reliability standards, implement monitoring, alerting and incident response practices and embed scalability and performance considerations into application design and delivery. Through tooling, automation, and best practices, you will help development teams build and operate services that meet agreed reliability objectives.

As a senior engineer in the organization, you will also provide mentorship within the SRE team and across peer engineering teams, helping elevate operational maturity, drive adoption of SRE practices, and strengthen reliability culture across our core initiatives.

Responsibilities
  • Support and operate production applications running on Kubernetes and AWS
  • Troubleshoot application-level issues using logs, metrics, traces, and runtime signals
  • Participate in incident response, root cause analysis, and post-incident reviews
  • Work closely with development teams to understand application architecture, dependencies, and data flows
  • Improve application observability by defining meaningful alerts, dashboards, and SLOs
  • Automate repetitive operational tasks to reduce toil
  • Support application deployments, rollbacks, and runtime configuration changes
  • Identify reliability, performance, and scalability gaps in application behavior
  • Drive continuous improvements in operational readiness, runbooks, and on-call practices
  • Influence application teams to adopt shift-left reliability practices
Must-Have Skills & Experience
  • Hands-on experience supporting Java applications in production
  • Strong understanding of JVM fundamentals (heap/memory management, garbage collection, OOM issues, thread analysis)
  • Proven experience with SRE practices, including:
    • Incident response and on-call support
    • Root cause analysis and postmortems
    • SLIs, SLOs, and reliability-driven operations
  • Strong experience troubleshooting using application logs, metrics, and monitoring tools
  • Experience operating Java applications on Kubernetes (EKS) from an application/runtime perspective
  • Experience with deployment strategies (rolling, blue/green, canary)
  • Ability to write automation and scripts (Python or any) to reduce operational toil
  • Solid understanding of application architecture and service dependencies (databases, messaging systems, external APIs)
  • Strong collaboration and communication skills; ability to work closely with development teams
  • Demonstrates accountability and sound judgment when responding to high-pressure incidents
Good-to-Have Skills & Experience
  • Exposure to platform or infrastructure concepts supporting application workloads
  • Experience with AWS services such as EKS, RDS/Aurora, S3, EFS, and CloudWatch
  • CI/CD pipeline experience (GitHub Actions, Jenkins)
  • Familiarity with GitOps practices
  • Experience with cloud migrations or modernization efforts
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE - Cloud Reliability & Observability
Senior SRE - Cloud Reliability & Observability

Capitolis • Sterling (VA)

On-site
USD 120,000 - 160,000
Staff Engineer
Staff Engineer

Capitolis • Sterling (VA)

Remote
USD 120,000 - 150,000
Software Engineer IV - (C/C++)
Software Engineer IV - (C/C++)

Capitolis • Sterling (VA)

On-site
USD 18,000 - 30,000
software Engineer III- C/C++
software Engineer III- C/C++

Capitolis • Sterling (VA)

On-site
USD 100,000 - 120,000
Competitive compensation
Strong work-life balance programs
Inclusive, diverse culture
Staff SW Engineer - Integration
Staff SW Engineer - Integration

Capitolis • Sterling (VA)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability engineering (SRE)
Site Reliability engineering (SRE)

TechDigital Group • San Leandro (CA)

On-site
USD 100,000 - 150,000
SW Staff Engineer - Java
SW Staff Engineer - Java

Capitolis • Sterling (VA)

On-site
USD 2,000 - 6,000
Software Staff Engineer
Software Staff Engineer

Capitolis • Sterling (VA)

On-site
USD 120,000 - 150,000
Competitive compensation
Growth opportunities
Transparent work culture
Staff Engineer - Banking Integrations Platform
Staff Engineer - Banking Integrations Platform

Capitolis • Sterling (VA)

On-site