Lead Site Reliability Engineer

SproutsAI

Bengaluru

On-site

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SproutsAI is seeking a Senior Application SRE to join our expanding Site Reliability Engineering team in Bengaluru, India. This role focuses on enhancing the reliability and performance of our trading environment.

The ideal candidate should have a strong background in Linux and Windows system administration, along with expertise in Docker, Kubernetes, and monitoring tools like Prometheus and Grafana. Join us to help shape our SRE strategy and contribute to mission-critical systems in a collaborative environment.

Qualifications

  • 5+ years of experience in application reliability and automation.
  • Proficiency in at least one scripting language.
  • Hands-on experience with various monitoring and observability tools.

Responsibilities

  • Ensure application reliability through monitoring and automation.
  • Develop real-time monitoring and alerting systems.
  • Automate application deployment and configuration.

Skills

Linux system administration
Windows system administration
Scripting (Python, Shell, etc.)
Docker
Kubernetes
Monitoring tools (Prometheus, Grafana, etc.)
Networking (TCP, IP, DNS)
Configuration management (Ansible, Puppet, etc.)
Debugging skills
Collaboration skills

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Kubernetes
Jenkins
Prometheus
Grafana
ELK Stack
Datadog

Job description

About the Role

Integral is committed to delivering best-in-class service reliability and performance. As part of this commitment, we are expanding our Site Reliability Engineering (SRE) team to ensure the reliability, performance, and availability of our software applications. We are looking for a highly motivated and technically talented Senior Application SRE to support our 24x7 FX trading environment. This role will focus on application monitoring, automation, and optimization to enhance system stability, minimize downtime, and improve overall user experience. The ideal candidate will bring strong problem-solving skills, experience in large-scale distributed systems, and a deep understanding of software and infrastructure reliability principles.

Responsibilities
  • Ensure the reliability, performance, and availability of Integral’s applications through proactive monitoring and automation.
  • Develop and maintain real-time monitoring, alerting, and logging systems to detect and resolve issues before they impact customers.
  • Automate manual operations, including application deployment, configuration, scaling, and recovery.
  • Collaborate with software engineering teams to integrate reliability best practices into the development lifecycle.
  • Conduct root cause analysis (RCA) and implement preventive measures to mitigate recurring issues.
  • Support a 24x7 distributed enterprise environment across multiple global data centers.
  • Work closely with Support to enhance incident response processes, ensuring fast and effective resolution of technical escalations.
  • Participate in on-call rotations to support critical application issues and outages.
  • Maintain and optimize CI,CD pipelines to ensure fast and reliable application releases.
  • Enhance system security by managing SSL certificates, encryption, and authentication mechanisms.
  • Foster a culture of continuous improvement by evaluating new tools, frameworks, and methodologies to enhance system reliability.
Requirements
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 5+ years of experience in a similar role, focusing on application reliability, automation, and performance optimization.
  • Strong expertise in Linux and Windows system administration.
  • Proficiency in at least one scripting language (e.g., Python, Shell, Perl, JavaScript).
  • Experience with Docker, Kubernetes, or containerization technologies.
  • Familiarity with CI,CD tools like Jenkins and deployment automation frameworks.
  • Hands‑on experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack, New Relic, Datadog).
  • Understanding of networking concepts (TCP, IP, DNS, load balancing, firewalls).
  • Experience with configuration management tools like Ansible, Salt, or Puppet.
  • Strong debugging and troubleshooting skills across application, database, and infrastructure layers.
  • Ability to work in a fast‑paced, high‑pressure environment with multiple priorities.
  • Excellent communication and collaboration skills to work effectively with engineering and support teams.
Nice-to-Have Skills
  • Experience in the financial services or trading industry.
  • Knowledge of distributed computing, cloud platforms (AWS, GCP, Azure).
  • Exposure to security best practices and compliance standards.
  • Familiarity with incident management frameworks (ITIL, SRE best practices, or similar methodologies).
Why Join Us?

Be a key player in shaping Integral’s SRE strategy and improving mission‑critical trading systems. Work in a collaborative, fast‑paced environment with top engineering talent. Enjoy career growth opportunities in an organization that values technical excellence and innovation. Competitive compensation and benefits package. If you're passionate about site reliability, automation, and scaling highly available applications, we'd love to hear from you! Apply now and help us build the future of reliable trading technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

FIS • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Competitive salary
Attractive range of benefits
Opportunity for skill growth
Site Reliability Engineer
Site Reliability Engineer

Bajaj Broking • Maharashtra

On-site
INR 400,000 - 600,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

FIS • Pune District

On-site
INR 3,500,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Bajaj Broking • Pune District

On-site
INR 4,000,000 - 7,000,000
Equity or performance-based bonus (if/
Lead Software Engineer - Site Reliability
Lead Software Engineer - Site Reliability

jobr.pro • Chennai District

On-site
INR 3,000,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Karnataka

On-site
INR 800,000 - 1,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site reliability engineer (Java,Unix,Dynatrace and Splunk)
Site reliability engineer (Java,Unix,Dynatrace and Splunk)

FIS • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Opportunities for professional education
Vibrant team environment
High degree of responsibility
+1
Engineering Division - SRE Platforms - Software Engineering - Vice President - Hyderabad
Engineering Division - SRE Platforms - Software Engineering - Vice President - Hyderabad

Goldman Sachs • Hyderabad

On-site
INR 2,500,000 - 3,500,000