Senior Systems Reliability Engineer I

Thoughtspot

Bengaluru

On-site

INR 1,200,000 - 2,800,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ThoughtSpot in Bengaluru is seeking a Technical & Customer Support and System Reliability professional to own customer-facing technical issues on our SaaS platform, including data connectivity, report errors, performance concerns, and integration challenges. You will translate complex tech into clear updates and drive issues through to resolution.

You will monitor cloud infrastructure with Grafana, Prometheus, Datadog, Splunk, and apply AI/ML-driven observability.

Qualifications

  • B.S./BTech in Computer Science or equivalent.
  • Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms.
  • Hands-on experience with monitoring tools such as Grafana, Prometheus, Datadog, or Splunk.
  • Demonstrated experience leveraging AI/ML to address SRE challenges.

Responsibilities

  • Act as the primary point of contact for customer-facing technical issues related to our SaaS platform.
  • Understand and empathize with customer challenges and offer tailored solutions.
  • Provide timely updates to customers and drive issues through to resolution.
  • Translate complex technical issues for both technical and non-technical stakeholders.
  • Create and maintain knowledge-base articles to empower self-service.

Skills

Linux troubleshooting
Cloud platforms
AI/ML for SRE
Incident management
On-call rotations
Scripting: Python Go Bash Java
Communication

Education

B.S. in Computer Science

Tools

Grafana
Prometheus
Datadog
Splunk

Job description

What Youll Do
Technical & Customer Support
  • Act as the primary point of contact for customer-facing technical issues related to our SaaS platform, including data connectivity, report errors, performance concerns, access problems, data inconsistencies, software bugs, and integration challenges.
  • Understand and empathize with the challenges ThoughtSpot users face, offering tailored solutions to improve their experience.
  • Provide timely, accurate, and clear updates to customers, consistently meeting SLAs and driving issues through to full resolution via tickets and calls.
  • Translate complex technical issues into clear, concise updates for both technical and non-technical stakeholders.
  • Create and maintain knowledge-base articles to empower customer self-service and improve support efficiency.
System Reliability & Monitoring
  • Maintain, monitor, and troubleshoot ThoughtSpot cloud infrastructure using tools like Grafana, Prometheus, Datadog, and Splunk.
  • Monitor system health and performance through metrics, logs, and dashboards to detect and prevent issues proactively.
  • Implement and leverage AI/ML-driven solutions for proactive observability, predictive anomaly detection, and intelligent alerting to enhance service reliability and reduce Mean Time to Resolution MTTR.
  • Understand and apply NetOps and SecOps principles for cloud and on-premise deployments.
  • Develop and implement automation and best practices to streamline operations and strengthen system reliability.
  • Optimize SRE workflows with AI tools to boost operational effectiveness.
Incident Management & Continuous Improvement
  • Participate in on-call rotations, lead incident reviews, and conduct thorough root cause analyses to drive continuous improvement.
  • Work cross-functionally with Engineering to define and implement tools that enhance debuggability, supportability, availability, scalability, and performance.
  • Be an expert in both cloud and on-premise infrastructure by developing automation and best practices.
What Youll Bring
  • B.S. in Computer Science or equivalent relevant experience.
  • Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms (VMware, AWS, Azure, GCP).
  • Hands-on experience with monitoring tools such as Grafana, Prometheus, Datadog, or Splunk.
  • Demonstrated experience and a keen interest in leveraging AI/ML principles to address SRE challenges - including AIOps, predictive maintenance, and intelligent automation.
  • Prior experience in enterprise customer support, including on-call rotations and incident management, with the ability to lead root cause analyses.
  • Strong problem-solving and algorithmic thinking with a solid understanding of system internals.
  • Excellent verbal and written communication skills with the ability to work independently and cross-functionally in fast-paced environments.
  • Familiarity with scripting and programming languages such as Python, Go, Bash, or Java.
  • Exposure to infrastructure and service monitoring frameworks with the ability to analyze data to ensure high availability.
Good to Have
  • Experience partnering with Engineering to design and implement mission-critical tooling and automation that advances system debuggability, high availability, elastic scalability, and performance.
  • Experience with alerting strategies and monitoring system tuning to minimize alert fatigue and optimize Mean Time to Acknowledge MTTA.
  • Familiarity with C/C++ or other low-level systems languages.
Ideal Candidate Profile

You have a balanced mix of technical expertise in cloud operations and a proven record of handling support incidents and end-user queries. This sets you apart from candidates with purely systems or cloud engineering backgrounds. You move fluidly between deep technical investigation and customer-facing communication - equally at home diagnosing a complex infrastructure issue and presenting findings clearly to an enterprise stakeholder.

Mandatory and Required Skills for All ThoughtSpot Roles

Spotters are expected to demonstrate AI literacy and workflow integration to include the ability to:

  • Comfortably and confidently integrate artificial intelligence into their daily workflow to increase productivity and quality.
  • Hands-on experience to leverage AI tools (industry-leading LLMs) to increase productivity, automate routine tasks, and improve work quality.
  • Speak to the experience of using AI for research, content creation, and document summarization while maintaining ownership of judgment and final decisions.
  • Write effective prompts to get the most accurate and creative results from AI tools.

Disclaimer: This job description has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Reliability Engineer I
Senior Systems Reliability Engineer I

ThoughtSpot • India

On-site
INR 1,800,000 - 2,400,000
Competitive salary
Career growth opportunities
Collaborative culture
Member of Technical Staff 4
Member of Technical Staff 4

ThoughtSpot • Hyderabad

On-site
INR 1,800,000 - 3,000,000
In-office 3+ days per week
Staff Engineer
Staff Engineer

ThoughtSpot • Bengaluru

Hybrid
INR 3,000,000 - 5,500,000
Staff Engineer
Staff Engineer

ThoughtSpot • India

On-site
INR 3,500,000 - 5,200,000
Staff Engineer - Backend
Staff Engineer - Backend

ThoughtSpot • Bengaluru

On-site
INR 1,800,000 - 3,000,000
ThoughtSpot Platform Engineer
ThoughtSpot Platform Engineer

Tekskills • Pune District

On-site
INR 3,500,000 - 5,000,000
Engineering Manager
Engineering Manager

ThoughtSpot • Bengaluru

On-site
INR 4,200,000 - 6,600,000
ThoughtSpot Platform Engineer
ThoughtSpot Platform Engineer

Tekskills • Dadri

On-site
INR 2,000,000 - 3,200,000
ThoughtSpot Platform Engineer
ThoughtSpot Platform Engineer

Tekskills • Bengaluru

On-site
INR 2,400,000 - 4,200,000
ThoughtSpot Platform Engineer
ThoughtSpot Platform Engineer

Tekskills • Chennai District

On-site
INR 1,800,000 - 3,200,000