Senior Systems Reliability Engineer II

Cerebras

Mountain View (CA)

Hybrid

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment

Job summary

Cerebras is looking for a Site Reliability Engineer to join our innovative team in Mountain View, California. You will ensure service reliability and act as a trusted customer partner, leveraging AI/ML for operational intelligence.

This role involves troubleshooting complex technical issues, maintaining cloud infrastructure, and providing exceptional customer support. Ideal candidates will have a strong background in Linux systems and excellent communication skills, joining the team in a hybrid work environment.

Qualifications

  • Proven experience troubleshooting complex Linux systems.
  • Hands-on experience with monitoring tools like Grafana and Datadog.
  • Strong problem-solving skills with a solid understanding of system internals.

Responsibilities

  • Act as the primary point of contact for customer-facing technical issues.
  • Maintain, monitor, and troubleshoot ThoughtSpot cloud infrastructure.
  • Participate in on-call rotations and lead incident reviews.

Skills

Troubleshooting complex Linux systems
Hands-on experience with monitoring tools
Excellent verbal and written communication
Problem-solving and algorithmic thinking

Education

B.S. in Computer Science or equivalent

Tools

Grafana
Prometheus
Datadog
Splunk
VMware
AWS
Azure
GCP
Python
Go
Java
Bash

Job description

The Role

As part of the ThoughtSpot SRE team, you will be on the cutting edge of operational intelligence. You will not only ensure service reliability but also act as a trusted partner for our customers — proactively leveraging AI/ML to deliver timely updates, meaningful solutions, and predictive improvements. You are the bridge between our customers and engineering, combining deep systems expertise with a genuine passion for customer success. If you thrive in dynamic environments and are committed to building resilient, self‑optimizing systems, this role is for you.

What You'll Do
Technical & Customer Support
  • Act as the primary point of contact for customer‑facing technical issues related to our SaaS platform, including data connectivity, report errors, performance concerns, access problems, data inconsistencies, software bugs, and integration challenges.
  • Understand and empathize with the challenges ThoughtSpot users face, offering tailored solutions to improve their experience.
  • Provide timely, accurate, and clear updates to customers, consistently meeting SLAs and driving issues through to full resolution via tickets and calls.
  • Translate complex technical issues into clear, concise updates for both technical and non‑technical stakeholders.
  • Create and maintain knowledge‑base articles to empower customer self‑service and improve support efficiency.
System Reliability & Monitoring
  • Maintain, monitor, and troubleshoot ThoughtSpot cloud infrastructure using tools like Grafana, Prometheus, Datadog, and Splunk.
  • Monitor system health and performance through metrics, logs, and dashboards to detect and prevent issues proactively.
  • Implement and leverage AI/ML‑driven solutions for proactive observability, predictive anomaly detection, and intelligent alerting to enhance service reliability and reduce Mean Time to Resolution (MTTR).
  • Understand and apply NetOps and SecOps principles for cloud and on‑premise deployments.
  • Develop and implement automation and best practices to streamline operations and strengthen system reliability.
  • Optimize SRE workflows with AI tools to boost operational effectiveness.
Incident Management & Continuous Improvement
  • Participate in on‑call rotations, lead incident reviews, and conduct thorough root‑cause analyses to drive continuous improvement.
  • Work cross‑functionally with Engineering to define and implement tools that enhance debuggability, supportability, availability, scalability, and performance.
  • Be an expert in both cloud and on‑premise infrastructure by developing automation and best practices.
What You'll Bring
  • B.S. in Computer Science or equivalent relevant experience.
  • Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms (VMware, AWS, Azure, GCP).
  • Hands‑on experience with monitoring tools such as Grafana, Prometheus, Datadog, or Splunk.
  • Demonstrated experience and a keen interest in leveraging AI/ML principles to address SRE challenges — including AIOps, predictive maintenance, and intelligent automation.
  • Prior experience in enterprise customer support, including on‑call rotations and incident management, with the ability to lead root‑cause analyses.
  • Strong problem‑solving and algorithmic thinking with a solid understanding of system internals.
  • Excellent verbal and written communication skills with the ability to work independently and cross‑functionally in fast‑paced environments.
  • Familiarity with scripting and programming languages such as Python, Go, Bash, or Java.
  • Exposure to infrastructure and service monitoring frameworks with the ability to analyze data to ensure high availability.
Good to Have
  • Experience partnering with Engineering to design and implement mission‑critical tooling and automation that advances system debuggability, high availability, elastic scalability, and performance.
  • Experience with alerting strategies and monitoring system tuning to minimize alert fatigue and optimize Mean Time to Acknowledge (MTTA).
  • Familiarity with C/C++ or other low‑level systems languages.
Ideal Candidate Profile

You have a balanced mix of technical expertise in cloud operations and a proven record of handling support incidents and end‑user queries. This sets you apart from candidates with purely systems or cloud engineering backgrounds. You move fluidly between deep technical investigation and customer‑facing communication — equally at home diagnosing a complex infrastructure issue and presenting findings clearly to an enterprise stakeholder.

What We Offer
  • Competitive salary and benefits package.
  • Opportunities for professional growth and career advancement.
  • A collaborative work environment where your input and expertise directly impact customer experience and platform reliability.
Hybrid Work at ThoughtSpot

Spotters are expected in‑office 3 days per week to experience the energy of their local office. This approach balances the benefits of in‑person collaboration and peer learning with the flexibility needed by individuals and teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Reliability Engineer
Senior Systems Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 90,000 - 120,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
Senior Systems Reliability Engineer
Senior Systems Reliability Engineer

ThoughtSpot • Mountain View (CA)

On-site
USD 140,000 - 170,000
Solutions Engineer
Solutions Engineer

Cerebras • Chicago (IL)

Hybrid
USD 150,000 - 198,000
Hybrid work arrangement
Sales Development Representative
Sales Development Representative

Cerebras • Chicago (IL)

On-site
USD 72,000 - 88,000
Equity
Company bonus or sales commissions
401(k) plan
+1
SRE Support Engineer - Observability
SRE Support Engineer - Observability

Gigster • Austin (TX)

Remote
USD 80,000 - 100,000
High autonomy in a remote-first environment
Real technical problem solving
Opportunity for scaling support
Senior SRE: AI-Driven Reliability & Customer Impact
Senior SRE: AI-Driven Reliability & Customer Impact

ThoughtSpot • Mountain View (CA)

On-site
USD 140,000 - 170,000
Enterprise Account Executive
Enterprise Account Executive

ThoughtSpot • California (MO)

Hybrid
USD 230,000 - 300,000
Manager, Web Production
Manager, Web Production

Sapphire Partners • Chicago (IL)

Hybrid
USD 145,000 - 195,000
Hybrid work
Sales Development Representative
Sales Development Representative

ThoughtSpot • Phoenix (AZ)

Hybrid
USD 60,000 - 90,000
Hybrid work model
Remote work available
Commercial Account Executive
Commercial Account Executive

Sapphire Partners • New York (NY)

Hybrid
USD 90,000 - 150,000
Hybrid work
Remote opportunities
Training and career growth