SRE Support Engineer - Observability

Gigster

Austin (TX)

Remote

USD 80,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High autonomy in a remote-first environment
Real technical problem solving
Opportunity for scaling support

Job summary

A leading technology company is looking for an Observability & Tools Support Engineer to provide high-impact technical support for its internal IaaS platform. This role involves monitoring, alerting, and troubleshooting, with a focus on helping customers seamlessly onboard and resolve issues swiftly. Ideal candidates will have strong Linux knowledge, experience with tools like Prometheus, and excellent communication skills. The position allows for high autonomy in a remote-first environment, fostering career growth in technical support and process improvement.

Qualifications

  • Several years supporting highly scalable applications and web services.
  • Strong understanding of the Linux operating system.
  • Excellent analytical capability and attention to detail.

Responsibilities

  • Manage Slack threads and tickets for customer support.
  • Troubleshoot monitoring and alerting issues.
  • Build customer-facing knowledge base articles.

Skills

Prometheus troubleshooting
AlertManager expertise
Linux system knowledge
TCP/IP fundamentals
Customer-facing communication

Tools

Kubernetes
OpenTelemetry

Job description

Role Overview

The Observability & Tools Support Engineer provides high-impact technical support for customers of a large technology company’s internal IaaS platform, with a focus on monitoring, alerting, telemetry, and operational tooling.

This role spans a wide range of support—from white-glove onboarding and end-to-end customer enablement, to deep technical troubleshooting across Linux, networking, and observability systems (especially Prometheus and AlertManager). You will also contribute to improving the support function itself: strengthening tooling, documentation, workflows, and feedback loops so the service scales.

Success depends on excellent troubleshooting, strong written communication, comfort working with highly technical customers, and the maturity to identify patterns and drive operational improvements beyond individual ticket resolution.

Business Outcome

Become a trusted frontline expert for the customer’s observability ecosystem and operational tooling - delivering fast, accurate support across Slack and tickets, improving monitoring reliability, and reducing incident impact through better triage, troubleshooting, onboarding, and knowledge capture.

Success Measures
  • Healthy volume of threads and tickets handled with high-quality outcomes
  • Consistent achievement of time-based SLAs
  • High customer satisfaction through surveys
  • Accurate classification of issue type, severity, and recurring patterns
  • Reduced repeat issues through better docs, tooling, and scalable onboarding
What Will Be True When You Succeed
  • Customers can onboard smoothly to monitoring/alerting with minimal friction
  • Monitoring and alerting issues are resolved quickly, with fewer escalations
  • Linux and networking-related incidents reach resolution faster due to strong troubleshooting and clean handoffs
  • Engineering and SRE teams receive clear, actionable feedback based on real customer trends
  • Knowledge base content prevents tickets and accelerates self-service
Core Work Units
  1. Frontline Support for Observability & Tooling
    • Manage Slack threads and tickets (roughly 50/50)
    • Handle a broad range of customer support: simple issue resolution through end-to-end onboarding
    • Provide clear, structured guidance to highly technical customers
    • Maintain strong attention to detail while managing multiple interactions in parallel
  2. Deep-Dive Troubleshooting & Incident Support
    • Troubleshoot, isolate, and resolve monitoring and alerting issues (especially Prometheus + AlertManager)
    • Troubleshoot complex Linux and networking issues (TCP/IP fundamentals required)
    • Support OpenTelemetry, tracing, and telemetry pipelines, including investigation of gaps in signals and instrumentation
    • Drive incidents to resolution in partnership with Engineering/SRE teams
  3. Documentation & Knowledge Development
    • Build and maintain customer-facing and internal knowledge base articles
    • Create informational posts for the community support platform
    • Turn repeated issues into reusable guides, checklists, and onboarding playbooks
  4. Trend Analysis & Feedback to Engineering
    • Analyze and categorize customer interaction trends
    • Provide accurate, meaningful feedback to Engineering and SRE orgs to improve product/tooling
    • Identify “top offenders” and propose practical fixes (tooling, docs, process, product)
  5. Operational Excellence & Continuous Improvement
    • Participate in post-mortem reviews and drive follow-through on improvements
    • Contribute meaningfully to team objectives and goals (process, tooling, and service scaling)
    • Bring creativity and discretion to resolve highly complex issues “outside the box”
High-Quality Work - what top performance looks like

Frontline Support

  • Moves smoothly from triage to deeper analysis without losing the customer
  • Communicates clearly and confidently with technical users
  • Maintains clean follow-ups and thread hygiene even with high context switching

Troubleshooting

  • Rapidly isolates issues across monitoring/alerting configs, Linux runtime behavior, and network connectivity
  • Uses structured approaches to incident handling: hypothesis → test → evidence → resolution
  • Produces high-signal writeups that accelerate downstream resolution

Documentation & Enablement

  • Documentation is clear enough that customers avoid opening tickets
  • Onboarding flows reduce time-to-value and prevent common misconfigurations
  • Captures “tribal knowledge” quickly and makes it reusable

Operational Excellence

  • Obsessing over details: correct severity, accurate tagging, clean timelines, strong handoffs
  • Spots patterns early and proactively proposes improvements that scale support
Typical Day / Work Patterns
  • ~50% Slack support, ~50% ticket handling
  • Deep-dive investigations during lower ticket volume periods
  • Documentation writing and lightweight tooling/process improvements when patterns emerge
  • Weekly team review of escalations, themes, and operational improvements
  • High rate of context switching and parallel issue management
Required Skills & Experience (Non-Negotiable)
  • Several years supporting highly scalable applications and web services
  • Hands-on experience with open-source observability and cloud-native tooling, including:
    • Kubernetes (and container fundamentals)
    • Prometheus and AlertManager troubleshooting
    • OpenTelemetry and distributed tracing concepts
  • Strong understanding of the Linux operating system (command line, process/network debugging, logs)
  • Good understanding of infrastructure observability principles (signals, alerting strategy, SLO thinking, noise reduction)
  • Good understanding of the TCP/IP suite and practical networking troubleshooting
  • Strong experience troubleshooting ambiguous, multi-layer issues
  • Excellent analytical capability and strong attention to detail
  • Strong written and verbal communication (clear, structured, customer-friendly)
  • Comfortable working with a very technical customer base
  • Passion for Technical Support and a service mindset
Nice-to-Haves
  • Experience improving or supporting internal support tooling or workflows (automation, templates, runbooks)
  • Experience operating at scale in a services environment (pattern detection, KPI/SLA awareness, operational process maturity)
  • Familiarity with Grafana, log aggregation, incident tooling, and production support practices
  • Prior SRE or platform support experience
Minimum Qualifications
  • 3–7+ years in Technical Support Engineering, SRE support, DevOps, Platform Support, or similar
  • Demonstrated experience supporting distributed systems, IaaS, or cloud platforms
  • Strong Linux, troubleshooting, and customer-facing communication background
  • Evidence of documentation, knowledge-base contributions, and process improvement mindset

Disqualifiers: weak Linux fundamentals, inability to troubleshoot systematically, poor written communication, or discomfort supporting highly technical users.

What You’ll Love
  • Real technical problem solving with tangible customer impact
  • A role that blends deep troubleshooting with scaling support via docs, tooling, and process
  • High autonomy in a remote-first environment
What May Be Challenging
  • High context switching and managing multiple threads in parallel
  • Repeated patterns that require discipline to convert pain into scalable improvements
  • Supporting high-visibility systems where speed and accuracy matter
Differentiation

Industry: Remote-first, trust-based culture; global team; autonomy; modern systems; meaningful technical challenges

Internal: High-impact, customer-facing observability support; direct influence on tooling and process maturity; opportunity to shape scalable support practices

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Lead Software Reliability Engineer
Lead Software Reliability Engineer

Jobtailor • Atlanta (TX)

On-site
USD 120,000 - 180,000
Technical Support Engineer
Technical Support Engineer

Jobtailor • California (MO)

On-site
USD 120,000 - 150,000
Senior Systems Reliability Engineer II
Senior Systems Reliability Engineer II

Cerebras • Mountain View (CA)

Hybrid
USD 100,000 - 150,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
Site Reliability Engineer
Site Reliability Engineer

Matlen Silver • Charlotte (NC)

Hybrid
USD 90,000 - 94,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000