Senior Observability Engineer: Automation & Reliability

Royal Cyber

United States

Remote

USD 120,000 - 180,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Royal Cyber is seeking a monitoring and observability specialist in the United States to design, implement, and maintain enterprise monitoring using LogicMonitor across infrastructure, apps, cloud services, networks, and critical systems. You will build dashboards, set KPIs, and align practices with governance standards.

You will drive automation, alert enrichment, ITSM integrations, and incident response improvements, collaborating with NOC and engineering teams to reduce manual toil and

Responsibilities

  • Monitoring & Observability Management: Develop, implement, and maintain enterprise monitoring standards and best practices. Design and configure monitoring solutions for infrastructure, applications, cloud services, networks, and business-critical systems using LogicMonitor.
  • Create and maintain operational dashboards, performance views, service health dashboards, and executive reporting dashboards.
  • Establish monitoring baselines, thresholds, KPIs, and health indicators to support proactive operations.
  • Ensure comprehensive monitoring coverage across enterprise environments and identify monitoring gaps.
  • Automation & Event Management: Design and implement automation solutions to streamline monitoring, event management, and operational workflows.
  • Develop automated event correlation and alert enrichment capabilities to improve incident response effectiveness. Support integration of monitoring platforms with ITSM, ticketing, notification, and automation tools.
  • Identify opportunities to eliminate manual operational tasks through automation and workflow optimization. Collaborate with NOC and engineering teams to improve operational efficiency through automation initiatives.
  • Alert Optimization & Operational Excellence: Analyze alert trends and monitoring data to identify opportunities for noise reduction and operational improvements.
  • Improve alert accuracy, reduce false positives, and enhance event detection capabilities. Develop and maintain alert tuning strategies to ensure actionable and meaningful notifications.
  • Support root cause analysis activities through effective monitoring and data correlation.
  • Establish monitoring governance practices and operational standards. Platform Administration & Continuous Improvement: Administer, maintain, and optimize the LogicMonitor platform and supporting integrations.
  • Manage monitoring configurations, device onboarding, templates, collectors, escalation chains, and alert rules. Monitor platform performance, availability, scalability, and reliability.
  • Ensure platform configurations align with operational requirements and governance standards.
  • Participate in capacity planning, system upgrades, feature adoption, and platform enhancement initiatives. Drive continual service improvement through monitoring maturity assessments and operational reviews.
  • Operational Support & Reporting: Provide subject matter expertise during incidents, major incidents, and service-impacting events. Develop regular reports and metrics related to monitoring effectiveness, alert quality, platform performance, and automation benefits.
  • Support operational reviews by providing trend analysis and recommendations for service improvements. Collaborate cross-functional teams to enhance overall observability and service reliability.

Job description

Royal Cyber is seeking a monitoring and observability specialist in the United States to design, implement, and maintain enterprise monitoring using LogicMonitor across infrastructure, apps, cloud services, networks, and critical systems. You will build dashboards, set KPIs, and align practices with governance standards.

You will drive automation, alert enrichment, ITSM integrations, and incident response improvements, collaborating with NOC and engineering teams to reduce manual toil and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

U.S. Bank • Hopkins (MN)

On-site
USD 98,000 - 116,000
Healthcare (medical, dental, vision)
Life insurance
Disability leave
+6
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

U.S. Bank • Town of Brookfield (WI), Northern (KY)

Hybrid
USD 98,000 - 116,000
Healthcare (medical, dental, vision)
Retirement plan (401(k))
Paid vacation
+2
Lead Observability & SRE Architect
Lead Observability & SRE Architect

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Senior SRE - Observability & Automation (Remote)
Senior SRE - Observability & Automation (Remote)

OnBoard Group • Northern (KY)

Hybrid
USD 120,000 - 150,000
Fully remote work with equipment
Competitive benefits: medical, dental,
401K with company match
+2
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Senior Observability Engineer
Senior Observability Engineer

Calance • Orlando (FL)

Hybrid
USD 80,000 - 120,000
Senior Observability Engineer: Build Reliability & Dashboards
Senior Observability Engineer: Build Reliability & Dashboards

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Sr Observability Engineer
Sr Observability Engineer

IT Associates • Irvine (CA)

Hybrid
USD 150,000 - 210,000
Observability & Reliability Engineer II
Observability & Reliability Engineer II

U.S. Bank • Chicago (IL)

On-site
USD 86,000 - 102,000
Healthcare
Retirement plan
Paid vacation
+2