Overview
At Blue Cross and Blue Shield of Nebraska, we are a mission‑driven organization dedicated to championing the health and well‑being of our members and the communities we serve. This Technical Analyst role focuses on Site Reliability Engineering with a strong emphasis on observability and AI‑assisted operations, bridging technical teams and business stakeholders to drive platform reliability.
What you’ll do
- Integrate and align enterprise observability across vendor‑managed and internal platforms, ensuring monitoring meets reliability policies, SLO/SLA targets, and MTTR goals.
- Own the service dependency map and asset criticality model—identify tier‑1 assets, fail‑over mechanisms, and single points of failure.
- Establish end‑user experience monitoring to detect member or provider experience issues, not just infrastructure alarms.
- Leverage AI and AIOps tooling for incident triage, root‑cause analysis, automated runbook execution, predictive anomaly detection, alert correlation, and self‑healing remediation.
- Define and track SLOs for platform reliability—availability, latency, error rates—into business‑aligned targets and error budgets.
- Build and maintain MTTR and reliability dashboards—platform health, incident trends, perfect‑day streaks, vendor SLA compliance.
Qualification – Must Have
- 3+ years of experience in Site Reliability Engineering, infrastructure engineering, platform operations, DevOps, or a related technical operations role supporting enterprise systems.
- Hands‑on experience with enterprise observability and monitoring platforms, ITSM/ticketing systems, operational dashboards, and at least one scripting language (Python, PowerShell, or Bash).
- Experience with AI/AIOps tools, telemetry, logs, traces, alert correlation, and incident response processes is strongly preferred.
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent education and experience.
- Strong communication, analytical problem‑solving, and knowledge of ITIL fundamentals (incident, change, and problem management).
Preferred
- Observability, SRE, DevOps, or cloud certifications.
- Healthcare payer or health plan environment experience.
- Agile/Scrum delivery experience and engineering partnership to translate reliability work into backlog items and sprint‑ready user stories.
- Experience with AI/ML operations tooling, CI/CD pipelines, infrastructure‑as‑code, CMDB or service dependency mapping, log aggregation or SIEM platforms, distributed systems reliability, chaos engineering, and report/dashboard authoring or data modeling.
EEO Statement
Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.