Job Description
What will you do?
- Design and implement observability solutions using industry-leading platforms, establishing logging standards that enable comprehensive system visibility across all systems
- Create and maintain monitoring dashboards that provide actionable insights into system health and performance, partnering with platform and application teams to integrate observability into architecture
- Evaluate and recommend observability tools and vendors to ensure the organization has access to best-in-class solutions
- Analyze logs, metrics, and traces to proactively identify system issues, performance bottlenecks, and translate observability data into insights about user patterns and system behavior
- Develop predictive monitoring strategies using AI tools to detect anomalies, identify emerging trends, and prevent incidents before they occur
- Conduct "what if" analysis using AI capabilities to model potential scenarios and their impact on system performance
- Apply machine learning-based anomaly detection to identify issues proactively and use predictive analytics to forecast system behavior and prevent failures
- Collaborate with cross-functional teams (Engineering, DevOps, Security, production support, product) to communicate complex observability concepts to both technical and non-technical stakeholders
- Lead observability initiatives and mentor junior team members on best practices while facilitating design and problem-solving discussions across the organization
Responsibilities
Lead and execute observability initiatives, establish standards, and partner with engineering teams to improve system visibility, reliability, and performance.
Qualifications
- 5+ years of experience with observability tools (ELK Stack, Dynatrace, Prometheus, Grafana, OpenTelemetry, Jaeger, Aternity or similar)
- 3+ years of software engineering or infrastructure experience
- Python, Java, Go
- Query languages
- Expert-level knowledge of logging requirements and best practices for enhanced observability
- Demonstrated experience building and optimizing monitoring dashboards
- Proven ability to use observability data to proactively identify and resolve system issues
- Experience using AI tools for anomaly detection and trend analysis
- Linux/Unix knowledge
Nice to have
- Expertise with Tableau or advanced visualization tools
- Experience in financial technology and financial services environments
- Previous experience with AI tools such as Anthropic, OpenAI, Devin, Copilot, etc.
- Experience with OpenShift, Kubernetes and containerized applications
- Background in DevOps or Site Reliability Engineering (SRE)
- Experience conducting "what if" analysis and scenario modeling
- Identify networking slowness and availability issues
What7s in it for you?
- A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, and stock where applicable
- Leaders who support your development through coaching and mentoring opportunities
- Ability to make a difference and lasting impact on system reliability and user experience
- Work in a dynamic, collaborative, progressive, and high-performing team
The expected salary range for this position is $90,000-$140,000, depending on your experience, skills, and market conditions.
You have the potential to earn more through discretionary variable compensation based on business performance and individual goals.
Job Details
- Address: 250 Nicollet Mall, Minneapolis, MN, United States
- City: Minneapolis
- Country: United States of America
- Work hours/week: 40
- Employment Type: Full time
- Platform: Technology and Operations
- Job Type: Regular
- Pay Type: Salaried
RBC is an equal opportunity employer. We are committed to fostering an inclusive workplace that values diverse perspectives. We provide policies and programs to support belonging and opportunity for all employees.