Senior Platform Engineer (Observability & Telemetry)

Ports North

Baltimore (MD)

On-site

USD 137,760 - 165,312

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Korn Ferry in Baltimore, MD is seeking a Senior Platform Engineer (Observability & Telemetry) to join the high‑performing Monitoring Engineering team in a fast‑paced financial technology environment.

You will design and operate OpenTelemetry pipelines, instrumentation standards, dashboards, and alerting to deliver end‑to‑end visibility, reliable services, and reduced MTTR through SLO‑driven practices. Collaboration with application and platform teams is essential to raise enterprise reliability.

Qualifications

  • Experience designing and implementing observability platforms across large, distributed environments.
  • Strong knowledge of instrumentation, alerting, dashboards, and tracing.
  • Proven ability to drive reliability through SLOs,SLIs, and incident reduction.

Responsibilities

  • Design, build, and maintain monitoring and observability solutions across production systems.
  • Develop instrumentation, telemetry, and alerting for enterprise monitoring centers using industry tools.
  • Collaborate with cross-functional teams to document observability requirements and ensure reliable release processes.
  • Apply SRE best practices to define SLIs/SLOs and health indicators for critical services.

Skills

Observability design
Telemetry pipelines
OpenTelemetry
SRE practices
Cross-team collaboration

Education

Bachelor's degree in Computer Science/IT

Tools

Grafana
OpsRamp
ElasticStack
BigPanda
AWS CloudWatch
Azure Monitor

Job description

Leading consumer finance company is seeking a Senior Platform Engineer (Observability & Telemetry) to join a high-performing Monitoring Engineering team within a fast-paced financial technology organization.

You will partner closely with application, platform, and development teams to implement data-driven alerting, SLO/SLA-based monitoring, telemetry pipelines, dashboards, correlations, and automated remediation. Your work will directly improve system reliability, reduce MTTR, and enhance enterprise-wide operational insight.

This role requires strong analytical thinking, systems engineering discipline, and a proactive approach to identifying risks, preventing incidents, and driving continuous improvement across the production ecosystem.

Key Responsibilities
Design, Build, and Maintain Monitoring & Observability Solutions
  • Architect, deploy, and operate OpenTelemetry-based telemetry pipelines, including instrumentation standards, collector configurations, sampling strategies, and routing to Elastic and other backends.
  • Develop and maintain instrumentation, telemetry, and alerting for the Enterprise Monitoring Center using industry-leading tools, such as:
    • Grafana, OpsRamp, ElasticStack, BigPanda
    • AWS CloudWatch, Azure Monitor
  • Drive observability standards and best practices across multiple engineering teams through influence, documentation, and partnership rather than direct authority.
  • Apply SRE best practices to ensure measurable SLIs/SLOs, reliability dashboards, and health indicators for critical systems.
  • Integrate and manage OpenTelemetry for distributed tracing and telemetry data collection, enabling end-to-end visibility of business-critical transactions.
Collaboration & Project Participation
  • Collaborate with application development teams to define and document observability requirements for each project or release, ensuring accurate and actionable monitoring and tracing are in place for every step of business-critical workflows.
  • Embed reliability considerations early in the SDLC, including SLO definitions, instrumentation needs, and failure-mode awareness.
  • Partner with product and engineering teams to use SLOs and error budgets to guide release decisions, prioritization, and toil reduction.
Alerting & Escalation Process
  • Define and maintain standardized alert payloads per engineering guidelines, ensuring alerts are actionable.
  • Partner with Level 2 and Level 3 support teams to reflect process changes in monitoring dashboards.
  • Maintain and optimize thresholds, ensuring seamless escalations via BigPanda as the central alert hub.
Dashboard Creation & Maintenance
  • Create and maintain intuitive, actionable dashboards for the Enterprise Monitoring Center and other finance teams.
  • Ensure dashboards are effectively monitored by Level 1 teams, presenting clear, actionable data that reduces MTTR.
Documentation, Governance & Reliability Standards
  • Develop and maintain technical documentation, runbooks, diagnostic guides, and observability standards across the enterprise.
  • Evaluate and refine release, deployment, and monitoring processes to support consistent, reliable delivery pipelines.
  • Mentor junior engineers and promote a culture focused on reliability, automation, and operational excellence.
Reliability Engineering, Automation & Continuous Improvement
  • Build automation frameworks for monitoring, alerting, self-healing workflows, and incident response to reduce toil and improve MTTR.
  • Drive system optimization through capacity analysis, performance tuning, and proactive detection of reliability risks.
  • Contribute to the automation of routine operational tasks to improve system reliability and engineer quality of life.
  • Advocate for and implement observability best practices across engineering teams.
  • Define, implement, and operationalize SLIs, SLOs, and error budgets for critical services.
  • Participate in and improve incident response processes, including detection, triage, escalation, and recovery.
Qualifications

Education bachelor's in computer science, IT, or related field.

Experience
  • 5+ years of experience in software, systems, or reliability engineering roles, with multiple years of hands-on experience owning production observability, monitoring, and SLOs in distributed systems.
Required Skills
  • Deep experience building scalable, reliable monitoring and observability solutions, including instrumentation, alerting, dashboarding, and configuration across large, complex environments.
  • Hands-on expertise and proficency with modern monitoring and observability tools, (e.g., OpsRamp, Grafana, Elastic, CloudWatch, Azure Monitor BigPanda (AIOps), and strong knowledge of metrics, logs, traces, and OpenTelemetry.
  • Strong scripting and programming capability (Bash, PowerShell, and one or more languages such as Python, C-family, or JavaScript) to automate telemetry, alerting, and platform workflows.
  • Strong expertise with cloud platforms (AWS and/or Azure) and container orchestration systems (Kubernetes, Docker).
  • Deep hands-on experience with Elastic Observability (APM, Logs, Metrics, Traces)
  • Understanding of distributed systems fundamentals, including networking, security, databases, DevSecOps principles, and performance/capacity engineering.
  • Strong communication skills, with the ability to clearly explain complex technical topics to both technical and non-technical audiences.
  • Exceptional problem-solving and troubleshooting abilities, especially in high-pressure or time-sensitive environments.
  • Effective prioritization and multitasking, able to manage competing deadlines while maintaining quality and focus.
  • Proven cross-functional collaboration, working seamlessly with diverse teams in large, complex IT environments and driving continuous improvement across systems.
Preferred Qualifications
  • Experience with CI/CD pipelines and tools like Jenkins, GitHub, GitLab CI, or CircleCI
  • Experience querying, manipulating, and visualizing time-series data.
  • Familiarity with Infrastructure as Code tools (e.g., Ansible, Terraform).
  • Knowledge of microservices architecture and event-driven systems.
  • Working knowledge of REST APIs, JSON, and ServiceNow.
  • Experience with cloud monitoring-particularly AWS or Azure.

Title : Sr Platform Engineer
Location Baltimore, MD
Client Industry Finance
Compensation $100-120/hr

About Korn Ferry

Korn Ferry unleashes potential in people, teams, and organizations. We work with our clients to design optimal organization structures, roles, and responsibilities. We help them hire the right people and advise them on how to reward and motivate their workforce while developing professionals as they navigate and advance their careers. To learn more, please visit Korn Ferry at www.Kornferry.com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Telemetry & Observability Platform Engineer
Senior Telemetry & Observability Platform Engineer

Ports North • Baltimore (MD)

On-site
Senior Platform Engineer
Senior Platform Engineer

Ww • United States

On-site
USD 200,000 - 215,000
Observability Engineer
Observability Engineer

The Fountain Group • Dallas (TX)

Hybrid
USD 103,000 - 117,000
Observability Operations Engineer
Observability Operations Engineer

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 100,000 - 120,000
Senior Observability Engineer
Senior Observability Engineer

Tata Consultancy Services • Los Angeles (CA)

On-site
USD 120,000 - 130,000
Back fill Engineer
Back fill Engineer

JPC TECHNO INC • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Principal Platform Engineer
Principal Platform Engineer

Flexential • United States

On-site
USD 180,000 - 210,000
Medical, Telehealth, Dental and Vision
HSA and FSA
AD&D
+2
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Solution Architect / Team Lead - Observability
Solution Architect / Team Lead - Observability

VOLTO Consulting • Irvine (CA)

On-site
USD 120,000 - 160,000