About the Role
The Software Engineering Specialist independently executes advanced activities to deliver the engineering strategy and roadmap that supports BT's commercial strategy through cross‑functional business partnering and the participation of a team that pursues innovation as well as engineering excellence.
What you’ll be doing
Architecture Ownership & Vision
- Own the overall system architecture of the observability platform across ingestion, processing, storage, and query layers.
- Work closely with Enterprise Architect to align system architecture with long‑term architectural vision and technical roadmap aligned with business and platform goals.
- Design and review high‑level system designs, data flows, integration patterns, and core technology choices.
- Act as the technical authority for complex architectural decisions, trade‑offs, and design reviews.
Platform & Domain Leadership
- Architect systems for infrastructure monitoring, metrics, logs, and distributed tracing at scale.
- Guide the evolution from infrastructure monitoring to a full observability platform including APM.
- Define architectural patterns for high‑throughput telemetry ingestion, real‑time processing, and query‑at‑scale.
- Ensure architectural consistency for multi‑tenant, cloud‑native distributed platforms.
Standards, Governance & Enablement
- Establish and evolve architecture standards, design principles, and best practices across teams.
- Identify architectural risks and proactively drive mitigation strategies.
- Enable teams through reference architectures, design frameworks, and technical guidance.
- Support modernization initiatives including scalability, performance optimization, resilience, and cost efficiency.
Hands‑On Architectural Validation
- Perform code reviews of critical and performance‑sensitive components.
- Build and guide proof‑of‑concepts (POCs) to validate architectural decisions and de‑risk new technologies.
- Develop reference implementations to demonstrate architectural intent.
- Collaborate closely with senior engineers to troubleshoot complex system‑level issues.
Delivery Collaboration & Execution Enablement
- Work closely with Software Engineering Managers to align architectural decisions with delivery and release plans.
- Assist in breaking down large architectural initiatives into phased, incremental deliverables.
- Identify architectural and technical risks early and proactively surface them to influence release planning.
- Support release readiness by validating that architecture, scalability, and non‑functional requirements are addressed ahead of key milestones.
The skills you’ll need
- Strong foundation in system design, distributed systems, and application scalability.
- Experience designing microservice‑based, event‑driven architectures.
- Ability to make architectural trade‑offs involving scalability, reliability, performance, and cost.
- Proven experience designing large‑scale, cloud‑based distributed platforms.
- Strong backend experience with Java (Spring Boot‑based microservices).
- Working knowledge of Python and/or Go for scripting, automation, and collectors.
- Ability to review and reason about performance‑critical backend code.
- Strong understanding of infrastructure monitoring and observability concepts with hands‑on experience on:
- Metrics, logs, and distributed tracing.
- Agent‑based and agentless monitoring models.
- Push / pull / subscription based data collection.
- Hands on or deep working knowledge of:
- Metric data models and storage (e.g., VictoriaMetrics or equivalent).
- Log aggregation and search platforms (Elasticsearch / OpenSearch).
- Familiarity with OpenTelemetry, exporters, and custom collectors.
- Understanding of common monitoring tools and protocols (e.g., SNMP, Prometheus‑style systems).
- Strong experience with:
- Time series databases for metrics.
- Relational databases (PostgreSQL/MySQL) for metadata and control‑plane services.
- NoSQL / in‑memory stores (e.g., Redis) for high volume or low latency workloads.
- Working knowledge of search and analytics engines for logs and traces.
- Understanding of data modeling, query patterns, and data lifecycle management.
- Experience with Kafka‑style distributed messaging systems.
- Designing and operating event‑driven telemetry ingestion pipelines.
- Understanding of throughput, backpressure, and reliability concerns in streaming systems.
- Strong hands‑on exposure to Docker and Kubernetes‑based platforms.
- Infrastructure automation using:
- Terraform (IaC).
- Ansible (deployment and configuration automation).
- CI/CD pipeline understanding and usage (e.g., GitLab CI or equivalent).
- Familiarity with Grafana, Kibana, or similar visualization platforms.
- Awareness of frontend technologies such as Angular (architecture‑level understanding).
- Experience with OpenStack, MAAS, Juju, or similar cloud/service orchestration platforms.
- Exposure to Canonical Observability Stack (COS) or equivalent platforms.
- Ability to perform code reviews for critical and performance‑sensitive components.
- Experience building or guiding proofs of concept (POCs) to validate architecture.
- Strong collaboration with Engineering Managers and senior engineers.
- Clear technical communication and mentoring capability.
Behaviours
- Customer First - Prioritize customer needs in every decision and action.
- Challengers - Challenge the status quo and bring innovative ideas to life.
- Committed - Own outcomes and deliver with integrity.
- Clear - Communicate openly and simply, ensuring alignment.
- Connected - Collaborate across teams to achieve shared goals.