Lead Principal Core Infrastructure Engineer

Oracle

Bengaluru

On-site

INR 3,500,000 - 7,000,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Oracle in Bengaluru, India seeks a Senior Network Observability Architect to lead the network monitoring and observability platform. You will define the architecture, ensure cloud-scale telemetry, and mentor engineers across teams.

Responsibilities include driving telemetry pipelines, SLOs, and performance improvements while collaborating with network, infrastructure, and service teams to detect and resolve issues, ensuring reliability and security.

Qualifications

  • Experience building observability or telemetry platforms for large-scale cloud or network infrastructure.
  • Strong software engineering in Java or Go, and experience with distributed systems.
  • Ability to mentor engineers and lead cross-team initiatives in a large organization.

Responsibilities

  • Define architecture and technical direction for OCI’s network monitoring and observability platform.
  • Architect distributed systems that collect, process, store, and visualize telemetry at cloud scale.
  • Drive telemetry ingestion, streaming pipelines, and alerting strategies for high-volume data.
  • Establish SLOs and standards for availability, latency, and data freshness.
  • Mentor senior engineers and collaborate with network, infrastructure, and service teams.

Skills

Distributed Systems
Observability
Streaming & Data Processing
Networking
Network Telemetry
Programming
Reliability Engineering
Automation
Security

Tools

Kubernetes
Prometheus
Grafana
Kafka
Flink
Java/Go

Job description

Job Description

Provides senior technical leadership for OCI’s network monitoring and observability platform, responsible for the systems that provide visibility into the health, performance, and behavior of OCI network infrastructure.

Job Description

Provides senior technical leadership for OCI’s network monitoring and observability platform, responsible for the systems that provide visibility into the health, performance, and behavior of OCI network infrastructure.

Architects highly scalable, reliable, and efficient distributed systems for collecting, processing, storing, and analyzing large volumes of network telemetry. Defines technical strategy and engineering standards for telemetry pipelines, real-time stream processing, metrics, alerting, visualization, and network health analytics.

Drives architecture for systems that must operate continuously at cloud scale, with strong requirements for availability, data integrity, latency, scalability, and operational efficiency. Identifies systemic performance and reliability constraints and leads solutions across multiple services and engineering teams.

Serves as a senior technical authority for network observability and distributed systems, leading complex architectural and production issues, influencing cross-organizational technical decisions, and mentoring senior engineers. Partners with network engineering, infrastructure, and service teams to continuously improve OCI’s ability to detect, understand, and resolve network problems.

Responsibilities
Network Observability Architecture
  • Define the architecture and technical direction for OCI’s network monitoring and observability platform.
  • Architect distributed systems that collect, process, aggregate, store, query, and visualize network telemetry at cloud scale.
  • Design scalable telemetry ingestion and streaming architectures capable of handling high-volume and high-cardinality data.
  • Define strategies for metrics, events, alerts, topology, network state, and other signals required to understand network health.
  • Drive architectures that enable rapid detection, correlation, diagnosis, and isolation of network failures and performance degradation.
  • Establish standards for telemetry quality, completeness, freshness, accuracy, retention, and availability.
Scalability & Distributed Systems
  • Lead the design of horizontally scalable and elastic systems supporting continued OCI infrastructure and traffic growth.
  • Identify performance, throughput, latency, storage, and scalability bottlenecks across telemetry pipelines and drive systemic improvements.
  • Architect high-throughput streaming and event-processing systems with appropriate partitioning, backpressure, buffering, aggregation, and failure-handling strategies.
  • Define resilient state-management, replication, synchronization, and recovery strategies for distributed monitoring systems.
  • Evaluate architectural trade-offs involving consistency, availability, latency, durability, and cost.
Reliability & Operational Excellence
  • Establish SLOs and engineering standards for availability, durability, latency, data freshness, and correctness of the observability platform.
  • Architect fault-tolerant systems that continue operating through infrastructure failures, network partitions, service disruptions, and software upgrades.
  • Define KPIs and telemetry that measure the health of the monitoring platform itself and identify gaps or blind spots in observability.
  • Lead diagnosis and resolution of complex production issues spanning networking, distributed systems, telemetry pipelines, and infrastructure.
  • Serve as a senior technical escalation point for critical incidents and drive root-cause analysis and systemic corrective actions.
  • Design systems and operational practices that support automated deployment, upgrade, rollback, recovery, and minimal customer-visible disruption.
Network Monitoring & Analytics
  • Define approaches for monitoring large-scale Layer 2 and Layer 3 network infrastructure and identifying changes in network state, topology, reachability, performance, and device health.
  • Drive correlation of telemetry across devices, network layers, and infrastructure services to improve fault localization and reduce time to detection and resolution.
  • Develop approaches for identifying abnormal network behavior, telemetry gaps, capacity risks, and emerging infrastructure failures.
  • Partner with network engineering teams to translate network behavior and operational requirements into scalable monitoring capabilities.
  • Improve signal quality and actionable alerting while reducing noise and unnecessary operational load.
Technical Leadership
  • Set technical direction and influence architecture decisions across Network Monitoring and dependent OCI organizations.
  • Lead complex and ambiguous initiatives spanning multiple systems and engineering teams.
  • Establish architectural patterns, engineering standards, and technical best practices for large-scale observability systems.
  • Mentor senior engineers and provide technical guidance on distributed systems, networking, reliability, and observability.
  • Evaluate emerging technologies and engineering approaches and drive adoption where they materially improve scale, reliability, performance, or operational efficiency.
  • Contribute to technical talent development through senior-level interviewing, candidate assessment, mentoring, and knowledge sharing.
Skills & Technologies

The candidate should have deep expertise in several of the following areas and sufficient breadth to lead architecture across the complete observability stack:

  • Distributed Systems & Cloud: Large-scale distributed systems, cloud infrastructure, Kubernetes, service-oriented architectures, distributed state management, high availability, fault tolerance, capacity planning, and performance engineering.
  • Observability: Metrics and telemetry architecture, Prometheus, Grafana, time-series data, high-cardinality metrics, alerting, dashboards, SLOs, telemetry pipelines, and monitoring large distributed environments.
  • Streaming & Data Processing: Kafka, Flink or equivalent technologies; real-time stream processing, event-driven architectures, partitioning, aggregation, backpressure, data pipelines, and large-scale telemetry processing.
  • Networking: Strong understanding of L2/L3 networking, routing and switching, network topology, BGP, LLDP, SNMP, gNMI, network failure modes, and network performance troubleshooting.
  • Network Telemetry: Experience designing or operating systems that collect and correlate telemetry from large fleets of network devices using streaming telemetry, counters, events, protocol state, and device health information.
  • Programming: Strong software engineering expertise in Java, Go, or comparable systems programming languages, with experience building highly concurrent, performance-sensitive production services.
  • Reliability Engineering: SLOs, availability and durability engineering, distributed failure handling, load shedding, throttling, rate limiting, incident response, root-cause analysis, and production readiness.
  • Automation: Infrastructure as Code, automated deployment and configuration management, safe rollout and rollback strategies, and operating large infrastructure fleets with minimal manual intervention.
  • Security: Secure multi-tenant cloud architectures, authentication and authorization, encryption, vulnerability remediation, and security considerations for infrastructure telemetry and management systems.

Experience building observability or telemetry platforms for large-scale cloud or network infrastructure is strongly preferred.

Qualifications

Career Level - IC5

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life‑saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation‑request_mb@oracle.com or by calling 1‑888‑404‑2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 4,000,000 - 9,000,000
Principal Network Developer
Principal Network Developer

Oracle • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle • Chennai District

On-site
INR 4,000,000 - 6,500,000
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle • Thiruvananthapuram

On-site
INR 1,800,000 - 2,400,000
Flexible medical
Life insurance
Retirement options
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle • Dadri

On-site
INR 4,000,000 - 6,000,000
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle • Ahmedabad District

On-site
INR 1,500,000 - 2,100,000
Core Infrastructure Engineer 2
Core Infrastructure Engineer 2

Oracle • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Network Developer
Senior Network Developer

Oracle • Thiruvananthapuram

On-site
INR 1,200,000 - 1,800,000
Director, Core Infrastructure Engineering
Director, Core Infrastructure Engineering

Oracle • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Lead Principal Systems Software Engineer
Lead Principal Systems Software Engineer

Oracle • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Competitive benefits
Volunteer programs