Senior DevOps Engineer

qualys

Pune District

On-site

INR 3,500,000 - 7,000,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Qualys is seeking a Senior DevOps Engineer, Observability to design, build, operate, and continuously improve a high-scale observability platform built around ClickHouse, HyperDX, OpenTelemetry, Kubernetes, and modern automation practices.

You will own production observability for a large-scale environment, implement telemetry pipelines, optimize ClickHouse, and mentor peers while collaborating with SRE, platform, and application teams to reduce incident time and improve visibility.

Qualifications

  • 6+ years of experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or Observability Engineering roles.
  • Strong production experience with ClickHouse at scale, including architecture and operations.
  • Hands-on experience with HyperDX integration for log search, tracing, dashboards, and troubleshooting.
  • Proficiency with OpenTelemetry and Collector configurations (receiver, processor, exporter, sampling, enrichment).
  • Working knowledge of Fluent Bit, Filebeat, Prometheus, Alertmanager, Grafana, and Kubernetes.
  • Automation experience with Jenkins, CI/CD, Terraform, Ansible, Consul, and Vault.
  • Strong understanding of Linux, networking, microservices, REST, gRPC, distributed systems, and high-availability.
  • Experience operating and troubleshooting high-volume production environments.

Responsibilities

  • Design, build, and operate scalable observability platforms using ClickHouse, HyperDX, OpenTelemetry, Kubernetes, Prometheus, Grafana, Alertmanager, Fluent Bit, and Filebeat.
  • Architect and optimize ClickHouse for high-volume observability workloads across logs, traces, and metrics.
  • Manage ClickHouse schemas, partitioning, TTLs, materialized views, and query optimization.
  • Develop and maintain HyperDX for log search, tracing, dashboards, and service analysis.
  • Build and maintain OpenTelemetry Collector pipelines for telemetry data processing.
  • Ensure reliable correlation across logs, metrics, and traces for faster troubleshooting.
  • Deploy observability infra on Kubernetes with scalability, HA, and capacity planning.
  • Automate deployment, configuration, upgrades, and onboarding using CI/CD, Terraform, and Ansible.
  • Collaborate with engineering, platform, SRE, and operations to improve coverage and incident response.
  • Participate in incident reviews, capacity planning, and remediation tracking; mentor junior engineers.

Skills

DevOps fundamentals
SRE/Platform engineering
Incident response
Mentoring
Automation mindset
Communication

Tools

ClickHouse
HyperDX
OpenTelemetry
Kubernetes
Prometheus
Grafana
Alertmanager
Fluent Bit
Filebeat
Terraform
Ansible
HashiCorp Consul
HashiCorp Vault
Jenkins

Job description

Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!

Role Overview

We are seeking a Senior DevOps Engineer, Observability to design, build, operate, and continuously improve a high-scale observability platform built around ClickHouse, HyperDX, OpenTelemetry, Kubernetes, and modern DevOps automation practices. This role is intended for a senior hands-on engineer who can take end-to-end ownership of observability infrastructure for production environments. The engineer will be responsible for building reliable telemetry pipelines for logs, metrics, and distributed traces, optimizing ClickHouse for large-scale observability workloads, operating HyperDX for troubleshooting and application performance analysis, and partnering with application, platform, and SRE teams to improve production visibility, reliability, and incident response. The ideal candidate is deeply technical, operationally disciplined, automation-oriented, and comfortable working in high-volume, production-critical environments. Senior-level expectation: This role requires ownership beyond task execution, including technical judgment, production accountability, automation-first delivery, mentoring, and clear communication during incidents and escalations.

Key Responsibilities
  • Design, build, and operate scalable observability platforms using ClickHouse, HyperDX, OpenTelemetry, Kubernetes, Prometheus, Grafana, Alertmanager, Fluent Bit, and Filebeat.
  • Architect and optimize ClickHouse for high-volume observability workloads, including logs, traces, metrics, and telemetry analytics.
  • Design and manage ClickHouse schemas, partitioning strategies, ordering keys, TTLs, materialized views, retention policies, storage efficiency, and query optimization.
  • Build, operate, and improve HyperDX for log search, distributed tracing, service analysis, dashboards, telemetry correlation, troubleshooting, and root-cause analysis.
  • Build and maintain scalable OpenTelemetry Collector pipelines for collecting, processing, enriching, filtering, sampling, and routing telemetry data.
  • Implement reliable correlation across logs, metrics, and traces to support faster application troubleshooting, service dependency analysis, and incident resolution.
  • Design resilient telemetry pipelines with batching, queuing, retries, backpressure handling, sampling, rate limiting, and cardinality controls.
  • Deploy and operate observability infrastructure on Kubernetes, with focus on scalability, high availability, capacity planning, resiliency, and operational safety.
  • Automate infrastructure deployment, configuration management, platform upgrades, application onboarding, and recurring operational tasks using DevOps best practices.
  • Build and maintain CI/CD workflows using Jenkins, infrastructure automation using Terraform and Ansible, and service discovery or secrets management integrations using HashiCorp Consul and Vault.
  • Partner with application engineering, platform engineering, SRE, and operations teams to troubleshoot production issues, improve observability coverage, and reduce mean time to detect and resolve incidents.
  • Analyze production performance issues across infrastructure, applications, telemetry pipelines, and ClickHouse queries, and drive corrective actions to closure.
  • Define and improve standards for telemetry instrumentation, log quality, metric hygiene, trace propagation, dashboard design, alert quality, and production readiness.
  • Participate in incident response, post-incident reviews, capacity planning, operational reviews, and remediation tracking for observability services.
  • Mentor junior engineers, review designs and automation changes, and raise the overall technical and operational maturity of the team.
Required Qualifications
  • 6+ years of experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or Observability Engineering roles.
  • Strong production experience with ClickHouse, including architecture, administration, schema design, performance tuning, query optimization, troubleshooting, retention management, and operations at scale.
  • Hands-on experience deploying and operating HyperDX, including integration with ClickHouse and usage for log search, distributed tracing, dashboards, service analysis, and troubleshooting.
  • Strong experience with OpenTelemetry and OpenTelemetry Collector, including receiver, processor, exporter, sampling, enrichment, batching, and routing configurations.
  • Strong working knowledge of Fluent Bit, Filebeat, Prometheus, Alertmanager, Grafana, ClickHouse, HyperDX, and OpenTelemetry.
  • Strong DevOps and automation experience with Jenkins, CI/CD pipelines, Ansible, Terraform, HashiCorp Consul, HashiCorp Vault, and Kubernetes.
  • Strong understanding of Linux, networking, microservices, REST, gRPC, distributed systems, high-availability architectures, and production operations.
  • Experience operating and troubleshooting high-volu
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineer - Data & Scalability Platform (8-12 Yrs)
Software Engineer - Data & Scalability Platform (8-12 Yrs)

Cisco • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Sr Software Engineer
Sr Software Engineer

Nvidia • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Senior DevOps Engineer
Senior DevOps Engineer

Daimler AG • Bengaluru

On-site
INR 3,000,000 - 4,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA AI • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000