Principal Data Platform Engineer · Office-first (Bangalore) · Engineering

Base14

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Flexible hours
Learning budgets
Premium AI tools

Job summary

Base14 is building Scout, an observability platform, and we’re hiring a Principal Data Platform Engineer to own the telemetry data platform end to end. You’ll run ClickHouse, design HA, and scale ingestion from OpenTelemetry, Kafka, and Parquet to S3, with pipelines feeding PostgreSQL, Neo4j and Redis.

You’ll lead architectural decisions, automate with GitOps/Argo CD, mentor engineers, and share on-call rotations.

Qualifications

  • Production ClickHouse at scale and terabytes of data.
  • HA database design and operation experience.
  • Proficient in Go and Rust for production services.
  • Understanding MergeTree, sorting keys and partitions.
  • On-call ownership with root-cause focus.
  • Experience with AI-assisted development and automation.
  • Principal-level judgment for ambiguous scaling problems.

Responsibilities

  • Run ClickHouse in production across cloud and data centers.
  • Design HA architecture for the data platform.
  • Keep queries fast via schemas, keys, partitions.
  • Build ingestion that stays sub-second with OTel/Kafka/S3 Parquet.
  • Develop services that write to/read from ClickHouse and move data to PostgreSQL/Neo4j/Redis.
  • Shape telemetry for fast agent queries and trustworthiness.
  • Automate infrastructure with GitOps and Argo CD.
  • Share on-call and lead incident response.
  • Set technical direction and mentor engineers.
  • Automate your workflows with AI tools.

Skills

ClickHouse production
HA infrastructure
Go and Rust
Schema design
On-call ownership
AI-assisted coding
Leadership judgement

Tools

Kubernetes
PostgreSQL
Neo4j
Redis

Job description

We are building autonomous infrastructure, and it runs on telemetry data.

base14 builds Scout, an observability platform built on OpenTelemetry that covers logs, metrics, traces, APM and RUM. Every signal our customers send lands in our telemetry data platform. When a customer's production is failing, Scout is how they see what is happening, so our platform has to be up when theirs is down.

We are also building agents that use this telemetry, both detailed and summarised, to manage production infrastructure. Self-healing, autonomous infrastructure is a company goal, and those agents are only as good as the data platform underneath them.

As we onboard more customers, we are investing in systems that hold 99.99% uptime and sub‑second ingestion for gigabytes of telemetry data. We are hiring a Principal Data Platform Engineer to own that platform and set its technical direction.

The Role

You will own the telemetry data platform end to end. Data arrives through OpenTelemetry Collectors, Kafka, or Parquet written directly to S3, depending on the scale and type of pipeline. Raw data lives in ClickHouse. Summary data flows from there into OLTP stores, graph databases and caches that serve different parts of the product.

ClickHouse will take a good share of your time: running clusters, designing them for high availability, and keeping queries fast as volume grows. The platform is wider than one store, though. It runs from the ingestion pipelines through to the summary stores and caches that serve the product, and you own how those pieces fit together. This is an engineering role as much as an operations role. You will write the services, tooling and automation around the data as well as run it.

As a principal engineer, you make the architectural calls for the data platform, write them down, and raise the bar for the engineers around you.

You will share the on‑call rotation for the data platform. We expect you to treat every page as a defect and engineer it out of the system.

We expect you to automate your own workflow. You will use AI tools and LLMs to write code, generate tests, analyze slow queries, draft migrations and runbooks, and investigate incidents, so you can focus on hard technical problems.

What You'll Do
  • Run ClickHouse in production: Operate clusters across AWS, GCP, Azure and our data centers. Own sharding, replication, capacity planning, upgrades, backups and recovery.
  • Engineer for four nines: Design the high‑availability architecture for the data platform. Plan for node, zone and region failures, test those plans, and remove single points of failure.
  • Keep queries fast: Design schemas, sorting keys, partitioning, materialised views and TTL policies for logs, metrics and traces. Find and fix slow queries before customers notice them.
  • Build ingestion that keeps up: Evolve our pipelines across OTel Collectors, Kafka and Parquet on S3 so ingestion stays sub‑second as volume and customer count grow. Handle bursts and backpressure without losing data.
  • Write the services around the data: Build the services that write to and read from ClickHouse, and the jobs that move summary data into PostgreSQL, Neo4j and Redis.
  • Make the data usable by agents: Shape detailed and summary telemetry so our agents can query it quickly and trust what they get back when they act on production infrastructure.
  • Automate the infrastructure: All of our infrastructure is automated and delivered through GitOps with Argo CD. Every infrastructure change you make ships as a reviewed commit.
  • Share on‑call: Take part in the shared rotation for the data platform. Lead incident response when the data layer is involved, write the post‑mortem, and fix the root cause. Automate the Incident response for areas owned.
  • Set technical direction: Decide how the data platform evolves. Review designs, write down the trade‑offs, and mentor engineers working on the data path.
  • Automate your work: Use AI assistants and automation tools to write, test and debug code, and to speed up operational work.
What We Look For
Must Have
  • Production ClickHouse at scale: You have run at least one ClickHouse cluster in production holding terabytes of data. You have dealt with replication issues, resharding, upgrades and node failures on a live system.
  • High‑availability experience: You have designed and operated HA infrastructure for a database, ClickHouse at minimum. You know replication, failover, coordination with ClickHouse Keeper or ZooKeeper, and backup and restore from having done them.
  • A programmer first: You have built production services that use ClickHouse. Your experience goes well beyond writing queries and operating the database. You write clean, legible and fast code. Most of our data path is Go and Rust.
  • Schema and query depth: You understand the MergeTree engine family, how sorting keys and partitions affect reads and merges, and how to find a bottleneck using query logs and system tables.
  • Operational ownership: You are comfortable carrying on‑call for the systems you build, and you prefer fixing root causes to restarting things.
  • AI‑assisted execution: You use generative AI tools to write code faster and to handle routine engineering and operational tasks.
  • Principal‑level judgment: You can take an ambiguous scaling problem, choose a direction, explain the trade‑offs in writing, and bring other engineers along.

Years matter less to us than what you have run. People at this level typically have 10 or more years of engineering experience, but depth with ClickHouse in production counts for more than the number.

Strong Pluses
  • Agent‑building experience: You have built AI agents or LLM‑driven systems that take real actions, and you are excited about infrastructure that heals itself. This is a massive bonus for us.
  • Kubernetes: You have run stateful workloads on Kubernetes, including operators, persistent storage and rolling upgrades of databases. This is a big advantage here.
  • Observability background: You have worked on logs, metrics or traces at scale, or with OpenTelemetry, and you understand the shape of telemetry data.
  • More scale: You have managed multiple ClickHouse clusters, or clusters holding petabytes.
Technical Stack
  • Ingestion: OpenTelemetry Collector, Kafka, Parquet direct to S3
  • Summary & Serving: PostgreSQL, Neo4j, Redis, both self‑managed and cloud‑managed
  • Languages: Go and Rust, with Python, Ruby and TypeScript where needed
  • Infrastructure: Managed Kubernetes on AWS, GCP and Azure. Talos Linux on VMs in our data centers.
  • Delivery: Argo CD, GitOps, 100% automated infrastructure
  • AI & Automation: AI coding assistants, LLM APIs and automation tooling
Why Join Us
  • Build autonomous infrastructure: We are building agents that read telemetry and manage production systems on their own. The platform you own is what they see and reason with, so your work decides how far self‑healing infrastructure can go.
  • Hard problems at real scale: Four nines of uptime and sub‑second ingestion on a multi‑cloud telemetry platform is a problem few engineers get to own.
  • Real ownership: You set the direction for the data platform and your decisions shape the product and the company.
  • Equity: Every employee owns a stake in the company.
  • Strong peers: You will work alongside engineers who have scaled massive infrastructure.
  • Fast feedback: Your work shows up in customer‑facing latency and uptime within days.
What We Offer
  • Competitive salary and equity.
  • Collaboration with experienced founders.
  • Hardware and software of your choice, including premium AI development tools.
  • Learning budgets and flexible hours focused on work output.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Engineer · Office-first (Bangalore) · Engineering
Product Engineer · Office-first (Bangalore) · Engineering

Base14 • Bengaluru

On-site
INR 900,000 - 1,300,000
Senior Data Engineer
Senior Data Engineer

super.money • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Database Reliability Engineer — ClickHouse / OLAP
Senior Database Reliability Engineer — ClickHouse / OLAP

VuNet Systems • Bengaluru

On-site
INR 900,000 - 1,300,000
Tech Lead at Saleshandy
Tech Lead at Saleshandy

Ikigai Infotech LLP (SalesHandy) • Ahmedabad District

On-site
INR 1,500,000 - 2,100,000
Lead ClickHouse Database Administrator
Lead ClickHouse Database Administrator

Persistent • Pune District

Hybrid
INR 4,000,000 - 7,000,000
Competitive salary
Education & certifications support
Cutting-edge technologies
+3
Senior ClickHouse Database Engineer
Senior ClickHouse Database Engineer

Persistent Systems Limited • Pune District

On-site
INR 4,000,000 - 6,000,000
Hybrid work culture
Flexible hours
Long Service awards
+1
Lead ClickHouse Database Administrator
Lead ClickHouse Database Administrator

Persistent Systems • Pune District

On-site
INR 4,000,000 - 7,000,000
Hybrid work
Flexible hours
Higher education sponsorship
+1
Senior Consulting Engineer - India
Senior Consulting Engineer - India

ClickHouse • India

Hybrid
INR 4,000,000 - 6,000,000
Flexible work environment
Healthcare contributions
Stock options
+3
Lead Data Engineer
Lead Data Engineer

Terrantic Inc. • India

On-site
INR 7,662,835 - 11,494,252
Competitive salary
Meaningful equity
Flexible remote work
Team Lead, Data Engineer
Team Lead, Data Engineer

Affinity Global • Maharashtra

On-site
INR 4,000,000 - 6,500,000