Distributed Systems Architect: Bare-Metal and High-Scale

Cercli Tech Limited

Abu Dhabi

On-site

AED 450,000 - 650,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cercli Tech Limited is seeking a Distributed Systems Architect to own the architectural blueprints for on-premises infrastructure, including multi-node clustering, replication, and failover across Kafka/Redpanda, PostgreSQL (Patroni), and storage layers. This is a hands-on role with architecture-quality output driving hardware-failure resilience in air-gapped client environments.

You will map Azure PaaS components to open-source equivalents, assess feature gaps, validate hardware specs, and

Qualifications

  • 10+ years architecting and operating distributed infrastructure at TB to PB scale in production environments.
  • Expert-level Kafka or Redpanda clustering: partitioning strategy, ISR tuning, consumer group management, compaction policies, and operational lifecycle at scale.
  • Deep Linux systems engineering: kernel networking subsystems, storage fabrics, and NUMA-aware process binding.
  • Proven track record deploying HA database clusters without cloud load balancers: Patroni, Pacemaker, or equivalent in production.
  • Experience deploying software-defined storage (Ceph, MinIO) on physical hardware in production, including CRUSH maps, erasure coding, and performance tuning.

Responsibilities

  • Map open-source equivalents for Azure PaaS components to on-premise infrastructure.
  • Produce side-by-side equivalence assessments documenting feature gaps and migration risks.
  • Validate client hardware specs and air-gap security compliance before deployments.
  • Design multi-node clustering, rack-aware replication, and quorum topologies for bare-metal clusters.
  • Architect resilience protocols for hardware and switch failures in air-gapped environments.
  • Tune Linux kernel parameters and NVMe I/O schedulers to remove bottlenecks for petabyte-scale workloads.
  • Deliver implementation-ready blueprints and runbooks for Terraform and Ansible automation.
  • Maintain reference architectures for single-rack, multi-rack, and geographically distributed topologies.
  • Perform pre-deployment reviews of hardware specs and provide go/no-go remediation guidance.

Skills

Distributed systems architecture
Kafka/Redpanda clustering
Linux systems engineering
HA database clusters
Software-defined storage
DevOps blueprinting & runbooks

Tools

Kafka
Redpanda
Patroni
Pacemaker
Ceph
MinIO

Job description

Distributed Systems Architect: Bare-Metal and High-Scale

Full time

On-site

Abu Dhabi - United Arab Emirates

Distributed Systems Architect: Bare-Metal and High-Scale

You will own the architectural blueprints for Analog’s on-premises open-source infrastructure, designing the clustering, replication, and failover topologies across messaging, streaming, and database layers that our DevOps team then automates at scale across client deployments. This is a hands-on, implementation-grade architecture role where the quality of your output directly determines whether a physically isolated or air-gapped client environment survives hardware failure at petabyte scale.

What You’ll Do

Open-Source Infrastructure Mapping

  • Map Analog’s Azure PaaS components to open-source equivalents: Event Hub to Kafka/Redpanda, IoT Hub to EMQX, ADX to ClickHouse or Apache Druid, and Blob Storage to Ceph/MinIO
  • Produce side-by-side equivalence assessments documenting feature gaps, operational differences, and migration risk for each component transition
  • Validate client hardware specifications and assess compliance with air-gap security requirements prior to each deployment

Cluster and Topology Design

  • Design multi-node clustering, rack-aware replication, and quorum topologies for bare-metal Kafka/Redpanda and PostgreSQL (Patroni) clusters
  • Design resilience protocols for hardware and network switch failures in physically isolated or air-gapped client environments
  • Architect RocksDB state backend tuning for stateful Apache Flink workloads, including compaction strategy, block cache sizing, and write-ahead log configuration

Linux and Hardware Optimization

  • Tune Linux kernel parameters, NUMA bindings, network ring buffers, and NVMe I/O scheduler configurations to eliminate hardware bottlenecks under petabyte-scale write workloads
  • Define CPU affinity, IRQ balancing, and huge page configurations for latency-sensitive broker and database processes
  • Author reproducible benchmark harnesses to validate configuration changes against client hardware before production rollout

Blueprints and DevOps Enablement

  • Deliver precise, implementation-ready configuration blueprints and runbooks for Terraform and Ansible automation, documents the DevOps team can execute without architectural interpretation
  • Maintain a library of parameterized reference architectures covering single-rack, multi-rack, and geographically distributed bare-metal topologies
  • Conduct pre-deployment reviews of client hardware specs and provide go/no-go assessments with remediation guidance

What You’ll Bring

  • 10+ years architecting and operating distributed infrastructure at TB to PB scale in production environments
  • Expert-level Kafka or Redpanda clustering: partitioning strategy, ISR tuning, consumer group management, compaction policies, and operational lifecycle at scale
  • Deep Linux systems engineering: kernel networking subsystems (TCP buffer tuning, interrupt coalescing), storage fabrics, and NUMA-aware process binding
  • Proven track record deploying HA database clusters without cloud load balancers: Patroni, Pacemaker, or equivalent in production
  • Experience deploying software-defined storage (Ceph, MinIO) on physical hardware in production, including CRUSH maps, erasure coding, and performance tuning
  • Ability to produce implementation-ready blueprints and runbooks; your output must be directly actionable by a DevOps automation team without architectural interpretation

About Analog

Analog builds industrial intelligence infrastructure for the physical world. We deliver high-throughput data pipelines, real-time analytics, and edge-to-cloud connectivity for mission-critical environments where cloud dependency is not an option. Our clients operate in regulated, air-gapped, and physically demanding settings and they depend on us to get the infrastructure right from day one.

This is a full-time, on-site role.

Job details

Job type

Full time

On-site

Location

Abu Dhabi - United Arab Emirates

Department

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Systems Architect: Bare-Metal and High-Scale
Distributed Systems Architect: Bare-Metal and High-Scale

Analog • Abu Dhabi

On-site
AED 400,000 - 760,000
On-site role
Full-time employment
Principle Solutions Architect
Principle Solutions Architect

Cercli Tech Limited • Abu Dhabi

On-site
AED 280,000 - 420,000
Senior Bare-Metal, Petabyte-Scale Distributed Architect
Senior Bare-Metal, Petabyte-Scale Distributed Architect

Analog • Abu Dhabi

On-site
AED 400,000 - 760,000
On-site role
Full-time employment
IT Infrastructure Engineer
IT Infrastructure Engineer

Cercli Tech Limited • Abu Dhabi

On-site
AED 320,000 - 520,000
Culture
Impact
Growth
+1
Senior Backend Engineer
Senior Backend Engineer

Cercli Tech Limited • Abu Dhabi

On-site
AED 300,000 - 520,000
Senior Bare-Metal Distributed Systems Architect, High Scale
Senior Bare-Metal Distributed Systems Architect, High Scale

Cercli Tech Limited • Abu Dhabi

On-site
AED 450,000 - 650,000
IT Infrastructure Engineer
IT Infrastructure Engineer

Analog • Abu Dhabi

On-site
AED 180,000 - 300,000
Healthcare insurance
Education support
Generous leave benefits
Senior Research Engineer
Senior Research Engineer

Cercli Tech Limited • Abu Dhabi

On-site
AED 300,000 - 460,000
Principle Solutions Architect
Principle Solutions Architect

Analog • Abu Dhabi

On-site
AED 200,000 - 320,000
Principle Solutions Architect
Principle Solutions Architect

analogai • Abu Dhabi

On-site
AED 360,000 - 600,000