Principal Elastic Platform Owner & Onsite Leader

Softility

Atlanta (GA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A technology services company in Atlanta, GA is seeking a technical lead to manage the Elastic Stack and Kubernetes infrastructure. The ideal candidate will have 10+ years of experience, focusing on onsite operations, architectural reliability, and incident management. Responsibilities include overseeing platform stability, defining logging strategies, and mentoring engineering teams. Strong expertise in cloud-based environments and infrastructure-as-code practices is essential, along with experience in large-scale corporate operations.

Qualifications

  • 10+ years of experience in technical leadership roles.
  • Expertise in Elastic Stack and Kubernetes for managing enterprise applications.
  • Strong background in incident management and SRE practices.

Responsibilities

  • Lead onsite operations for platform stability and manage engineering backlogs.
  • Oversee Elastic Stack architecture deployed on Elastic Cloud.
  • Design cluster topology for performance and cost-efficiency.

Skills

Elastic Stack
Kubernetes
Infrastructure-as-Code
Incident Management
CI/CD pipelines

Job description

Employment Type : Full-time

Experience : 10+ Years

Required Skills :

Technical Skills

  • Elastic Expert : 5+ years of production experience with the Elastic Stack (Elasticsearch, Kibana, Logstash, Beats).
  • Kubernetes Mastery : 3+ years managing Elastic Cloud on Kubernetes (ECK) or similar operators on enterprise K8s distributions (Anthos, GKE, or EKS).
  • Ingest & Pipelines: Deep knowledge of log ingest architectures, index templates, sharding strategies, and cluster tuning.

Leadership & Experience

  • Onsite Leadership : Proven ability to run day-to-day operations, manage technical rosters, and lead cross-functional troubleshooting sessions.
  • Enterprise Scale : Experience supporting large-scale platforms (e.g., ETL jobs, microservices, anomaly detection) in a complex corporate environment.
  • Incident Management : Familiarity with SRE practices, including incident response (P1-P3), MTTR tracking, and root cause analysis.

Preferred Skills

  • Experience with Infrastructure-as-Code (Terraform, Helm) and CI/CD pipelines (Jenkins, GitLab).
  • Knowledge of migration strategies between legacy logging platforms (e.g., Splunk) and Elastic.
Responsibilities
Platform Architecture & Onsite Leadership
  • Onsite Operational Lead : Act as the primary point of contact for platform stability. Manage daily stand‑ups, prioritize the engineering backlog, and coordinate between offshore and onsite teams.
  • Architecture & Reliability : Own the Elastic Stack (Elasticsearch, Kibana, Ingest components) deployed on Elastic Cloud on Kubernetes (ECK).
  • Capacity Planning : Design cluster topology (nodes, roles, zones) and manage resource quotas (CPU, Heap, Disk) to ensure cost‑efficiency and performance.
  • SLO Management : Define and track Service Level Objectives (SLOs) for ingestion latency, search availability, and data retention.
Logging Strategy & Data Modeling
  • Standardization : Define enterprise logging and index templates, including field conventions (service, environment, tenant, correlation IDs) to ensure reliable event correlation across the observability stack.
  • Schema Design : Work with application teams to implement standardized mappings and Index Lifecycle Management (ILM) policies (Hot/Warm/Cold tiers, rollover, and retention).
  • Data Quality : Own ingest patterns using Filebeat, Fluent Bit, Logstash, or Elastic Agent. Design parsing pipelines (JSON, Grok), enrichment logic, and dead‑letter queue strategies.
Kubernetes & Infrastructure Operations
  • Cluster Management : Oversee daily health across Kubernetes‑based clusters (e.g., Anthos/GKE). Resolve pod‑level issues such as CrashLoopBackOffs, memory spikes, and disk usage alerts.
  • Storage Operations : Lead the management and migration of Persistent Volume Claims (PVCs) for stateful sets, ensuring high availability during infrastructure upgrades.
  • Security & Governance : Implement RBAC for indices and Kibana spaces. Enforce data governance and retention policies based on data classification (Infra logs vs. App logs vs. Sensitive data).
Observability & Enablement
  • Self‑Service Enablement : Deliver Kibana dashboards and "Golden Queries" to support SRE and NOC teams in rapid incident triaging.
  • Documentation & Runbooks : Author and maintain operational runbooks, disaster recovery scenarios, and "Self‑Help" guides for onboarding new log sources.
  • Mentorship : Provide technical guidance and training sessions for platform engineers and application developers on effective search and logging practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Elastic Administrator / Operator
Elastic Administrator / Operator

Silicontek Inc • New York (NY)

On-site
USD 90,000 - 120,000
Senior Elasticsearch Database Admin
Senior Elasticsearch Database Admin

CatchProbe Intelligence Technologies • San Francisco (CA)

Remote
USD 130,000 - 180,000
Elastic Platform Lead & Onsite Architect
Elastic Platform Lead & Onsite Architect

Softility • Atlanta (GA)

On-site
USD 120,000 - 160,000
Director - IT Software Engineering
Director - IT Software Engineering

Elasticsearch B.V. • Northern (KY)

Hybrid
USD 150,000 - 210,000
Sr. Elastic Engineer
Sr. Elastic Engineer

ECS • Bedford (MA)

On-site
USD 180,000 - 210,000
Director, Platform Engineering
Director, Platform Engineering

CyberMaxx, Inc. • Linthicum (MD)

On-site
USD 150,000 - 190,000
Flexible PTO
401k with company match
Medical, Dental & Vision
+5
Director, Platform Engineering
Director, Platform Engineering

CyberMaxx • Linthicum (MD)

On-site
USD 180,000 - 260,000
Flexible PTO
401k with match
Medical, Dental and Vision Coverage
+5
Sr. Elastic Engineer
Sr. Elastic Engineer

ECS • Shiloh (IL)

On-site
USD 100,000 - 130,000
Principal Software Engineer I - Serverless - Platform Control Plane Elastic
Principal Software Engineer I - Serverless - Platform Control Plane Elastic

Neura Market • Northern (KY)

Hybrid
USD 160,000 - 253,000
Stock options
401k matching
Health coverage
+4
Principal Software Engineer I - Serverless - Platform Control Plane
Principal Software Engineer I - Serverless - Platform Control Plane

Elastic • Mountain View (CA)

On-site
USD 160,000 - 253,000
Health coverage
Stock program (company equity)
401k matching
+3