DataOps Engineer

Woongjin, Inc

Englewood Cliffs (NJ)

On-site

USD 99,000 - 121,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Woongjin, Inc. is seeking a mid‑level engineer to build and operate a Docker‑based data platform using Apache Iceberg as the lake‑house format. You will own end‑to‑end delivery pipelines, monitoring, security, and incident response to ensure reliable scale.

The role covers Iceberg operations, Docker‑based image creation and testing, ETL/ELT pipeline development, and CI/CD automation. You’ll work with Spark/Flink/Presto, Python/Ansible, and observability tools while collaborating with security

Qualifications

  • Bachelor’s degree in Computer Science, IT, Data Engineering, or related field.
  • 5+ years of hands‑on experience building and operating large‑scale data platforms (lake‑house, data‑warehouse, or big‑data ecosystems).
  • Proven production experience with Apache Iceberg (table creation, partition management, schema evolution, catalog integration).
  • Strong Docker skills: multi‑stage builds, Docker‑Compose testing, routine image security scanning.
  • Experience with at least one major data‑processing engine (Spark, Flink, or Presto/Trino) and its connection to Iceberg tables.
  • Proficiency in Python and/or Ansible for automating infrastructure and platform tasks.
  • Experience building CI/CD pipelines that include Docker linting, vulnerability scanning, and automated deployment of data‑pipeline code.
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, Loki) and ability to create useful alerts and dashboards.
  • Ability to respond to incidents, write clear root‑cause analysis reports, and contribute to post‑mortem actions.
  • Willingness to participate in an on‑call rotation as a first‑line responder.
  • Availability to work on‑site in New Jersey for the initial assignment and relocate to Dallas by October 2026.

Responsibilities

  • Iceberg operations: support tables, manage schema changes, partitions, snapshot retention, and keep the catalog synchronized.
  • Docker image creation & testing: write multi‑stage Dockerfiles for Spark/Flink/Presto, run local test environments with Docker‑Compose, and conduct vulnerability scans.
  • Data pipeline development: build ETL/ELT jobs that ingest raw data and write to Iceberg tables; add simple streaming components using Kafka, Pulsar, or Kinesis when needed.
  • CI/CD automation: configure pipelines to lint Dockerfiles, scan images, version Iceberg metadata, and deploy pipelines without downtime.
  • Automation with Ansible/Python: script cluster provisioning, catalog configuration, vacuum/compaction, and other routine housekeeping tasks.
  • Observability: instrument services with OpenTelemetry, Prometheus, Grafana, and Loki; create dashboards showing pipeline latency, resource usage, table health, and error rates; set up basic alerts.
  • SLA monitoring: measure data freshness, job success rates, and query response times against targets and report deviations.
  • Incident response: join on‑call rotation, diagnose and resolve pipeline failures or Iceberg metadata issues; write root‑cause analyses and suggest improvements.
  • Security & compliance support: help enforce image signing, mTLS, IAM roles, and bucket policies; coordinate with security for GDPR, HIPAA, or ISO 27001 requirements.
  • Knowledge sharing: keep internal docs up to date and run short demos on Iceberg, Docker, and automation techniques.

Skills

Python
Ansible
CI/CD
Incident response
Docker

Education

Bachelor’s degree in Computer Science or related field
Master’s degree a plus

Tools

Docker
Docker-Compose
Spark
Flink
Presto/Trino
Apache Iceberg
Hive Metastore
AWS Glue
Nessie
Prometheus
Grafana
OpenTelemetry
Loki
Kafka
Pulsar
Kinesis

Job description

Job Description

We are looking for a mid‑level engineer to build and operate a data platform that uses Apache Iceberg as the lake‑house table format and Docker‑based micro‑services (Spark, Flink, Presto, etc.). you will own the end‑to‑end delivery pipeline, monitoring, security, and incident response, ensuring the platform runs reliably at scale.

Key Responsibilities
  • Iceberg operations: support tables, manage schema changes, partitions, snapshot retention, and keep the catalog (Hive Metastore, AWS Glue, Nessie, …) synchronized.
  • Docker image creation & testing: write multi‑stage Dockerfiles for Spark/Flink/Presto, run local test environments with Docker‑Compose, and conduct vulnerability scans (Trivy, Snyk, …).
  • Data pipeline development: build ETL/ELT jobs that ingest raw data and write to Iceberg tables; add simple streaming components using Kafka, Pulsar, or Kinesis when needed.
  • CI/CD automation: configure pipelines (GitHub Actions, GitLab CI, Azure DevOps, …) to lint Dockerfiles, scan images, version Iceberg metadata, and deploy pipelines without downtime.
  • Automation with Ansible/Python: script cluster provisioning, catalog configuration, vacuum/compaction, and other routine housekeeping tasks.
  • Observability: instrument services with OpenTelemetry, Prometheus, Grafana, and Loki; create dashboards showing pipeline latency, resource usage, table health, and error rates; set up basic alerts.
  • SLA monitoring: measure data freshness, job success rates, and query response times against agreed‑upon targets and report deviations.
  • Incident response: join the on‑call rotation, perform first‑line diagnosis and resolution of pipeline failures, Iceberg metadata issues, or container crashes; write concise root‑cause analyses and suggest improvements.
  • Security & compliance support: help enforce image signing, mTLS, IAM roles, and bucket policies; collaborate with the security team to meet GDPR, HIPAA, or ISO 27001 requirements.
  • Knowledge sharing: keep internal documentation up to date and run short tech demos or brown‑bag sessions on Iceberg, Docker best practices, and automation techniques.
Qualifications
  • Bachelor’s degree in Computer Science, IT, Data Engineering, or a related field (Master’s a plus).
  • 5+ years of hands‑on experience building and operating large‑scale data platforms (lake‑house, data‑warehouse, or big‑data ecosystems).
  • Proven production experience with Apache Iceberg (table creation, partition management, schema evolution, catalog integration).
  • Strong Docker skills: multi‑stage builds, Docker‑Compose testing, routine image security scanning.
  • Experience with at least one major data‑processing engine (Spark, Flink, or Presto/Trino) and its connection to Iceberg tables.
  • Proficiency in Python and/or Ansible for automating infrastructure and platform tasks.
  • Experience building CI/CD pipelines that include Docker linting, vulnerability scanning, and automated deployment of data‑pipeline code.
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, Loki) and ability to create useful alerts and dashboards.
  • Ability to respond to incidents, write clear root‑cause analysis reports, and contribute to post‑mortem actions.
  • Willingness to participate in an on‑call rotation as a first‑line responder.
  • Availability to work on‑site in New Jersey for the initial assignment and relocate to Dallas by October 2026.
Preferred Qualifications
  • Experience with cloud‑native data services on AWS, Azure, or GCP (EMR, Dataproc, Synapse, etc.).
  • Familiarity with other lake‑house formats such as Delta Lake or Apache Hudi and ability to evaluate trade‑offs against Iceberg.
  • Knowledge of streaming platforms (Kafka, Pulsar, Kinesis) and real‑time processing patterns.
  • Relevant certifications (Databricks Lakehouse Associate, Google Professional Data Engineer, AWS Certified Data Analytics – Specialty, etc.).
  • Background supporting data platforms in regulated industries (pharma, finance, healthcare) and understanding of associated compliance frameworks.
Additional Information

All your information will be kept confidential according to EEO guidelines.

*** NO C2C ***

Compensation

$110,000-$110,000 per year

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DataOps Engineer
DataOps Engineer

SBT Global, Inc. • Englewood Cliffs (NJ)

On-site
USD 140,000 - 220,000
Apache Iceberg Engineer
Apache Iceberg Engineer

Smart IT Frame LLC • Sunnyvale (CA)

On-site
USD 90,000 - 150,000
Data Engineer
Data Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 90,000 - 120,000
DataOps Engineer: Iceberg Lakehouse Platform Lead
DataOps Engineer: Iceberg Lakehouse Platform Lead

Woongjin, Inc • Englewood Cliffs (NJ)

On-site
USD 110,000
Data Engineer
Data Engineer

Eliassen Group • Dallas (TX)

Hybrid
Medical benefits
Dental benefits
401k with matching
+1
Data Engineer
Data Engineer

Blutic • Dallas (TX)

Hybrid
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Peyton Resource Group • Houston (TX)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Tavant • Houston (TX)

On-site
USD 110,000 - 160,000
Sr. Staff Software Engineer - Apache Iceberg
Sr. Staff Software Engineer - Apache Iceberg

Cloudera • Washington

On-site
USD 184,000 - 230,000
Generous PTO
Unplugged days
Flexible WFH
+6
Data Pipeline Engineer
Data Pipeline Engineer

Cylake Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 250,000
Comprehensive benefits package
Competitive compensation