Senior Data Engineer

8020rei

Bogotá

Presencial

COP 291.658.564 - 453.691.100

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Competitive Base Compensation
Profit Share Bonus
Flex PTO (up to 26 days per year)
Home Office Upgrade Bonus
HMO Bonus
Full-time Remote Work
Growth and team-building opportunities
Budget for skill development

Descripción de la vacante

8020REI is seeking a Senior Data Engineer to own the design, reliability, and cost-efficiency of our AWS data platform. Build and optimize PySpark ETL/ELT pipelines, operate a medallion lakehouse, and manage data-serving APIs with a focus on production reliability and revenue impact.

You will partner with Data Science to turn models into production systems, enforce data quality standards, and own cloud cost and infrastructure as code in a fast-growing real estate tech environment.

Formación

  • 4+ years of hands-on data engineering with production ownership.
  • Advanced PySpark performance tuning on large-scale workloads.
  • Strong Python with production-grade coding standards.
  • Deep AWS experience across EMR, Glue, Lambda, S3, Athena.
  • Advanced SQL for data warehousing and lakehouse querying.
  • Experience with open table formats (Hudi preferred).
  • Data quality gates, monitoring, and QA processes.
  • Infrastructure as Code in CI/CD workflows.
  • English and Spanish proficiency.
  • Bachelor’s degree in CS/SE or related field.

Responsabilidades

  • Own end-to-end Big Data pipelines on EMR and Glue.
  • Operate and evolve lakehouse on S3 with Hudi and manage partitions.
  • Automate orchestration with Step Functions, EventBridge, and Lambda.
  • Enforce data quality with QA audits, gating, and monitoring.
  • Manage production databases (Aurora PostgreSQL, DynamoDB) and APIs.
  • Own cloud cost optimization for the data platform.
  • Ship infrastructure as code via Terraform/CloudFormation in CI/CD.
  • Collaborate with Data Science to productionize models.
  • Document runbooks and data dictionaries.

Conocimientos

Advanced PySpark
Strong Python
Deep AWS
Advanced SQL
Lakehouse experience
Data quality mindset
Infrastructure as Code
English and Spanish
4+ years of data engineering

Educación

Bachelor’s degree in Computer Science, Data Engineering, or related field

Herramientas

Terraform
CloudFormation
AWS EMR
Airflow/Step Functions
Aurora PostgreSQL
DynamoDB

Descripción del empleo

What is 8020REI?

8020REI is a pioneering B2B data and SaaS platform that empowers professional real estate investors to generate consistent leads from Cold Call, SMS, and Direct Mail without spending more. We leverage AI, machine learning, and advanced analytics to identify homeowners likely to sell at a discount, and we use a proven strategy to maximize our clients' ROI. As part of our network, we also operate alongside three sister companies: DMForce, 8020Recruit, and 8020CRM, each contributing specialized services to the real estate sector.

Core Values at 8020REI:
  • Find a Better Way: We dig until we understand the real problem, then improve or innovate to solve it. Focused on the 20% that drives the biggest result.

  • Have Each Other's Back: We work as a team. We support, step in, and go the extra mile to create a great experience for each other and our clients.

  • Honor Your Word: We honor our promises, hold ourselves accountable, and do the right thing even when nobody’s looking.

  • Own Your Growth: We're obsessed with personal growth and take pride in the work we deliver. Intrinsically motivated, we hold ourselves to a high standard and never coast.

  • Bring Your A-Game: We take care of our recovery, health, and family time, and that’s why we show up focused, energized, and ready to perform at our best.

Role Overview

We are looking for a Senior Data Engineer to own the design, reliability, and cost-efficiency of our AWS data platform. You will build and optimize the PySpark/EMR pipelines that power our deal-scoring engine (Apollo), our property valuation product (BestEstimate), and our permits and roofing intelligence lines; harden our Hudi-based medallion lakehouse; and operate the data-serving APIs our clients and internal teams depend on.

This is a hands‑on senior role with real ownership: you will make architecture decisions, carry the cloud budget for the data platform, enforce our data quality standard, and partner directly with Data Science to turn models into reliable production systems. If you like nationwide‑scale data, pragmatic engineering, and seeing your pipelines drive real revenue, this seat is for you.

Key Responsibilities:
  • Own Big Data pipelines end to end. Design, build, and optimize PySpark ETL/ELT pipelines on Amazon EMR and AWS Glue that process nationwide, county‑partitioned property data (First American, BuildZoom permits, market comps) on daily and monthly cadences.

  • Run our lakehouse. Operate and evolve our Bronze → Silver → Gold data lake on S3 with Apache Hudi and the AWS Glue Data Catalog, queried through Athena, including schema contracts, partitioning strategy, compaction, and performance tuning.

  • Orchestrate and automate. Build reliable orchestration with AWS Step Functions, EventBridge, and Lambda; make reruns, backfills, and failure recovery boring and documented.

  • Enforce data quality. Implement and extend our Data QA Audit Standard: layer contracts, write‑audit‑publish gating, quarantine flows, drift monitoring, and actionable Slack alerting, so bad data never reaches a client list.

  • Operate production databases and APIs. Manage Aurora PostgreSQL and DynamoDB workloads, and run data‑serving APIs (API Gateway, SQS‑backed async workers), such as our Address Resolution Service, to production SLAs with dashboards and runbooks.

  • Own cloud cost. Monitor, report, and reduce the AWS data‑platform bill (EMR cluster sizing, Glue/Lambda usage, S3 lifecycle, Athena scan costs) as a first‑class engineering responsibility.

  • Ship infrastructure as code. Define infrastructure with Terraform and CloudFormation, delivered through GitHub Actions CI/CD with tests, linting, and coverage gates, we run a disciplined PR, branch‑policy, and code‑review culture.

  • Partner with Data Science. Build the feature pipelines, training datasets, and serving paths behind our ML scoring and valuation models (scikit‑learn/XGBoost‑family stack), and co‑own the handoff contracts between DS and DE.

  • Document like a pro. Maintain runbooks, architecture docs, and data dictionaries (Confluence) so any teammate can operate what you build.

Qualifications:
  • 4+ years of hands‑on data engineering with large‑scale distributed data systems and a track record of production ownership (not just development).

  • Advanced PySpark performance tuning, partitioning strategy, and cost‑aware cluster sizing on real workloads (EMR or equivalent).

  • Strong Python clean, tested, production‑grade code (we use pytest, ruff, mypy, and coverage gates in CI).

  • Deep AWS experience EMR, Glue, Lambda, S3, Athena, Step Functions, EventBridge, IAM, and VPC networking; comfort operating (not just deploying to) these services.

  • Advanced SQL complex analytical queries, query optimization, and data modeling on both a warehouse/lake engine (Athena/Presto) and PostgreSQL.

  • Lakehouse experience hands‑on production work with at least one open table format (Apache Hudi strongly preferred; Iceberg or Delta Lake also valued) and medallion‑style architecture.

  • Data quality mindset experience implementing validation, quality gates, monitoring, and incident response for production data.

  • Infrastructure as Code Terraform and/or CloudFormation in a CI/CD workflow.

  • English and Spanish professional working proficiency in both (B2+).

  • Bachelor’s degree in Computer Science, Systems Engineering, Data Engineering, or equivalent practical experience.

Nice to Have:
  • Real estate, property, or geospatial data experience (county/FIPS‑partitioned datasets, address standardization, parcel/permit data).

  • Building or operating public/internal data APIs (API Gateway, SQS, DynamoDB caching, SLAs).

  • Observability tooling (CloudWatch, Grafana dashboards, structured logging).

  • ML‑adjacent engineering: feature pipelines, training‑data reproducibility, model‑serving data paths.

  • Modern Python tooling (uv, Poetry) and monorepo/template‑driven repo governance.

  • Agile/SCRUM experience and a habit of writing documentation others actually use.

Why Join Us?

At 8020REI, you’ll play a key role in shaping the future of data analytics for the real estate investment industry. As a team member, you’ll have the opportunity to lead initiatives, work with cutting‑edge technologies, and collaborate with a dynamic team passionate about innovation and results.

What We Offer:
  • Competitive Base Compensation

  • Profit Share Bonus

  • Flex PTO (up to 26 days per year)

  • Home Office Upgrade Bonus

  • HMO Bonus

  • Full‑time Remote Work

  • Opportunity for growth and team‑building potential

  • Ongoing support and budget to develop new skills

Join us at 8020REI and contribute directly to revolutionizing the real estate investment landscape!

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine • Pereira

Presencial
COP 288.055.306 - 480.092.178
Growth opportunities
Remote work with flexible hours
Annual learning budget
+1
Senior Data Engineer Remote Production Pipelines
Senior Data Engineer Remote Production Pipelines

8020rei • Bogotá

Híbrido
COP 291.658.000 - 453.692.000
Competitive Base Compensation
Profit Share Bonus
Flex PTO (up to 26 days per year)
+5
Senior Data Engineer ID81743
Senior Data Engineer ID81743

AgileEngine • Bogotá

A distancia
COP 120.000.000 - 180.000.000
Growth without limits
Competitive compensation
Remote work 100%
+3
Senior Data Engineer ID81743
Senior Data Engineer ID81743

AgileEngine • Cartagena de Indias

Presencial
COP 374.895.000 - 562.342.000
Growth budget
Competitive pay
Remote work
+3
Senior Data Engineer ID82269
Senior Data Engineer ID82269

AgileEngine • Centrosur

Presencial
COP 90.000.000 - 120.000.000
Growth without limits
Competitive compensation
100% remote
+4
Data Engineer (Lead) ID41785
Data Engineer (Lead) ID41785

AgileEngine • Bucaramanga

Híbrido
COP 374.895.000 - 562.342.000
Professional growth opportunities
Competitive USD-based compensation
Exciting project selection
+1
SR Data Engineer - Snowflake (LATAM)
SR Data Engineer - Snowflake (LATAM)

Pearster • Bogotá

Híbrido
COP 250.196.000 - 357.424.000
Work from anywhere
Paid time off
International certifications
+2
Senior Full-stack Engineer (React/Node) (Backend-Focused) - Real Estate - LATAM
Senior Full-stack Engineer (React/Node) (Backend-Focused) - Real Estate - LATAM

Truelogic • Colombia

A distancia
COP 376.825.000 - 565.238.000
100% Remote Work
Highly Competitive USD Pay
Paid Time Off
+2
Senior Data Engineer
Senior Data Engineer

AgileEngine • Colombia

Híbrido
COP 376.530.000 - 564.794.000
Professional growth
Competitive USD-based compensation
Flextime
Senior Data Engineer ID71670
Senior Data Engineer ID71670

AgileEngine • Bogotá

Híbrido
COP 281.171.000 - 406.136.000
Professional growth
Competitive compensation
Exciting projects
+1