SENIOR DATA ENGINEER

Svitla Systems, Inc.

Colombia

Presencial

COP 200.541.462 - 300.812.193

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Christmas Bonus (50% of monthly pay)
Bonuses for writing/public talks
Personalized learning program
Free tech webinars and meetups
Remote-friendly culture
Team celebrations

Descripción de la vacante

Svitla Systems, Inc. is seeking an experienced Senior Data Engineer to own a Python/Scrapy-based legal data ingestion pipeline in Colombia and Mexico. The role combines large-scale data engineering, web scraping against complex sources, and cost-efficient cloud design, plus maintaining a React frontend.

Key tasks include building 50+ jurisdiction scrapers, preserving original document structure, and enabling robust search and change detection while optimizing resources.

Formación

  • 5+ years in data engineering or backend development with production pipelines.
  • Strong experience in Python with production web scraping and structured-data processing.
  • Practical understanding of anti-bot measures within legal bounds.
  • Understanding of relational and non-relational databases; schema design for hierarchical data.
  • SQL and data modeling for change detection and display.
  • Experience with cloud infrastructure (AWS, GCP, or Azure), IaC (Terraform/CloudFormation), and containers (Docker).
  • Knowledge of Azure services (VMs, Blob storage, App Service, containers).
  • Knowledge of search architectures (full-text, semantic search — Elasticsearch/OpenSearch).
  • Familiarity with messaging/notification systems (email APIs, Slack/Teams).
  • Infrastructure-as-code for automated deployments.
  • Be comfortable owning a system end-to-end with minimal hand-holding.
  • Ability to handle complex legacy systems with limited docs.
  • Meticulous attention to detail for legal text processing.
  • Cost- and efficiency-oriented mindset.
  • Comfort with ambiguity across diverse data sources.

Responsabilidades

  • Take over and understand an existing Python/Scrapy ingestion pipeline and codebase.
  • Build and maintain web scrapers across 50+ jurisdictions with varied sources and barriers.
  • Consolidate tooling to robust scraping capabilities; avoid sprawl.
  • Normalize content into structured plain-text for hashing while preserving original HTML for display.
  • Preserve source hierarchy (nested clauses, numbering, Unicode markers, tables).
  • Maintain multi-destination architecture and flattened schema for search.
  • Operate change detection and notifications, aiming for paragraph-level diffs.
  • Design cost-efficient infrastructure with autoscaling and right-sizing.
  • Maintain React/Vite frontend for regulations display.

Conocimientos

Python
Web scraping
SQL
Cloud platforms
IaC
Docker
Data pipelines
Legacy systems
Vulnerability to anti-bot measures
Schema design
Elasticsearch/OpenSearch
Message systems
Cost optimization
Team ownership

Herramientas

Terraform/CloudFormation
Docker
AWS
GCP
Azure
Elasticsearch/OpenSearch

Descripción del empleo

Svitla Systems Inc. is looking for a Senior Data Engineer for a full-time position (40 hours per week) in Colombia, Mexico.

You'll take full ownership of a Python/Scrapy-based legal/regulatory data ingestion pipeline. This role combines large-scale data engineering, advanced web scraping against hostile sources, complex legal document processing, and cost-efficient cloud infrastructure design. It also includes maintenance of a React frontend application.

Requirements
  • 5+ years of experience in data engineering or backend development, with a focus on production pipelines.
  • Strong experience in Python, including production web scraping (Scrapy or equivalent) and structured-data processing.
  • Practical understanding of defeating or working around anti-bot measures (rate limiting, Cloudflare, paywalled/gated legal databases) within legal and ethical bounds.
  • Understanding of relational and non-relational databases; schema design for hierarchical data.
  • Solid understanding of SQL and data modeling, with the judgment to design schemas that serve both change detection and human-readable display.
  • Experience with cloud infrastructure (AWS, GCP, or Azure), IaC (Terraform/CloudFormation), and containers (Docker).
  • Knowledge of Microsoft Azure, including VMs, Blob storage, App Service plans, and container services, with the ability to provision and tear down resources programmatically.
  • Knowledge of search architectures (full-text, semantic search — Elasticsearch, OpenSearch, or vector DBs).
  • Familiarity with messaging/notification systems (email APIs, Slack/Teams webhooks).
  • Understanding of Infrastructure-as-code for repeatable, automated deployments.
  • Be comfortable owning a system end-to-end with minimal hand-holding, including reading and improving inherited code and documentation.
  • The ability to take ownership of complex legacy systems without extensive documentation.
  • Meticulous attention to detail: this work handles legal text where structural precision is critical.
  • Cost- and efficiency-oriented mindset.
  • Be comfortable working with ambiguity across heterogeneous, changing data sources.
Nice to have
  • Experience with React and Vite for maintaining and extending the display layer.
  • Prior experience in legal tech, regtech, or compliance software.
  • Familiarity with anti-detection scraping techniques (proxy rotation, fingerprinting, simulated human behavior).
  • Familiarity with deterministic hashing and content versioning systems.
  • Experience designing systems for horizontal scale (clustering, multi-VM, multi-threaded, or containerized workloads).
  • Experience AWS alongside Azure (the concepts transfer; multi-cloud).
  • Familiarity with semantic / cross-jurisdictional search techniques.
  • Exposure to LLM or vision-AI integration — e.g., using AI to parse document structure, generate summaries and action statements, or pre-classify regulatory changes as impacting vs. non-impacting for human review.
Responsibilities
  • Take over and fully understand an existing Python/Scrapy ingestion pipeline, its codebase, and its documented database schemas, then make it your own.
  • Build and maintain web scrapers across 50+ jurisdictions, each with its own structure, format, and obstacles — XML feeds, HTML pages, DOCX-only sources (e.g., West Virginia), and content behind Westlaw, LexisNexis, Cloudflare, and similar barriers.
  • Consolidate tooling where possible: prefer one or two robust, broadly capable scraping tools over a sprawl of point solutions, falling back to specialized tooling only for genuinely esoteric sources.
  • Normalize ingested content into a structured, plain-text format for deterministic hashing and change detection, while preserving the original, formatted HTML so documents can be displayed exactly as the issuing agency intended.
  • Faithfully preserve each source's original hierarchy (nested clauses, sub-paragraphs, lettered and numbered subsections, Roman numerals, Unicode markers, appendices, and tables), so stored content remains a true representation of the statute.
  • Maintain and extend a multi-destination architecture, including the nested-hierarchy schema used by the SaaS compliance platform (white-labeled in some deployments) and a flattened schema supporting full-text, semantic, and cross-jurisdictional search.
  • Operate and evolve change detection and notification (email, Slack, Teams), with a path toward more granular, paragraph- and clause-level change tracking and side-by-side visual diffs.
  • Design the infrastructure for cost efficiency: model how many VMs/containers are needed and for how long, automate spin-up and shutdown so nothing runs idle, and right-size compute based on real scraping timings rather than assumptions.
  • Maintain the React and Vite single-page application that displays ingested regulations, including a searchable table-of-contents navigation pattern.
We offer
  • US and EU projects based on advanced technologies.
  • Competitive compensation based on skills and experience.
  • Remote-friendly culture and no micromanagement.
  • Christmas Bonus in the amount of 50% of the monthly payment.
  • Bonuses for article writing, public talks, other activities.
  • Personalized learning program tailored to your interests and skill development.
  • Free tech webinars and meetups organized by Svitla.
  • Fun corporate online/offline celebrations and activities.
  • Awesome team, friendly and supportive community!
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Data Engineer — Remote, End-to-End Data Pipelines
Senior Data Engineer — Remote, End-to-End Data Pipelines

Svitla Systems, Inc. • Colombia

Híbrido
COP 200.541.000 - 300.813.000
Christmas Bonus (50% of monthly pay)
Bonuses for writing/public talks
Personalized learning program
+3
SENIOR SOFTWARE ENGINEER (REACT+.NET)
SENIOR SOFTWARE ENGINEER (REACT+.NET)

Svitla Systems, Inc. • Colombia

Presencial
COP 120.000.000 - 230.000.000
Christmas Bonus 50% of monthly payment
Referral bonuses
Personalized learning program
+1
Data Scrapping Engineer
Data Scrapping Engineer

Trabajosihay • Colombia

Híbrido
Remote Business Immigration Paralegal for Boutique Law Firm
Remote Business Immigration Paralegal for Boutique Law Firm

Pearl Talent • Bogotá

Presencial
COP 131.889.000 - 204.114.000
Fully remote
Competitive salary
Mentorship
+4
Lead FullStack Software Engineer
Lead FullStack Software Engineer

N-iX • Colombia

A distancia
COP 294.691.000 - 442.038.000
Flexible working format
Competitive salary
Personalized career growth
+2
Senior Data Engineer ID81743
Senior Data Engineer ID81743

AgileEngine • Cartagena de Indias

Presencial
COP 374.895.000 - 562.342.000
Growth budget
Competitive pay
Remote work
+3
Senior Data Engineer ID81743
Senior Data Engineer ID81743

AgileEngine • Bogotá

A distancia
COP 120.000.000 - 180.000.000
Growth without limits
Competitive compensation
Remote work 100%
+3
Associate Software Engineer
Associate Software Engineer

Foundation Inc. • Bogotá

Presencial
COP 91.151.000 - 145.842.000
Sr Software Data Engineer - Python
Sr Software Data Engineer - Python

MPS Group LLC • Perímetro Urbano Medellín

Híbrido
COP 240.288.000 - 320.385.000
Compensatory days off
Certification support
Birthday bonuses
+2
Senior Full Stack Developer ID71007
Senior Full Stack Developer ID71007

AgileEngine • Sur

Presencial
COP 376.825.000 - 565.238.000
Growth opportunities
Competitive compensation
Remote work 100%
+3