Senior Data & Python Software Engineer

Ceartas DMCA

Town of Poland (NY)

On-site

USD 110,000 - 170,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ceartas DMCA is seeking a Data & Python Software engineer to drive data pipelines and crawling technologies for scalable web data extraction and brand protection workflows.

You will own data quality, governance, and the reliability of ingestion, processing, and storage layers while collaborating with the CTO and Head of Engineering to advance our data engineering capabilities.

Qualifications

  • Experience in designing and building scalable web scraping systems and data pipelines.
  • Strong SQL skills and experience with relational databases (PostgreSQL or similar).
  • Proficient Python development with backend API design experience.

Responsibilities

  • Design, build, and maintain high-performance web scraping systems and data pipelines.
  • Develop scraping-focused APIs and backend services powering internal products.
  • Ensure data ingestion, processing, and storage workflows for large-scale web data.

Skills

Web scraping
SQL
Python
PostgreSQL
API design
Docker
Cloud platforms

Education

Masters preferred

Tools

FastAPI
Django
Airflow
Playwright
Selenium
DBT
CI/CD

Job description

At Ceartas, we lead the way in AI-powered brand protection, copyright law, and digital security, safeguarding the integrity of content creators, brands, and enterprises worldwide. As we scale rapidly, we're looking for a Data & Python Software engineer to drive innovation in our data pipelines and crawling technologies. In this pivotal role, you'll collaborate with our CTO and Head of Engineering, steering our Data Engineering Team toward developing groundbreaking solutions for digital security challenges.

Build Scalable Web Data Extraction Pipelines:

Design and develop web scraping systems that support large-scale web data extraction and brand protection workflows. Ensure that data moves reliably from collection through processing to storage while maintaining performance, resilience, and operational stability at scale.

Ensure Data Quality and Governance:

Own data validation, consistency, and governance across ingestion, storage, and serving layers. Establish clear standards for schema design, transformation logic, and monitoring to guarantee trustworthy, production-grade datasets that can be reliably consumed across the organization.

Optimize Performance and Reliability:

Continuously improve scraping system efficiency through performance tuning, cost optimization, and architectural enhancements. Implement logging, metrics, and tracing to monitor production systems, diagnose issues quickly, and maintain high reliability under growing workloads.

Responsibilities:
  • Design, build, and maintain high-performance web scraping systems as well backend services and data pipelines supporting web data extraction and brand protection use cases
  • Implement and maintain scraping focused APIs and other data services that power internal products and external integrations
  • Build reliable ingestion, processing, and storage workflows for large-scale web data
  • Handle cleaning of web data and ensure data quality, validation, and governance across ingestion, storage, and serving layers
  • Optimize scraping systems for performance, scalability, reliability, and cost efficiency
  • Monitor, debug, and improve scraping system reliability using observability tools (logging, metrics, tracing)
  • Collaborate closely with product and engineering teams to deliver features from design through full end-to-end production deployment
  • Take independent ownership of systems in production, including maintenance,
  • iteration and performance management
Core Technical Requirements:
  • Experience with web scraping
  • Strong SQL skills
  • Strong Python experience
  • Experience with PostgreSQL or similar relational databases
  • Experience designing and building scalable APIs and backend services (e.g. FastAPI, Django, or similar frameworks)
  • Experience designing efficient, scalable data models and database schemas
  • Hands-on experience deploying and operating systems in the cloud (AWS, GCP, or Azure)
  • Experience working with Docker and containerized environments
Preferred Technical Requirements:
  • Experience with workflow orchestration tools such as Airflow
  • Experience with browser-based automation tools (Playwright, Selenium, or similar)
  • Experience with DBT or analytics-focused data transformation workflows
  • Experience building or operating high-concurrency systems and task queues
  • Experience designing and deploying cloud-native workflows on AWS
  • Familiarity with CI/CD pipelines and production deployment practices
  • Experience working in a high-growth, early-stage startup environment
  • Experience - University education in a technical field such as Computer Science, Engineering or similar. Masters level preferred. 4+ years ( or 2 year+ in a early stage startup)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data & Python Software Engineer
Senior Data & Python Software Engineer

Ceartas • Town of Poland (NY)

On-site
USD 120,000 - 180,000
25 paid days off
Culture of innovation
Diversity and equal opportunity
+3
Senior Data & Python Engineer — Scalable Web Scraping
Senior Data & Python Engineer — Scalable Web Scraping

Ceartas • Town of Poland (NY)

On-site
USD 120,000 - 180,000
25 paid days off
Culture of innovation
Diversity and equal opportunity
+3
Senior Data Engineer
Senior Data Engineer

Rachel Paul Recruiting • New York (NY)

Hybrid
USD 120,000 - 150,000
Senior Data Engineer - Web Scraping
Senior Data Engineer - Web Scraping

Jobgether • United States

Remote
USD 120,000 - 170,000
Fully remote
Full-time
Autonomy
+1
Director of Software Engineering (Node.js & Web Scraping Expert)
Director of Software Engineering (Node.js & Web Scraping Expert)

Roman Health Pharmacy LLC • Los Angeles (CA)

On-site
USD 150,000 - 200,000
Health insurance
Flexible working hours
Professional development opportunities
ETL/DB/Data Extraction Engineer
ETL/DB/Data Extraction Engineer

WebDataGuru • Indiana (PA)

On-site
USD 110,000 - 170,000
Senior Data Engineer
Senior Data Engineer

deploy • Atlanta (GA)

On-site
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

AgileEngine, LLC • United States

On-site
USD 140,000 - 200,000
100% remote work with flexible hours
Senior Data Engineer ID75059
Senior Data Engineer ID75059

AgileEngine • Miami (FL)

Hybrid
USD 120,000 - 180,000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Data Engineer ID75059
Senior Data Engineer ID75059

AgileEngine • Dallas (TX)

Hybrid
USD 140,000 - 190,000
Professional growth
Competitive compensation
Exciting projects
+1