Senior Engineer - Data Scraping & Python Engineering

Forage AI

India

On-site

INR 4,200,000 - 6,500,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Forage AI is seeking a hands-on Engineering Lead with deep Python expertise to own the data acquisition platform, from architecting resilient crawlers to leading a team of engineers. Your core strength is Python craftsmanship and scalable data extraction.

You will drive end-to-end engineering delivery, ensure production reliability, and shape best practices across the crawling, parsing, and ETL pipelines. Cloud/DevOps experience is valued but not primary.

Qualifications

  • 5–7 years of software engineering with leadership experience.
  • Expert-level Python: async patterns, memory management, packaging.
  • Production-scale web scraping experience with Scrapy/Selenium/Playwright.
  • Strong SQL and NoSQL experience with unstructured/semi-structured data.
  • Experience leading engineers and delivering reliable, scalable systems.

Responsibilities

  • Own end-to-end delivery across requirements, design, implementation, testing and production operations.
  • Lead and grow a team of engineers through mentorship, code reviews, and pairing on hard problems.
  • Translate business and client needs into clear technical plans; manage risks and timelines.
  • Establish and enforce engineering best practices: branching, test discipline, incident/RCA processes.
  • Architect and build scalable crawling and parsing systems (Scrapy/Selenium/Playwright).
  • Write production-grade Python and provide reference implementations and reviews.
  • Build robust anti-bot evasion, proxy rotation, rate limiting, and content fingerprinting.
  • Design parsers and extractors for structured/unstructured content; operate ETL/ELT pipelines.
  • Ensure data quality, observability, and pipeline reliability across the stack.
  • Drive pipeline reliability: monitoring, alerts, graceful failure handling, recovery.

Skills

Python mastery
Async/concurrency
Web scraping
Leadership
Cloud/AWS
DevOps
SQL/NoSQL

Tools

Scrapy
Selenium
Playwright
Docker
Kubernetes
Airflow

Job description

About Forage AI

Forage AI delivers large-scale data collection and processing platforms, including web crawlers, document parsers, data pipelines, AI/Agentic Solutions, and AI-assisted workflows.


Role Overview

We are looking for a hands‑on Engineering Lead who combines deep Python expertise with proven experience building high-scale web scraping and data collection systems. You will own the full lifecycle of our data acquisition platform — from architecting resilient crawlers and parsers to leading a team of engineers and ensuring production reliability.


Cloud/DevOps experience is valued but not the primary focus; your core strength is in Python craftsmanship and data extraction at scale.


Key Responsibilities

Engineering Leadership


  1. Own end-to-end delivery across requirements, design, implementation, testing, and production operations.

  2. Lead and grow a team of engineers through mentorship, code reviews, pairing on hard problems, and raising overall engineering quality.

  3. Translate business and client needs into clear technical plans; manage risks, trade-offs, and timelines.

  4. Establish and enforce engineering best practices: branching strategy, code standards, testing discipline, and incident\/RCA processes.

  5. Architect and build scalable, fault‑tolerant crawling and parsing systems (Scrapy, Selenium, Playwright, or equivalent frameworks).

  6. Write production‑grade Python and set the bar via reference implementations, design docs, and thorough code reviews.

  7. Build robust mechanisms for anti‑bot evasion, proxy rotation, rate limiting, and content fingerprinting.

  8. Design parsers and extractors for structured and unstructured content across diverse and changing web targets.

  9. Design and operate ETL\/ELT pipelines for large‑scale data enrichment, transformation, and delivery.

  10. Ensure data quality, consistency, and observability across the collection and processing stack.

  11. Drive pipeline reliability: monitoring, alerting, graceful failure handling, and recovery strategies.


Infrastructure & Operations (Supporting Role)


  1. Work with cloud infrastructure (primarily AWS) for deployment, scheduling, and storage — own enough to be self‑sufficient.

  2. Collaborate on CI/CD, containerisation (Docker\/Kubernetes), and environment management.

  3. Contribute to observability: structured logging, metrics, and alerting for scraping and pipeline workloads.


Required Qualifications


  • 5-7 years of software engineering experience, with meaningful team or project leadership in a production environment.

  • Expert-level Python: strong command of async\/concurrency patterns, memory management, packaging, and writing idiomatic, maintainable code.

  • Hands‑on, production experience with web scraping frameworks — Scrapy and\/or Selenium (or Playwright) — at meaningful scale.

  • Experience designing and operating resilient crawlers: retry logic, deduplication, scheduling, and distributed crawl management.

  • Solid SQL and NoSQL experience: schema design, query optimisation, and working with unstructured\/semi-structured data.

  • Proven ability to lead engineers: code reviews, technical mentorship, driving delivery discipline and engineering quality.

  • Strong DS\/Algo fundamentals; ability to reason clearly about system design, performance, and correctness.

  • Excellent communication: clear design documents, strong stakeholder communication, and actionable technical feedback.


Preferred / Good to Have

Data Pipelines (High Priority)


  • Experience building or operating data pipeline tooling — Airflow, Prefect, Luigi, or similar orchestrators.

  • Familiarity with streaming systems (Kafka, Kinesis) for real‑time data ingestion workflows.

  • Experience with Spark or distributed compute frameworks for large‑scale data transformation.

  • Working knowledge of AWS services relevant to data workloads: S3, Lambda, ECS, SQS, RDS\/DynamoDB.

  • Familiarity with CI/CD (GitHub Actions, Jenkins) and containerised deployments (Docker, Kubernetes).

  • Basic infrastructure‑as‑code experience (Terraform or CloudFormation) is a plus, not a requirement.


Additional Nice-to-Haves


  • Familiarity with vector databases or semantic search systems. Security awareness: OWASP basics, secrets management, and least‑privilege patterns.

  • Experience with NLP\/ML tooling for content extraction, classification, or enrichment.

  • Hiring, interviewing, and onboarding experience; talent development mindset.

  • Frontend\/JS exposure (nice‑to‑have only; useful for understanding scraping targets).


What Sets You Apart

You are a craftsperson who takes pride in clean, performant Python. You understand the web deeply, not just how to scrape it, but why it behaves the way it does. You've operated systems that collect data at scale and you know how to make them reliable. You lead by example, elevate the people around you, and communicate clearly with both engineers and non‑technical stakeholders.



  • High‑speed internet for calls and collaboration.

  • A capable, reliable computer (modern CPU, 16GB+ RAM).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)

FullStackTechies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Nextyn • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Challenging projects
High growth potential
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)

AIMLEAP Inc • India

Remote
INR 1,000,000 - 1,500,000
Senior Web Scraping Engineer
Senior Web Scraping Engineer

Alternative Path • India

On-site
INR 1,200,000 - 1,800,000
Technical Recruiter
Technical Recruiter

Forage AI • India

On-site
INR 600,000 - 1,200,000
Data Engineer - Web Scraping
Data Engineer - Web Scraping

Alternative Path • India

On-site
INR 900,000 - 1,300,000
Docker-enabled workflows
Lead Python Developer Data Analytics Visualization
Lead Python Developer Data Analytics Visualization

Udyora Careers • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Technical Lead - Bangalore
Technical Lead - Bangalore

Juniper Square • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Technical Lead - Bangalore
Technical Lead - Bangalore

Apply • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Python Developer
Senior Python Developer

Neurealm • Gurugram District

On-site
INR 2,500,000 - 4,000,000