Data Engineer

AiLogic Neural Network Pvt Ltd

Hyderabad

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AiLogic Neural Network Pvt Ltd is seeking a Data Engineer to build scalable data pipelines for NLP/LLM applications in Hyderabad. You will ingest, parse, and transform documents from PDFs, HTML, and other text sources, leveraging OCR and Spark-based processing. You’ll collaborate with AI/ML teams to prepare data for model training and inference.

The role emphasizes pipeline scalability, performance optimization, and cost-efficient data workflows, with strong Python and NLP toolkits demanded.

Qualifications

  • 2-3 years of experience in Data Engineering, NLP Engineering, or AI Data Processing.
  • Experience with large-scale datasets and distributed computing frameworks.
  • Strong Python and NLP fundamentals required.

Responsibilities

  • Design, develop, and maintain scalable data pipelines for high-volume data.
  • Build document ingestion for PDFs, scanned docs, HTML pages, and text sources.
  • Implement OCR, PDF parsing, HTML parsing, and text extraction workflows.
  • Develop chunking and preprocessing frameworks for NLP/LLM apps.
  • Collaborate with AI/ML engineers for dataset preparation.
  • Optimize data processing for performance, scalability, and cost.

Skills

Python programming
NLP concepts
Text Processing
Hugging Face Transformers
PDF Parsing
OCR
HTML Parsing
Text Extraction
Document Chunking
Apache Spark & Spark SQL
Vector Databases
Git & CI/CD
ETL pipelines
Debugging & problem solving

Education

Bachelor's or Master's degree in Computer Science / Data Science / Information Technology

Tools

Apache Spark
Spark SQL
Vector Databases
Git
CI/CD

Job description

AiLogic Neural Network Pvt Ltd is an AI-driven product company focused on building advanced language technology solutions, including machine translation, document intelligence, and large-scale NLP systems. We are looking for a highly motivated Data Engineer to join our growing AI team and contribute to the development of scalable data processing pipelines for NLP and LLM applications.

  • Design, develop, and maintain scalable data pipelines for processing large volumes of structured and unstructured data.
  • Build document ingestion and processing workflows for PDFs, scanned documents, HTML pages, and other text sources.
  • Implement OCR, PDF parsing, HTML parsing, and text extraction pipelines.
  • Develop document chunking and preprocessing frameworks for NLP and LLM-based applications.
  • Work with Hugging Face models and NLP libraries for text processing tasks.
  • Create and optimize data transformation workflows using Python, Apache Spark, and Spark SQL.
  • Develop and manage Vector Database pipelines for embedding storage and retrieval.
  • Implement text normalization, sentence segmentation, deduplication, and data quality processes.
  • Design and implement data masking, classification, and categorization solutions.
  • Collaborate with AI/ML engineers to prepare datasets for model training and inference.
  • Optimize large-scale data processing workflows for performance, scalability, and cost efficiency.
  • Maintain CI/CD pipelines and follow software engineering best practices.
  • Monitor, troubleshoot, and improve production data processing systems.

Mandatory Skills

  • Strong experience with Python programming.
  • Hands-on experience in NLP concepts such as:
  • Text Processing
  • Hugging Face Transformers
  • Experience in:
  • PDF Parsing
  • OCR
  • HTML Parsing
  • Text Extraction
  • Document Chunking
  • Experience with Apache Spark and Spark SQL.
  • Working knowledge of Vector Databases.
  • Good understanding of Git and CI/CD practices.
  • Experience building data pipelines and ETL workflows.
  • Strong debugging and problem-solving skills.

Preferred Skills

  • Sentence Segmentation
  • Exact Deduplication and Near Deduplication
  • Data Masking
  • Data Classification & Categorization
  • Embedding Generation and Retrieval Pipelines
  • RAG (Retrieval-Augmented Generation) Pipelines

Performance Optimization Skills

  • CPU Distribution and Parallel Processing
  • Chunking Optimization Strategies
  • GPU and CPU Parallel Distribution
  • CUDA Optimization
  • PyTorch Performance Tuning
  • Spark Performance Optimization

Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Science, Information Technology, or a related field.
  • 2-3 years of experience in Data Engineering, NLP Engineering, or AI Data Processing.
  • Experience working with large-scale datasets and distributed computing frameworks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Engineer
Senior AI Engineer

Genzeon Technology Solutions • Maharashtra

On-site
INR 3,500,000 - 6,000,000
Data Scientist/ NLP
Data Scientist/ NLP

Thompsons HR Consulting Pvt Ltd • Bengaluru

On-site
INR 1,200,000 - 2,200,000
AI/ML & Data Engineer
AI/ML & Data Engineer

Congruent Software Inc. • Chennai District

On-site
INR 800,000 - 1,200,000
AI / Machine Learning Engineer
AI / Machine Learning Engineer

DigitalXNode • Maharashtra

On-site
INR 800,000 - 1,200,000
AI/ML Engineer â Data Science & Data Engineering
AI/ML Engineer â Data Science & Data Engineering

TopGrep Tech Private Limited • Bengaluru

On-site
INR 1,000,000 - 1,800,000
Senior AI Engineer
Senior AI Engineer

BigStep Technologies • Gurugram District

On-site
INR 1,000,000 - 1,500,000
Data Scientist
Data Scientist

Thompsons Hr Consulting • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Lead Engineer - AI/ML
Lead Engineer - AI/ML

Mindfire Solutions • India

On-site
INR 2,000,000 - 3,000,000
AI and Data Engineering Tech Lead
AI and Data Engineering Tech Lead

Carelon Global Solutions • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Data Engineer (Databricks + AI)
Data Engineer (Databricks + AI)

Durapid Technologies Pvt Ltd • Bengaluru Urban

On-site
INR 1,500,000 - 2,300,000