Data Engineer

Codvo.ai

Pune District

On-site

INR 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Codvo.ai is seeking a Data Engineer in Pune, India, to own data pipelines from BMS ingestion to analytics. The role involves designing resilient data ingestion and storage solutions, ensuring data quality, and integrating multiple systems.

The ideal candidate will have 5+ years in data engineering, expertise in ETL/ELT pipelines, and strong Python skills. A solid understanding of time-series databases and familiarity with industrial protocols will be advantageous. Join us to help drive innovation through AI.

Qualifications

  • 5+ years in data engineering including ETL/ELT pipelines and time-series databases.
  • Strong Python skills with experience in PostgreSQL/TimescaleDB.
  • Familiarity with industrial data protocols like OPC-UA and BACnet is a strong advantage.

Responsibilities

  • Build and maintain BMS protocol bridge for data ingestion.
  • Implement MQTT ingestion and monitor ingestion health.
  • Design and maintain TimescaleDB schema for data storage.
  • Manage the MinIO/S3 data lake and implement data retention policies.
  • Build data quality monitoring pipeline and track data lineage.

Skills

ETL/ELT pipelines
Streaming data
Python
PostgreSQL/TimescaleDB
MQTT
Docker
Kubernetes
Data quality frameworks

Job description

About us: Codvo.ai is a next-gen AI and engineering company helping global enterprises transform through Generative AI, Cloud-native platforms, and Product Engineering. With proprietary platforms like NeIO and Pulse, we’re enabling faster, smarter, and scalable digital transformation for industries including Energy, Retail, Travel, BFSI, and Healthcare.

As we gear up to launch new AI-powered products and expand global presence, we are seeking a marketing leader to define how Codvo.ai influences the market, shapes perception, and creates a movement.

Role Summary

Owns the data pipeline from BMS ingestion through to the analytics layer. Responsible for data reliability, quality, and the real-time streaming infrastructure that feeds the ML models.

Responsibilities
Data Ingestion & Streaming
  • Build and maintain the BMS protocol bridge — OPC-UA, BACnet, and MQTT connectors
  • Implement the MQTT ingestion pipeline — topic subscription, message parsing, schema validation, TimescaleDB insertion
  • Monitor ingestion health — message rates, latency, dropped messages, reconnection logic
  • Implement the tag auto-mapping engine — pattern-based matching, confidence scoring, manual override workflow
  • Build the historian adapter — bulk data extraction from customer historian systems (AVEVA Historian, OSIsoft PI, InfluxDB) for baseline profiling
Data Storage & Management
  • Design and maintain the TimescaleDB schema — hypertables, continuous aggregates, retention policies, compression
  • Implement data partitioning strategy — per-tenant, per-site isolation
  • Build the training data snapshot pipeline — versioned dataset extraction with SHA-256 checksums and manifest tracking
  • Manage the MinIO/S3 data lake — dataset storage, MLflow artifact storage, backup strategy
  • Implement data retention and archival policies per customer requirements
Data Quality & Monitoring
  • Build the data quality monitoring pipeline — null rate tracking, stale timestamp detection, out-of-range value flagging, schema violation alerting
  • Implement the tag freshness monitor — detect when sensors stop reporting, alert on stale data
  • Build the data lineage tracking system — from raw BMS reading through feature computation through model prediction
  • Monitor database performance — query latency, storage growth, index health, connection pool utilization
Integration
  • Build and maintain the CMMS integration — ServiceNow / Maximo API connectors, work order creation, status sync
  • Implement the feedback ingestion worker — poll CMMS for work order outcomes, match to predictions, update ground truth labels
  • Build the Prometheus/Grafana metrics export pipeline — platform health metrics, data quality dashboards
Expected Background
  • 5+ years in data engineering — ETL/ELT pipelines, streaming data, time-series databases
  • Strong Python skills, experience with PostgreSQL/TimescaleDB, MQTT, and message broker systems
  • Experience with Docker, Kubernetes, and CI/CD pipelines
  • Familiarity with OPC-UA, BACnet, or industrial data protocols is a strong advantage
  • Experience with data quality frameworks and monitoring

Note- Please apply via our official careers portal only, as applications sent directly to executives may not be considered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Codvo Private Limited • Pune District

On-site
INR 800,000 - 1,200,000
Solution Architect (Remote)
Solution Architect (Remote)

Codvo.ai • Pune District

Remote
INR 2,000,000 - 2,500,000
Data Engineer Lead (OT Data)( Oil & Gas) (India)
Data Engineer Lead (OT Data)( Oil & Gas) (India)

Codvo Private Limited • India

On-site
INR 1,400,000 - 2,400,000
Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Data Engineer
Data Engineer

Navikenz India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineering Consultant (Category - Architect)
Data Engineering Consultant (Category - Architect)

Codvo.ai • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Data Engineer
Senior Data Engineer

Confidential • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Hybrid options
Learning budget
Competitive pay
+1
Data Engineer Lead (OT Data)( Oil & Gas) (India)
Data Engineer Lead (OT Data)( Oil & Gas) (India)

Codvo.ai • Pune District

On-site
INR 1,200,000 - 1,800,000
Sr Engineer - DataOps
Sr Engineer - DataOps

Tridiagonal Ai • Pune District

On-site
INR 1,200,000 - 2,000,000
Data Engineer
Data Engineer

Meril • Vapi

On-site
INR 800,000 - 1,200,000