MonoEdge develops an advanced intelligence layer for the Indian mid-market manufacturing sector. We provide data-driven insights and strategic recommendations to plant leadership and operations supervisors across diverse linguistic contexts, including English, Hindi, and Marathi.
As a technically-driven, bootstrapped organization, we are defining a new category in industrial optimization.
What you'll work on
- Ingestion — pulling data in from many sources: PLC and sensor streams (OPC UA, Modbus), lab systems, ERP exports, and flat files, each on its own schedule and format.
- Cleaning and validation — handling missing values, unit mismatches, duplicate rows, and silent gaps, and catching bad data before it reaches a model or a report.
- Joining heterogeneous sources — aligning data recorded at very different frequencies and timestamps into tables that can actually be analysed together.
- Storage and access — shaping how time-series and relational data are stored so that queries are fast and the schema makes sense to the people using it.
- Reliable pipelines — building jobs that run on a schedule, recover from failure, and tell someone when a source stops sending data.
What you'll do
- Write Python and SQL to move, clean, and join data from real industrial sources — not tidy sample datasets.
- Design table schemas with the senior team so that heterogeneous data lands somewhere sensible and stays queryable.
- Build validation and monitoring so a broken or missing feed is caught early, not discovered in a wrong report a week later.
- Work closely with the data scientist to give models the clean, joined tables they need, in the shape they need them.
- Occasionally go to source — a plant or an automation partner — to understand what a field actually means before you build a pipeline around it.
Who should apply
- A recent graduate in any engineering discipline or a related field. What matters is comfort with data and code, not the exact branch.
- Solid Python and SQL: you can write a script that reads, transforms, and writes data, and query a database confidently. You should be able to reason about a join, not just run one.
- A feel for data structure: you understand tables, keys, and types, and why a well-shaped schema saves everyone downstream a lot of pain.
- Patience with messy data: missing values, inconsistent units, undocumented sources. You fix it methodically rather than complain about it — and you know that most of the work is here.
- A reliability mindset: you care whether a job ran, whether it ran correctly, and how you would know if it did not.
- Something to show: a project where you moved or wrangled real data — a scraper, an ETL script, a dataset you cleaned and analysed. Anything you can walk us through.
Nice to have
- Any exposure to time-series or industrial data — sensors, IoT, logs, or PLC/SCADA tags.
- Familiarity with a workflow or scheduling tool (Airflow, cron, or similar) and the idea of an ETL/ELT pipeline.
- Comfort with pandas, and with databases beyond a single table — Postgres, a time-series DB, or a warehouse.
- Basic cloud or Linux comfort — running a job somewhere other than your laptop.
- Familiarity with Git, the command line, and working in a small team.
What you will learn
You will learn how data engineering works on genuinely hard, real-world data — the kind with no schema and no documentation — from a senior team that builds analytics and machine learning on top of what you produce. You will see exactly how your pipelines feed real decisions.
You will work directly with the Founder and senior engineers, with real ownership of the data layer from early on. Because we are small, the tables you build are used almost immediately.
How we work
We operate with agility and a rigorous focus on real, EBITDA-level impact at the customer, not on isolated metrics. You will be given ownership early, and the support to grow into it.
MonoEdge speaks the way a reliable colleague would — calm, direct, and never overselling itself. We would rather tell you plainly what we do not yet know than pretend we know it, and we expect our data to hold to the same standard.