Big Data Engineer

NOV

Ernakulam

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NOV is hiring a Big Data Engineer to design, build, and optimize scalable data ingestion, transformation, and analytics pipelines for IoT, operational, and business data using Spark and Delta Lake. You will develop reusable frameworks for ingesting, validating, enriching, and processing large-scale datasets.

The role requires 5+ years of experience with distributed systems, Timeseries data, cloud technologies, and CI/CD practices.

Qualifications

  • Bachelor’s degree in computer science, information systems, or a related field. However, relevant experience will be considered.
  • A minimum of 5 years of relevant experience required.
  • Advanced understanding with spark and/or Databricks(preferred), including spark cluster tuning.
  • Great understanding of distributed systems and partitioning.
  • Automation and scripting using .net, C#, Python, Javascript, GoLang, AWS Cloud APIs.
  • Database technologies and Timeseries databases like OSIPI, PostgreSQL, Timescale and SQL..
  • Developing applications and/or scripts utilizing Timeseries data.
  • Linux, Windows OS, Containers.
  • Data management and data integration strategies.
  • Amazon Web services and/or cloud technologies.
  • Visualization/reporting tools such as Tableau, Spotfire, Grafana, etc.
  • Understanding of CI/CD practices and technologies. GitHub and GitHub Actions(preferred), Jenkins, TeamCity, Udeploy, etc.
  • Optimally communicating technical information.
  • Working knowledge of ML and AI model development and the data science development life cycle is an added advantage.
  • Experience with large datasets and understanding of stream and batch processing , Datadog, Brokers like Kafka or NATS , Container technologies, and terraform.
  • Ability to relate architectural decisions and recommendations to business needs.
  • Strong analytical and problem-solving skills.
  • Highly independent and ability to communicate across all levels of the organization and work with diverse projects teams.
  • Willingness to demonstrate a growth-mindset when faced with new challenges and opportunities.

Responsibilities

  • Design, develop, and optimize scalable data ingestion, transformation, and analytics pipelines for IoT, operational, and business data using Spark and Delta Lake technologies.
  • Create and tune Spark transformation, aggregation, streaming, and machine learning workloads, including optimization of Spark clusters and processing infrastructure.
  • Develop reusable frameworks for ingesting, validating, enriching, and processing large-scale IoT and event-driven datasets.
  • Build and maintain analytical data models, curated datasets, and KPI calculation frameworks that enable product analytics, operational reporting, business intelligence, and self-service analytics.
  • Collaborate with Product, Engineering, Operations, and business stakeholders to define, implement, and maintain meaningful KPIs, metrics, and data products that support decision-making.
  • Design and implement real-time and batch processing solutions using event-based and streaming technologies.
  • Build and execute large-scale data migration solutions for transferring historical and operational data from source systems such as OSI PI and TimescaleDB into enterprise data platforms.
  • Develop monitoring, observability, reconciliation, and automation tools that ensure the reliability, quality, and performance of data pipelines, migrations, analytics workloads, and MLOps/DataOps processes.
  • Support the delivery of related platform capabilities, including APIs, data services, and integrations, while working within Agile and DevOps methodologies.
  • Document solution architecture, data lineage, business logic, KPI definitions, and operational procedures to support long-term maintainability and governance.

Skills

Spark
Databricks
Distributed systems
Linux
Windows OS
Containers
Python
C#
JavaScript
Go
AWS
Cloud APIs
CI/CD
GitHub Actions
Kafka
NATS
Timeseries databases
PostgreSQL
SQL
MLOps/DataOps
Terraform
Kubernetes
Docker
TimescaleDB
Tableau
Grafana

Education

Bachelor’s degree in CS/IS

Tools

Databricks
Spark
TimescaleDB
PostgreSQL
Tableau
Grafana
Terraform
Kafka
NATS
Docker
Kubernetes
Jenkins
GitHub Actions

Job description

Job Description

Are you passionate about helping people improve the impact of their business through the use of data and analytics? If you enjoy the challenges associated with making complex information systems more user-friendly and effective, we are looking for you to join a team of creative and highly motivated professionals who are driving innovation in the upstream oil field services industry!

Are you passionate about helping people improve the impact of their business through the use of data and analytics? If you enjoy the challenges associated with making complex information systems more user-friendly and effective, we are looking for you to join a team of creative and highly motivated professionals who are driving innovation in the upstream oil field services industry! At National Oilwell Varco, we strive to lead technology innovation that delivers significant value to our customers. We are hiring a Big Data Engineer to help us build and support our data and analytics delivery pipeline.

Responsibilities
  • Design, develop, and optimize scalable data ingestion, transformation, and analytics pipelines for IoT, operational, and business data using Spark and Delta Lake technologies.
  • Create and tune Spark transformation, aggregation, streaming, and machine learning workloads, including optimization of Spark clusters and processing infrastructure.
  • Develop reusable frameworks for ingesting, validating, enriching, and processing large-scale IoT and event-driven datasets.
  • Build and maintain analytical data models, curated datasets, and KPI calculation frameworks that enable product analytics, operational reporting, business intelligence, and self-service analytics.
  • Collaborate with Product, Engineering, Operations, and business stakeholders to define, implement, and maintain meaningful KPIs, metrics, and data products that support decision-making.
  • Design and implement real-time and batch processing solutions using event-based and streaming technologies.
  • Build and execute large-scale data migration solutions for transferring historical and operational data from source systems such as OSI PI and TimescaleDB into enterprise data platforms.
  • Develop monitoring, observability, reconciliation, and automation tools that ensure the reliability, quality, and performance of data pipelines, migrations, analytics workloads, and MLOps/DataOps processes.
  • Support the delivery of related platform capabilities, including APIs, data services, and integrations, while working within Agile and DevOps methodologies.
  • Document solution architecture, data lineage, business logic, KPI definitions, and operational procedures to support long-term maintainability and governance.
Requirements
  • Bachelor’s degree in computer science, information systems, or a related field. However, relevant experience will be considered.
  • A minimum of 5 years of relevant experience required.
  • Advanced understanding with spark and/or Databricks(preferred), including spark cluster tuning.
  • Great understanding of distributed systems and partitioning.
  • Automation and scripting using .net, C#, Python, Javascript, GoLang, AWS Cloud APIs.
  • Database technologies and Timeseries databases like OSIPI, PostgreSQL, Timescale and SQL..
  • Developing applications and/or scripts utilizing Timeseries data.
  • Linux, Windows OS, Containers.
  • Data management and data integration strategies.
  • Amazon Web services and/or cloud technologies.
  • Visualization/reporting tools such as Tableau, Spotfire, Grafana, etc.
  • Understanding of CI/CD practices and technologies. GitHub and GitHub Actions(preferred), Jenkins, TeamCity, Udeploy, etc.
  • Optimally communicating technical information.
  • Working knowledge of ML and AI model development and the data science development life cycle is an added advantage.
  • Experience with large datasets and understanding of stream and batch processing , Datadog, Brokers like Kafka or NATS , Container technologies, and terraform.
  • Ability to relate architectural decisions and recommendations to business needs.
  • Strong analytical and problem-solving skills.
  • Highly independent and ability to communicate across all levels of the organization and work with diverse projects teams.
  • Willingness to demonstrate a growth-mindset when faced with new challenges and opportunities.
About Us

Every day, the oil and gas industry’s best minds put more than 150 years of experience to work to help our customers achieve lasting success.

We Power the Industry that Powers the World

Throughout every region in the world and across every area of drilling and production, our family of companies has provided the technical expertise, advanced equipment, and operational support necessary for success—now and in the future.

Global Family

We are a global family of thousands of individuals, working as one team to create a lasting impact for ourselves, our customers, and the communities where we live and work.

Purposeful Innovation

Through purposeful business innovation, product creation, and service delivery, we are driven to power the industry that powers the world better.

Service Above All

This drives us to anticipate our customers’ needs and work with them to deliver the finest products and services on time and on budget.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer Lead (OT Data)( Oil & Gas) (India)
Data Engineer Lead (OT Data)( Oil & Gas) (India)

Codvo.ai • Pune District

On-site
INR 1,200,000 - 1,800,000
Software Engineer
Software Engineer

NOV • Ernakulam

On-site
INR 1,100,000 - 1,500,000
Data Engineer Lead (OT Data)( Oil & Gas) (India)
Data Engineer Lead (OT Data)( Oil & Gas) (India)

Codvo Private Limited • India

On-site
INR 1,400,000 - 2,400,000
Lead Big Data Engineer
Lead Big Data Engineer

S&P Global, Inc. • Rangareddy

On-site
INR 1,800,000 - 2,600,000
Health & Wellness
Flexible downtime
Continual learning
+3
DT Advisor– AI Ops & Observability
DT Advisor– AI Ops & Observability

bakerhughes • Balewadi

On-site
INR 3,500,000 - 7,000,000
Private medical care options
Wellbeing activities
Global exposure with a leading energy/
Sr Analytics Engineer
Sr Analytics Engineer

GE Vernova • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Relocation Assistance Provided
Applications Software Engineer
Applications Software Engineer

NOV • Ernakulam

On-site
INR 700,000 - 1,200,000
Advanced Data Scientist
Advanced Data Scientist

Exxon Mobil • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Digital Cloud & Data Platform Sr. Engineer
Digital Cloud & Data Platform Sr. Engineer

Black & Veatch • Pune District

On-site
INR 1,500,000 - 2,500,000
Health and welfare benefits
Performance-based bonus program
Paid time off
Senior Data Engineer
Senior Data Engineer

United States Digital Space LLC • Karnataka

On-site
INR 2,500,000 - 4,200,000