Data Scientist II - Big Data Engineer

Cogent IBS, Inc

United States

Remote

USD 120,000 - 150,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cogent IBS, Inc. seeks a Data Scientist (Big Data Engineer) II to design and optimize scalable data pipelines with Apache Spark on Databricks, and to integrate with Azure services.

You will own ETL/ELT workflows, data models, governance, security, and CI/CD deployments, collaborating with data scientists and analysts in Agile, multicultural teams. Proficiency in Python and SQL, plus Azure Data Lake/Delta Lake, is required.

Qualifications

  • 4+ years ETL/ELT workflows for structured and unstructured data.
  • 4+ years automating deployments using CI/CD tools.
  • 4+ years collaborating with data scientists, analysts, stakeholders, and cross-functional teams.
  • 4+ years designing and maintaining data models, schemas, and database structures.
  • 4+ years working with data storage solutions, including Azure Data Lake Storage and data warehouses.
  • 4+ years implementing data validation and data quality checks.
  • 4+ years contributing to data governance, metadata management, data lineage, and data cataloging.
  • 4+ years implementing data security measures, including encryption, access controls, and auditing.
  • 4+ years of proficiency in Python and R programming languages.
  • 4+ years of strong SQL querying and data manipulation experience.
  • 4+ years of experience with the Microsoft Azure cloud platform.
  • 4+ years of experience with DevOps, CI/CD pipelines, and version control systems.
  • 4+ years working in Agile and multicultural environments.
  • 4+ years of strong troubleshooting and debugging capabilities.
  • 3+ years designing and developing scalable data pipelines using Apache Spark on Databricks.
  • 3+ years optimizing Spark jobs for performance and cost efficiency.
  • 3+ years integrating Databricks with Azure Data Factory.
  • 3+ years ensuring data quality, governance, and security using Unity Catalog or Delta Lake.
  • 3+ years of strong understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL.
  • 3+ years of hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks.
  • Implement ETL/ELT workflows for structured and unstructured data.
  • Develop and optimize Spark jobs for performance and cost efficiency.
  • Build and maintain data models, schemas, and database structures supporting analytical and operational use cases.
  • Integrate Databricks solutions with Azure Data Factory and other Azure cloud services.
  • Work with Azure Data Lake Storage and data warehouse solutions.
  • Implement data validation and quality checks to ensure data accuracy, consistency, and reliability.
  • Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging.
  • Implement data security measures, including encryption, access controls, and auditing.
  • Support compliance with applicable regulations, security requirements, and industry best practices.
  • Automate deployments using CI/CD pipelines, DevOps practices, and version control systems.
  • Work with Databricks notebooks, clusters, jobs, and Delta Lake.
  • Utilize Unity Catalog and/or Delta Lake to support data quality, governance, and security.
  • Troubleshoot and debug data pipelines, Spark applications, and related technical issues.
  • Collaborate with data scientists, data analysts, stakeholders, and cross-functional teams.
  • Work effectively within Agile and multicultural environments.

Skills

Python
SQL
Spark
Databricks
Azure
CI/CD
Data Modeling
Delta Lake
Data Governance
Unix/Linux

Tools

Azure Data Factory
Azure Data Lake Storage
Delta Lake
Unity Catalog

Job description

ONLY LOCAL TO TEXAS

Data Scientist (Big Data Engineer) II -- Databricks / Azure
Position: Data Scientist (Big Data Engineer) II
Openings: 2
Location: 100% Remote
Duration: 12 Months ( Up to 3 years extension )

Rate : $65/Hr on C2C

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks.
  • Implement ETL/ELT workflows for structured and unstructured data.
  • Develop and optimize Spark jobs for performance and cost efficiency.
  • Build and maintain data models, schemas, and database structures supporting analytical and operational use cases.
  • Integrate Databricks solutions with Azure Data Factory and other Azure cloud services.
  • Work with Azure Data Lake Storage and data warehouse solutions.
  • Implement data validation and quality checks to ensure data accuracy, consistency, and reliability.
  • Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging.
  • Implement data security measures, including encryption, access controls, and auditing.
  • Support compliance with applicable regulations, security requirements, and industry best practices.
  • Automate deployments using CI/CD pipelines, DevOps practices, and version control systems.
  • Work with Databricks notebooks, clusters, jobs, and Delta Lake.
  • Utilize Unity Catalog and/or Delta Lake to support data quality, governance, and security.
  • Troubleshoot and debug data pipelines, Spark applications, and related technical issues.
  • Collaborate with data scientists, data analysts, stakeholders, and cross-functional teams.
  • Work effectively within Agile and multicultural environments.

Required Qualifications

  • 4+ years of experience implementing ETL/ELT workflows for structured and unstructured data.
  • 4+ years of experience automating deployments using CI/CD tools.
  • 4+ years collaborating with data scientists, analysts, stakeholders, and cross-functional teams.
  • 4+ years designing and maintaining data models, schemas, and database structures.
  • 4+ years working with data storage solutions, including Azure Data Lake Storage and data warehouses.
  • 4+ years implementing data validation and data quality checks.
  • 4+ years contributing to data governance, metadata management, data lineage, and data cataloging.
  • 4+ years implementing data security measures, including encryption, access controls, and auditing.
  • 4+ years of proficiency in Python and R programming languages.
  • 4+ years of strong SQL querying and data manipulation experience.
  • 4+ years of experience with the Microsoft Azure cloud platform.
  • 4+ years of experience with DevOps, CI/CD pipelines, and version control systems.
  • 4+ years working in Agile and multicultural environments.
  • 4+ years of strong troubleshooting and debugging capabilities.
  • 3+ years designing and developing scalable data pipelines using Apache Spark on Databricks.
  • 3+ years optimizing Spark jobs for performance and cost efficiency.
  • 3+ years integrating Databricks with Azure Data Factory.
  • 3+ years ensuring data quality, governance, and security using Unity Catalog or Delta Lake.
  • 3+ years of strong understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL.
  • 3+ years of hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake.

Preferred Qualifications

  • Knowledge of machine learning libraries such as:

  • MLflow

  • Scikit-learn

  • TensorFlow

  • Databricks Certified Associate Developer for Apache Spark certification.

  • Microsoft Certified: Azure Data Engineer Associate certification.

Role Overview:
The position involves designing, developing, and optimizing scalable data pipelines and big data solutions using Azure, Databricks, and Apache Spark. The role also involves data quality, governance, security, CI/CD, and collaboration with data engineering, analytics, and business teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer II - Databricks, Azure & Spark
Data Engineer II - Databricks, Azure & Spark

Cogent IBS, Inc • United States

Remote
USD 120,000 - 150,000
Senior Data Engineer on-site)
Senior Data Engineer on-site)

Ziosk • Dallas (TX)

On-site
USD 140,000 - 190,000
Data Engineer(Ony locals)
Data Engineer(Ony locals)

Sophus IT Solutions • Atlanta (GA)

On-site
USD 120,000 - 150,000
Databricks Data Engineer
Databricks Data Engineer

VOLTO Consulting • Irving (TX)

On-site
USD 120,000 - 150,000
Azure Databricks Engineer
Azure Databricks Engineer

ZEUS SOLUTIONS INC • Houston (TX)

On-site
USD 140,000 - 180,000
Fulltime Only- Sr. Azure Data Engineer/Architect - Local to CA only
Fulltime Only- Sr. Azure Data Engineer/Architect - Local to CA only

E-Solutions • Los Angeles (CA)

On-site
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Stelvio Inc. • Plano (TX)

On-site
USD 110,000 - 170,000
On-site role
Fridays from home after 90 days
Sr Data Engineer
Sr Data Engineer

Golden Technology • Cincinnati (OH)

On-site
USD 100,000 - 130,000
Senior Data Engineer (Azure, Databricks, PySpark)
Senior Data Engineer (Azure, Databricks, PySpark)

The Planet Group • Addison (TX)

On-site
USD 194,848,000 - 209,175,000
Data Engineer - Databricks
Data Engineer - Databricks

Open Systems Technologies • New York (NY)

On-site
USD 137,760 - 179,088