Data Architect (Mid Level)

Cornerstone Global Partners (CGP Group)

United States

On-site

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cornerstone Global Partners (CGP Group) is seeking a Data Architect to design the unified data model powering Sentinel, an AI-driven risk and asset integrity platform. You will establish data lineage, governance, and quality standards while collaborating with data engineers to implement scalable pipelines.

The role emphasizes hands-on work with Databricks, Delta Lake, and Unity Catalog, shaping how data supports probabilistic risk models and cross-product analytics.

Qualifications

  • 5 to 8 years in data architecture with end-to-end data modeling.
  • Hands-on Databricks with Delta Lake and Unity Catalog.
  • Strong medallion/lakehouse architecture and dimensional modeling.
  • Advanced SQL; Python or PySpark proficiency.
  • Experience with at least one major cloud platform; Azure preferred.
  • CDC, SCD, schema evolution, and data-quality frameworks.
  • Data governance, cataloging, lineage, RBAC, and encryption.

Responsibilities

  • Design the schema and unified data model for the platform across products.
  • Apply medallion architecture patterns and define semantic layer for downstream use.
  • Model data for probabilistic risk models with defined sources and quality checks.
  • Design cross-product entity resolution and ingestion patterns (batch, streaming, CDC).
  • Produce architecture documentation, diagrams, and data dictionaries.
  • Define GIS layer, geospatial relationships, and linear referencing components.
  • Collaborate with data engineers, analysts, and product teams to implement.
  • Maintain documentation and defend design decisions with evidence.

Skills

Advanced SQL
Python or PySpark
Data modeling
Dimensional modeling
Semantic layer design
Data governance
RBAC / access control
Data quality frameworks

Education

Bachelor's degree in CS or related

Tools

Databricks
Delta Lake
Unity Catalog
SQL Warehouses

Job description

About Our Client

Our client is a SaaS company providing pipeline integrity management software to oil and gas pipeline operators, spanning asset integrity, damage prevention, land management, and related risk workflows.

About Our Client

Our client is a SaaS company providing pipeline integrity management software to oil and gas pipeline operators, spanning asset integrity, damage prevention, land management, and related risk workflows.

Role Overview

We're looking for a Data Architect to design the schema and data model powering Sentinel, a new AI-driven risk and asset integrity platform that unifies data across multiple products into one governed data platform. You will design a unified structure spanning asset integrity, damage prevention, and land and stakeholder management, capable of supporting probabilistic risk models that require dozens of precisely defined input attributes per pipeline segment. This is a hands-on design role: you will establish standards for data lineage, governance, and quality while working closely with the data engineering team to implement them. This is a skill-first role: hands-on experience with Databricks, Delta Lake, Unity Catalog, and medallion/lakehouse architecture matters far more than prior exposure to the oil and gas or pipeline integrity industry, which is a nice-to-have, not a requirement.

Key Responsibilities
  • Design the schema and unified data model for the platform, spanning asset, inspection, incident, geospatial, and consequence data across multiple products.
  • Apply medallion architecture patterns (Bronze, Silver, Gold) and define the semantic layer consumed by downstream models, reporting, and applications.
  • Model the data required by probabilistic risk models, ensuring every required input attribute has a defined source, transformation path, and data-quality expectation.
  • Design cross-product entity resolution so records from separate products reconcile to shared assets and entities.
  • Define ingestion patterns for batch, streaming, and change-data-capture (CDC) sources, including schema evolution and slowly changing dimensions.
  • Produce architecture documentation, diagrams, data dictionaries, and decision records that engineers can implement without ambiguity.
  • Design the shared GIS layer as a first-class component of the data model, including pipeline centerlines, consequence-area polygons, right-of-way corridors, and parcel data.
  • Model linear referencing and dynamic segmentation so risk results align with centerline geometry and can be aggregated across differing segmentations.
  • Define data structures for inspection data, repair history, and external enrichment sources such as weather history, soil characteristics, and satellite-derived data.
  • Define cataloging, classification, and metadata standards in Unity Catalog and establish expectations for lineage coverage across the data estate.
  • Establish data-quality rules, validation gates, and profiling standards, including processes for quarantining and remediating failures.
  • Define access-control patterns, including role-based and attribute-based access, tenant isolation, and appropriate handling of sensitive data.
  • Work directly with data engineers to translate architectural designs into production pipelines and review implementations for conformance.
  • Partner with application engineers and data scientists to determine how the data model is consumed through serving layers, feature stores, and APIs.
  • Interface directly with other product teams to understand where their data lives and develop the data contracts needed to move data into the unified platform.
  • Participate in design reviews and architecture alignment sessions, and defend design decisions using data and evidence.
  • Maintain architecture documentation as the model evolves, treating stale or inaccurate documentation as a defect.
Requirements
  • 5 to 8 years of experience in data architecture, data modeling, or senior data engineering, including end-to-end ownership of a non-trivial data model; up to 10 years is welcome.
  • Hands-on experience with Databricks, including Delta Lake, Unity Catalog, SQL Warehouses, and pipeline orchestration.
  • Strong command of medallion and lakehouse architecture patterns, dimensional modeling, and semantic layer design.
  • Advanced SQL skills and proficiency in Python or PySpark.
  • Experience with at least one major cloud platform; Azure experience preferred.
  • Practical experience with change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data-quality frameworks.
  • Working knowledge of data governance, including cataloging, lineage, classification, role-based access control (RBAC), and encryption.
Nice-to-Have
  • Strong documentation and diagramming skills, with the ability to explain data models clearly to both technical and non-technical stakeholders.
  • Prior exposure to pipeline integrity or oil and gas regulatory frameworks.
  • Experience designing geospatial data models and working with GIS tooling, spatial joins, and linear referencing.
  • Experience with multi-tenant platforms, data residency requirements, or regulated environments.
  • Familiarity with metadata and lineage tools such as Unity Catalog, Microsoft Purview, or equivalent platforms.
  • Experience modeling data specifically for machine learning or probabilistic risk models.
  • Experience designing semantic layers for business intelligence (BI) consumption.
  • Databricks or cloud data certifications.
  • Experience using AI-assisted coding tools such as Cursor or GitHub Copilot, and/or agentic coding tools such as Claude Code, as part of a professional development workflow.
  • Familiarity with pipeline or utility asset data, including inline inspection results, alignment sheets, facilities, and centerline geometry.
  • Understanding of asset integrity concepts such as corrosion growth, defect tracking, and consequence-of-failure modeling.
  • Awareness of regulatory reporting requirements applicable to pipeline integrity data.
  • Experience migrating data from legacy desktop or spreadsheet-based systems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Data Engineer
Principal Data Engineer

Nodi • Minnesota

On-site
USD 150,000 - 190,000
Senior Data Engineer / Data Architect | Hybrid | Camp Murray, WA | Contingent
Senior Data Engineer / Data Architect | Hybrid | Camp Murray, WA | Contingent

Ares Enterprise • Northern (KY)

Hybrid
USD 120,000 - 180,000
Data Architect
Data Architect

EXL • Chicago (IL)

On-site
USD 130,000 - 190,000
Director Sr. Information Architect
Director Sr. Information Architect

Ledgent Technology • California (MO)

Remote
USD 130,000 - 160,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Data Architect
Data Architect

SiloSmashers • Washington

On-site
USD 140,000 - 170,000
Data Architect
Data Architect

Capital Technology Alliance • Tallahassee (FL)

On-site
USD 130,000 - 160,000
Health insurance
401(k) retirement plan
Flexible working hours
AI/ML Data & ETL Data Architect
AI/ML Data & ETL Data Architect

DATAECONOMY Inc • Charlotte (NC)

On-site
USD 130,000 - 180,000
Data Architect
Data Architect

Aureon • West Des Moines (IA)

On-site
USD 120,000 - 180,000
Data Architect
Data Architect

SiloSmashers, Inc. • Washington

On-site
USD 150,000 - 190,000