Senior Data Architect

Arkhya Tech. Inc.

Dallas (TX)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech firm in Dallas, Texas, seeks an experienced data engineer to lead the migration of a SQL data warehouse to a Data Lake on Google Cloud. The successful candidate will design data models, build ingestion pipelines, and utilize advanced analytics for data-driven decision-making. A Bachelor's or Master's in Computer Science and 10+ years of relevant experience is needed. Google Cloud certification is required, and a competitive salary is offered, reflecting experience and expertise.

Qualifications

  • 10–14 years in data engineering/architecture required.
  • 5+ years of experience designing solutions on GCP.
  • Google Cloud Professional Cloud Architect certification required.

Responsibilities

  • Migrate data warehouse to a Data Lake on GCP.
  • Design data models optimized for analytics.
  • Build batch and streaming ingestion pipelines using GCP services.

Skills

Data engineering
Cloud Computing (GCP)
BigQuery SQL
Data modeling
Python programming
Data ingestion pipelines

Education

Bachelor’s/Master’s in Computer Science, Information Systems

Tools

BigQuery
Dataflow
Apache Beam
Cloud Composer
Apache Spark

Job description

Overview

Contract

Project/Program

Identity & Access Management (IAM) Data Modernization – migration of an on‑premises SQL data warehouse to a target‑state Data Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, and GenAI use cases (natural language querying, accelerated summarization, cross‑domain trend analysis).

About Program/Project

The IAM Data Modernization project involves migrating an on-premises SQL data warehouse to a target state Data Lake in GCP cloud environment. Key highlights include:

  • Integration Scope: 30+ source system data ingestions and multiple downstream integrations
  • Capabilities: Metrics, reporting, and Gen AI use cases with natural language querying, advanced pattern/trend analysis, faster summarizations, and cross-domain metric monitoring
  • Scalability and access to advanced cloud functionality
  • Highly available and performant semantic layer with historical data support
  • Unified data strategy for executive reporting, analytics, and Gen AI across cyber domains

This modernization establishes a single source of truth for enterprise-wide data-driven decision-making.

Data Lake Architecture & Storage
  • Proven experience designing and implementing data lake architectures (e.g., Bronze/Silver/Gold or layered models).
  • Strong knowledge of Cloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls

Experience with Hadoop/HDFS architecture, distributed file systems, and data locality principles

  • Hands-on experience with columnar data formats (Parquet, Avro, ORC) and compression techniques
  • Expertise in partitioning strategies, backfills, and large-scale data organization
  • Ability to design data models optimized for analytics and BI consumption
Qualifications
  • Experience: [10–14]+ years in data engineering/architecture, 5+ years designing on GCP at scale; prior on‑prem → cloud migration a must.
  • Education: Bachelor’s/Master’s in Computer Science, Information Systems, or equivalent experience.
  • Certifications: Google Cloud Professional Cloud Architect (required or within 3 months). Plus: Professional Data Engineer, Security Engineer.
Data Ingestion & Orchestration
  • Experience building batch and streaming ingestion pipelines using GCP-native services
  • Knowledge of Pub/Sub-based streaming architectures, event schema design, and versioning
  • Strong understanding of incremental ingestion and CDC patterns, including idempotency and deduplication
  • Hands-on experience with workflow orchestration tools (Cloud Composer / Airflow)
  • Ability to design robust error handling, replay, and backfill mechanisms
Data Processing & Transformation
  • Experience developing scalable batch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc)
  • Strong proficiency in BigQuery SQL, including query optimization, partitioning, clustering, and cost control.
  • Hands-on experience with Hadoop MapReduce and ecosystem tools (Hive, Pig, Sqoop)
  • Advanced Python programming skills for data engineering, including testing and maintainable code design
  • Experience managing schema evolution while minimizing downstream impact
Analytics & Data Serving
  • Expertise in BigQuery performance optimization and data serving patterns
  • Experience building semantic layers and governed metrics for consistent analytics
  • Familiarity with BI integration, access controls, and dashboard standards
  • Understanding of data exposure patterns via views, APIs, or curated datasets
Data Governance, Quality & Metadata
  • Experience implementing data catalogs, metadata management, and ownership models
  • Understanding of data lineage for auditability and troubleshooting
  • Strong focus on data quality frameworks, including validation, freshness checks, and alerting
  • Experience defining and enforcing data contracts, schemas, and SLAs
  • Familiarity with audit logging and compliance readiness
  • Strong hands-on experience with Google Cloud Platform (GCP), including project setup, environment separation, billing, quotas, and cost controls
  • Expertise in IAM and security best practices, including least-privilege access, service accounts, and role-based access
  • Solid understanding of VPC networking, private access patterns, and secure service connectivity
  • Experience with encryption and key management (KMS, CMEK) and security auditing
DevOps, Platform & Reliability
  • Proven ability to build CI/CD pipelines for data and infrastructure workloads
  • Experience managing secrets securely using GCP Secret Manager
  • Ownership of observability, SLOs, dashboards, alerts, and runbooks
  • Proficiency in logging, monitoring, and alerting for data pipelines and platform reliability
Good to have
  • Security, Privacy & Compliance
Security, Privacy & Compliance
  • Hands-on experience implementing fine-grained access controls for BigQuery and GCS
  • Experience with VPC Service Controls and data exfiltration prevention
  • Knowledge of PII handling, data masking, tokenization, and audit requirements
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Google Cloud Data Architect IAM Data Modernization
Google Cloud Data Architect IAM Data Modernization

Vytwo • Prosper (TX)

Hybrid
USD 150,000 - 180,000
Flexible work from home options
GCP Data Architect
GCP Data Architect

KANINI • United States

On-site
USD 120,000 - 160,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Dearborn (MO)

On-site
USD 100,000 - 130,000
Data Architect
Data Architect

SysTechCorp Inc • New York (NY)

On-site
USD 180,000 - 260,000
GCP Architect
GCP Architect

TechDigital Group • Secaucus (NJ)

On-site
USD 90,000 - 140,000
GCP Data Engineer
GCP Data Engineer

TechDigital Group • Seattle (WA)

On-site
USD 130,000 - 160,000
Senior Data Analytics Architect
Senior Data Analytics Architect

Codinix Consulting Services • Southlake (TX)

On-site
USD 120,000 - 150,000
GCP Data Engineer
GCP Data Engineer

Fixity Technologies • Charlotte (NC)

On-site
USD 120,000 - 180,000
GCP Data Engineer
GCP Data Engineer

Rivago Infotech Inc • New York (NY)

On-site
USD 120,000 - 160,000
Staff Data Platform and Products Architect
Staff Data Platform and Products Architect

Jobtailor • Dearborn (MO)

On-site
USD 140,000 - 190,000