Google Cloud Data Architect IAM Data Modernization

Vytwo

Prosper (TX)

Hybrid

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work from home options

Job summary

Vytwo is seeking a Google Cloud Data Architect in Texas to lead IAM Data Modernization projects, converting on-premises SQL data warehouses to Google Cloud. This role demands 10–14+ years in DevOps and strong skills in PySpark and GCP.

Candidates should have expertise in designing scalable data solutions with CI/CD pipelines, data modeling, and Google Cloud tools. The company offers flexible work from home options and expects certifications in GCP.

Qualifications

  • 10–14+ years in DevOps and Data Architecture focusing on data modernization.
  • 5+ years designing on PySpark, GCP, and OCP at scale.
  • Must possess Google Cloud Professional Cloud Architect certification or be obtained within 3 months.

Responsibilities

  • Implement CI/CD pipelines for data and analytics workloads.
  • Design data platforms on Google Cloud Platform.
  • Build batch and streaming ingestion pipelines using GCP-native services.

Skills

CI/CD pipelines for data and analytics workloads
OpenShift Container Platform (OCP)
PySpark for ETL/ELT
Google Cloud Platform (GCP)
Data modeling for analytics
BigQuery SQL optimization
Cloud Storage design
Hands-on experience with Hadoop/HDFS

Education

Bachelor’s/Master’s in Computer Science, Information Systems

Tools

Git-based source control
Cloud Composer / Airflow

Job description

Role

Google Cloud Data Architect – IAM Data Modernization

Location

Dallas, TX / Charlotte, NC / Iselin, NJ / Chandler, AZ / Ohio, Delaware (Hybrid)

Eligibility

Must be a US Citizen/GC only

About Position

Identity & Access Management (IAM) Data Modernization – migration of an on‑premises SQL data warehouse to a target‑state Data Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, and GenAI use cases (natural language querying, accelerated summarization, cross‑domain trend analysis) leveraging PySpark‑based processing, cloud‑native DevOps CI/CD pipelines, and containerized deployments on OpenShift (OCP) to deliver scalable, secure, and high‑performance data solutions.

What You'll Do
  • Experience implementing CI/CD pipelines for data and analytics workloads.
  • Familiarity with Git‑based source control, build automation, and deployment strategies.
  • Experience with OpenShift Container Platform (OCP) for deploying data workloads and services.
  • Understanding of containerized architecture, scaling, and environment management.
  • Proven ability to build CI/CD pipelines for data and infrastructure workloads.
  • Experience managing secrets securely using GCP Secret Manager.
  • Ownership of observability, SLOs, dashboards, alerts, and runbooks.
  • Proficiency in logging, monitoring, and alerting for data pipelines and platform reliability.
  • Hands‑on experience with PySpark for ETL/ELT, data transformation, and performance optimization.
  • Solid understanding of distributed data processing concepts.
  • Strong experience designing data platforms on Google Cloud Platform (GCP).
  • Experience with Data Lakes, data warehousing, and large‑scale migration programs.
  • Proven experience designing and implementing data lake architectures (e.g., Bronze/Silver/Gold or layered models).
  • Strong knowledge of Cloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls.
  • Experience with Hadoop/HDFS architecture, distributed file systems, and data locality principles.
  • Hands‑on experience with columnar data formats (Parquet, Avro, ORC) and compression techniques.
  • Expertise in partitioning strategies, backfills, and large‑scale data organization.
  • Ability to design data models optimized for analytics and BI consumption.
  • Experience building batch and streaming ingestion pipelines using GCP-native services.
  • Knowledge of Pub/Sub‑based streaming architectures, event schema design, and versioning.
  • Strong understanding of incremental ingestion and CDC patterns, including idempotency and deduplication.
  • Hands‑on experience with workflow orchestration tools (Cloud Composer / Airflow).
  • Ability to design robust error handling, replay, and backfill mechanisms.
  • Experience developing scalable batch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc).
  • Strong proficiency in BigQuery SQL, including query optimization, partitioning, clustering, and cost control.
  • Hands‑on experience with Hadoop MapReduce and ecosystem tools (Hive, Pig, Sqoop).
  • Advanced Python programming skills for data engineering, including testing and maintainable code design.
  • Experience managing schema evolution while minimizing downstream impact.
  • Expertise in BigQuery performance optimization and data serving patterns.
  • Experience building semantic layers and governed metrics for consistent analytics.
  • Familiarity with BI integration, access controls, and dashboard standards.
  • Understanding of data exposure patterns via views, APIs, or curated datasets.
  • Experience implementing data catalogs, metadata management, and ownership models.
  • Understanding of data lineage for auditability and troubleshooting.
  • Strong focus on data quality frameworks, including validation, freshness checks, and alerting.
  • Experience defining and enforcing data contracts, schemas, and SLAs.
Good to have
Security, Privacy & Compliance
  • Hands‑on experience implementing fine‑grained access controls for BigQuery and GCS.
  • Experience with Sprint planning and helping team technically.
  • Strong stakeholder communication and solution‑architecture skills.
Expertise You'll Bring
  • Experience: 10–14+ years in DevOps and Data Architecture, 5+ years designing on Pyspark/GCP/OCP at scale; prior on‑prem → cloud migration a must.
  • Education: Bachelor’s/Master’s in Computer Science, Information Systems, or equivalent experience.
  • Certifications: Google Cloud Professional Cloud Architect/DevOps/OCP (required or within 3 months). Plus: Professional Data Engineer, Security Engineer.
  • Flexible work from home options available.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Architect
Senior Data Architect

Arkhya Tech. Inc. • Dallas (TX)

On-site
USD 120,000 - 150,000
Data Architect
Data Architect

SysTechCorp Inc • New York (NY)

On-site
USD 180,000 - 260,000
GCP Data Engineer
GCP Data Engineer

TechDigital Group • Seattle (WA)

On-site
USD 130,000 - 160,000
GCP Data Engineer
GCP Data Engineer

Fixity Technologies • Charlotte (NC)

On-site
USD 120,000 - 180,000
Staff Data Platform and Products Architect
Staff Data Platform and Products Architect

Jobtailor • Dearborn (MO)

On-site
USD 140,000 - 190,000
GCP Data Architect
GCP Data Architect

KANINI • United States

On-site
USD 120,000 - 160,000
Big Data Architect (GCP)
Big Data Architect (GCP)

SoftServe • Town of Poland (NY)

On-site
USD 120,000 - 190,000
GCP Data Engineer (Fulltime)
GCP Data Engineer (Fulltime)

ASB Resources • Princeton (NJ)

On-site
USD 120,000 - 160,000
GCP Data Engineer
GCP Data Engineer

Rivago Infotech Inc • New York (NY)

On-site
USD 120,000 - 160,000
Sr Data Engineer
Sr Data Engineer

TechDigital Group • St. Louis (MO)

On-site
USD 80,000 - 120,000