GCP Data Engineer with Gen AI | Code1 | Remote

Code1

United States

Remote

USD 165,000 - 248,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Code1 is seeking an experienced GCP Data Engineer with Gen AI expertise for a remote, full-time contract. You will design, build, and optimize large-scale data pipelines and data warehouses on Google Cloud Platform, integrating Gen AI workflows.

Role requires 7–13 years of software engineering experience, strong Python/PySpark skills, and proficiency with BigQuery, Dataflow, Pub/Sub, and Cloud Composer. Immediate availability and CI/CD practices are essential.

Qualifications

  • Minimum of 7 to 13 years of professional software engineering experience with strong focus on GCP data engineering.
  • Hands-on production experience designing and optimizing data architectures on Google Cloud Platform.
  • Proficiency in Python and PySpark for large-scale distributed data processing.
  • Advanced SQL and data modeling for data warehousing.
  • Experience with ETL/ELT pipelines and workflow orchestration tools (Cloud Composer / Airflow).
  • Familiarity with CI/CD pipelines and Git version control.

Responsibilities

  • Architect, build, and optimize end-to-end data pipelines on Google Cloud Platform.
  • Integrate Gen AI models and workflows with cloud data pipelines and data warehouses.
  • Develop and tune ETL/ELT pipelines using BigQuery, Dataflow, and Dataproc.
  • Manage real-time data ingestion via Pub/Sub and storage via GCS.
  • Write clean Python and PySpark code and maintain CI/CD practices.
  • Design scalable data models and enterprise analytics layers on GCP.
  • Monitor performance, optimize queries, and govern costs across BigQuery and GCP storage.
  • Lead data engineering efforts with cross-functional analytics and ML teams.

Skills

Python
PySpark
SQL

Tools

BigQuery
Dataflow
Dataproc
Pub/Sub
Airflow/Cloud Composer
CI/CD with Git

Job description

Position Summary:

Code1 is actively seeking an experienced, highly skilled, and innovative GCP Data Engineer with Generative AI expertise and 7 to 13 years of professional engineering experience to join our team on a full-time contract / freelancing remote basis. In this critical data architecture and AI integration role, you will be responsible for designing, developing, optimizing, and scaling advanced data engineering solutions on Google Cloud Platform integrated with cutting-edge Gen AI workflows. You will leverage your deep expertise in BigQuery, Dataflow, Pub/Sub, Dataproc, Python, and PySpark to build robust data pipelines, large-scale data warehouses, and intelligent cloud architectures. This opportunity requires an elite data professional with immediate availability for remote contract work and strong email application protocols.


Detailed Job Description:

As a GCP Data Engineer with Gen AI at Code1, you will take full technical ownership of designing, building, and optimizing large-scale data processing pipelines, data warehousing models, and cloud-native integration flows on Google Cloud Platform. Your day-to-day responsibilities include architecting ETL/ELT pipelines using Cloud Composer and Apache Airflow, managing real-time data streams via Pub/Sub, executing advanced performance tuning in BigQuery, and integrating Generative AI models into enterprise data workflows. You will write clean PySpark and Python code, maintain rigorous CI/CD Git version control, ensure data modeling best practices, and collaborate with cross-functional analytics and machine learning teams to deliver scalable data solutions.


Key Responsibilities:


  • Design, build, optimize, and scale end-to-end data engineering solutions and data pipelines on Google Cloud Platform (GCP).

  • Integrate Generative AI (Gen AI) models and machine learning workflows with cloud data pipelines and data warehouse architectures.

  • Architect, execute, and monitor high-performance ETL/ELT data pipelines using BigQuery, Dataflow, and Dataproc.

  • Manage real-time data ingestion and streaming messaging architectures utilizing Google Cloud Pub/Sub and Cloud Storage (GCS).

  • Develop clean, efficient, and maintainable data processing scripts using Python and PySpark.

  • Design scalable data models, dimensional data warehouses, and GCP data integration layers for enterprise analytics.

  • Conduct rigorous performance tuning, query optimization, and cost governance across BigQuery and GCP data storage environments.

  • Implement robust workflow orchestration using Cloud Composer or Apache Airflow, supported by CI/CD and Git version control.


Required Qualifications & Skills:


  • Minimum of 7 to 13 years of professional software engineering experience, with strong focused specialization in GCP data engineering.

  • Proven hands‑on production experience designing and optimizing data architectures on Google Cloud Platform (BigQuery, Dataflow, Pub/Sub, Dataproc, GCS).

  • Strong practical expertise in Python and PySpark programming for large-scale distributed data processing.

  • Advanced SQL proficiency and deep understanding of data modeling, data warehousing, and relational/non-relational database design.

  • Proven experience building, orchestrating, and monitoring complex ETL/ELT pipelines and workflow orchestration tools (Cloud Composer / Airflow).

  • Practical experience with large-scale data processing, performance tuning, and query optimization.

  • Solid understanding of CI/CD pipelines, Git version control, and collaborative software engineering practices.

  • Mandatory eligibility to work on a remote contract / freelancing basis.


Nice-to-Have Skills:


  • Hands‑on experience integrating Generative AI (Gen AI) models, LLMs, and vector embeddings into GCP data pipelines.

  • Professional cloud certifications such as Google Cloud Professional Data Engineer or Google Cloud Professional Cloud Architect.

  • Familiarity with containerization tools (Docker, Kubernetes) and infrastructure-as-code (Terraform).

  • Prior consulting or freelancing experience delivering enterprise data migration and cloud modernization projects.

  • Strong technical documentation and stakeholder communication capabilities.


Salary/Rate:

Market Competitive / Contract Standard (Freelancing)


Deadline:

Open until filled


Notice Period:

Immediate


Contract Duration:

Contract / Freelancing (Remote Work Mode)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote GCP Data Engineer - Gen AI & Data Pipelines
Remote GCP Data Engineer - Gen AI & Data Pipelines

Code1 • United States

Remote
USD 165,000 - 248,000
GCP Data Engineer
GCP Data Engineer

Capgemini • New York (NY)

On-site
USD 120,000 - 180,000
GCP Data Engineer - feelance
GCP Data Engineer - feelance

Lingaro • United States

Remote
USD 120,000 - 180,000
GCP Data Engineer
GCP Data Engineer

Capgemini • Charlotte (NC)

On-site
USD 110,000 - 150,000
Senior Data AI Engineer
Senior Data AI Engineer

KADIT Innovations • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
GCP Data Engineer
GCP Data Engineer

Programmers.io • California (MO)

On-site
USD 100,000 - 130,000
Hybrid Cloud & Generative AI Developer (GCP)
Hybrid Cloud & Generative AI Developer (GCP)

Experienced Recruiting Partners • City of Albany (NY)

Hybrid
USD 140,000 - 190,000
GCP Data Engineer
GCP Data Engineer

Rivago Infotech Inc • New York (NY)

On-site
USD 120,000 - 160,000
Agentic Data Cloud Customer Engineer II, Google Cloud
Agentic Data Cloud Customer Engineer II, Google Cloud

Google • Sunnyvale (CA)

Hybrid
USD 127,000 - 184,000
Health insurance
Dental insurance
Vision insurance
+2
GCP Architect
GCP Architect

Coltech • United States

On-site
USD 100,000 - 160,000