Virtues is seeking a GCP BigQuery Data Engineer to design cloud-native, scalable data platforms and build real-time ingestion pipelines on Google Cloud. This onsite role in Irving, TX focuses on event-driven streaming architectures that support near real-time analytics and enterprise reporting.
Responsibilities
- Design, develop, and optimize enterprise data warehouse solutions using Google BigQuery.
- Design, develop, and maintain real-time data ingestion pipelines using Apache NiFi to extract, transform, and route streaming data into Apache Kafka topics.
- Build scalable event-driven streaming architectures using Apache NiFi, Apache Kafka, and Google BigQuery to enable high-throughput, fault-tolerant, low-latency processing.
- Configure and optimize Apache NiFi processors for extraction, transformation, routing, filtering, schema validation, error handling, retry mechanisms, and reliable delivery to Kafka.
- Develop streaming ingestion solutions that allow Google BigQuery to consume Kafka event streams and transform near real-time events into analytical tables and reporting datasets.
- Design data models and implement ETL/ELT processes to move data from raw to curated and published layers.
- Design, create, and manage large-scale BigQuery datasets, including temporary/permanent and internal/external tables.
- Optimize BigQuery workloads using query tuning, partitioning, clustering, and cost optimization techniques to improve performance and reduce cloud costs.
- Monitor, troubleshoot, and optimize streaming workloads by tuning NiFi flows, Kafka topics/partitions, and BigQuery streaming ingestion to support high availability, data integrity, minimal latency, and cost-efficient processing.
- Build scalable data lake frameworks and ingestion pipelines using cloud-native GCP technologies.
- Develop reporting and visualization solutions using Looker, Looker Studio (Data Studio), Connected Sheets, and other BigQuery reporting tools.
- Collaborate with data architects, modelers, developers, DevOps engineers, project managers, and business stakeholders to deliver scalable enterprise analytics solutions and continuous improvements to the data platform.
Requirements
- 7+ years of experience in Data Engineering, Data Warehousing, or Big Data platforms.
- Must-have skills: Apache NiFi and Apache Kafka.
- Strong hands-on experience with Google Cloud Platform (GCP), including BigQuery, Cloud Dataflow, Pub/Sub, and Google Cloud Storage (GCS).
- Hands-on experience designing and supporting real-time streaming pipelines using Apache NiFi, Apache Kafka, and Google BigQuery.
- Experience using BigQuery Console/Query Editor for data management, performance tuning, and SQL development.
- Strong experience designing scalable data models, ETL/ELT pipelines, data lake architectures, and enterprise data warehouse solutions.
- Experience with batch and streaming data ingestion using GCP services.
- Thorough understanding of BigQuery cost structure, including storage, ingestion, and query costs, with experience implementing query optimization, partitioning, clustering, and other cost optimization techniques.
- Experience managing large-scale datasets, including temporary/permanent and internal/external BigQuery tables.
- Experience developing reporting and visualization solutions using Looker, Looker Studio (Data Studio), and Connected Sheets.
- Excellent verbal and written communication skills, with the ability to collaborate with technical teams, business stakeholders, senior management, and executive leadership.
- Experience working in Agile environments and collaborating with cross-functional teams to deliver enterprise data engineering solutions.
Preferred Qualifications
- Google Cloud certifications.
- Experience delivering enterprise-scale analytics and cloud data solutions.
Technologies
- Google BigQuery
- Apache NiFi
- Apache Kafka
- Google Cloud Dataflow
- Google Pub/Sub
- SQL
- ETL/ELT
- Google Cloud Storage (GCS)
- Looker
- Looker Studio (Data Studio)
- Connected Sheets
- CI/CD
- DevOps
- Agile/Scrum
- Python
- Spark
- Data Modeling
- Data Warehousing
- Data Lakes
- Batch & Streaming Pipelines
Compensation and Location
Salary: USD 100,000 - 110,000 per year.
Work location: In person (Irving, TX).