QuestionQ54

Data Pipeline Orchestration

You work for a gaming company that gathers real-time player-activity data. The data is streamed into Pub/Sub and must be processed and loaded into BigQuery for analysis. Processing includes filtering, enriching, and aggregating the data before loading it into partitioned BigQuery tables. You need to design a pipeline that provides low latency and high throughput while using a Google-recommended approach. What should you do?

Explanation

Dataflow provides a managed, scalable streaming pipeline for reading Pub/Sub events, applying per-record transformations and windowed aggregations, and writing results to BigQuery. Google Cloud recommends a Dataflow pipeline when complex transformations or aggregations are required before Pub/Sub data is stored in BigQuery; its Pub/Sub-to-BigQuery guidance also covers performance for high-throughput streaming workloads.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!