QuestionQ28

Preparing and using data for analysis

A shipping company has live package-tracking data that is sent to an Apache Kafka stream in real time and then loaded into BigQuery. Analysts at your company want to query the tracking data in BigQuery to analyze geospatial trends throughout a package’s lifecycle. The table was initially created with ingest-date partitioning. Over time, query processing time has increased. You need to implement a change that improves query performance in BigQuery. What should you do?

Explanation

BigQuery clustering sorts storage blocks by the clustering column. Clustering the existing ingestion-time-partitioned table by package-tracking ID colocates a package’s tracking events and lets BigQuery prune irrelevant blocks when queries filter by that ID, reducing data scanned for lifecycle analysis. Introduction to clustered tables

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!