QuestionQ24

Maintaining and automating data workloads

You are monitoring your organization’s data lake, which is hosted on BigQuery. The ingestion pipelines read data from Pub/Sub and write it into BigQuery tables. After deploying a new version of the ingestion pipelines, the daily stored data grew by 50%. Pub/Sub data volumes remained unchanged, and only some tables had their daily partition data size double. You need to investigate and resolve the cause of the data increase. What should you do?

Explanation

Duplicate data in only a subset of tables after a pipeline deployment indicates that more than one pipeline version may be ingesting the same Pub/Sub data into those tables. Duplicate rows confirm the symptom; BigQuery audit logs identify the jobs associated with table writes, and Dataflow monitoring correlates those jobs with their start times and pipeline versions. Stopping all but the latest version removes the concurrent writer and prevents further duplicate ingestion. BigQuery audit logs record table and job activity, while the Dataflow monitoring interface provides job status and timing information.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!