QuestionQ35

Ingesting and processing the data

You plan to load some existing on-premises data into BigQuery on Google Cloud. Depending on the use case, you want to stream or batch-load the data. You also want to mask sensitive data before it is loaded into BigQuery. You must accomplish this programmatically while minimizing costs. What should you do?

Explanation

Dataflow, using the Apache Beam SDK for Python, supports both batch and streaming pipelines in code and can write to BigQuery. The pipeline can call Cloud DLP (Sensitive Data Protection) to de-identify data before it reaches the BigQuery sink, satisfying the requirement to mask sensitive data before loading while avoiding the persistent Cloud Data Fusion instance costs.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!