QuestionQ26

Data Pipeline Orchestration

You need to build a data pipeline for a new application. The application will stream data that must be enriched and cleaned. Ultimately, the data will be used to train machine learning models. You need to identify the appropriate data-manipulation methodology and the Google Cloud services to use in this pipeline. What should you select?

Explanation

ETL transforms data before it is loaded into its analytics destination. Dataflow provides scalable batch and streaming processing for transformations such as cleaning and enrichment, and Google Cloud identifies Dataflow-to-BigQuery as a typical ETL workflow. BigQuery provides the destination for prepared data used in analytics and machine-learning workflows.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!