QuestionQ36

Designing data processing systems

Your data science team must run interactive SQL queries on large datasets in Apache Parquet format that are stored in a Cloud Storage bucket. The team knows Apache Hive and wants to use its existing HiveQL queries.

You need to provide an environment where the team can execute interactive HiveQL queries directly against the Cloud Storage data while keeping operational overhead to a minimum. What should you do?

Explanation

Dataproc is a managed Hadoop ecosystem service that includes Apache Hive and integrates with Cloud Storage through its installed Cloud Storage connector. It can run Hive queries, including HiveQL that defines external tables located in gs:// paths, while Google manages cluster provisioning and operation. BigQuery external tables can read Parquet files in Cloud Storage, but they are queried with GoogleSQL rather than HiveQL.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!