QuestionQ37

Designing data processing systems

You created an analytics environment on Google Cloud so your data scientist team can explore data without affecting the on-premises Apache Hadoop solution. Data in the on-premises Hadoop Distributed File System (HDFS) cluster is stored in Optimized Row Columnar (ORC) formatted files with multiple columns of Hive partitioning. The data scientist team must be able to explore the data in a way similar to how they used the on-premises HDFS cluster with SQL on the Hive query engine. You need to select the most cost-effective storage and processing solution. What should you do?

Explanation

Cloud Storage can act as durable storage that is compatible with Hadoop-style access through Dataproc’s automatically installed HDFS-compatible Cloud Storage connector. A Dataproc cluster supplies the Hadoop/Hive environment, allowing the existing ORC files and Hive partition layout to be queried with Hive SQL while keeping persistent data separate from transient cluster storage.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!