You created an analytics environment on Google Cloud so your data scientist team can explore data without affecting the on-premises Apache Hadoop solution. Data in the on-premises Hadoop Distributed File System (HDFS) cluster is stored in Optimized Row Columnar (ORC) formatted files with multiple columns of Hive partitioning. The data scientist team must be able to explore the data in a way similar to how they used the on-premises HDFS cluster with SQL on the Hive query engine. You need to select the most cost-effective storage and processing solution. What should you do?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion