QuestionQ38

Maintaining and automating data workloads

Your company has expanded rapidly and is now ingesting data at a substantially higher rate than before. You manage the daily batch MapReduce analytics jobs in Apache Hadoop. However, the recent data increase has caused the batch jobs to fall behind. You have been asked to recommend ways the development team can improve analytics responsiveness without increasing costs. What should you recommend?

Explanation

Apache Spark is a fast processing engine that is compatible with Hadoop data and can run on Hadoop clusters. Rewriting batch analytics in Spark can improve responsiveness without purchasing additional cluster capacity; expanding the Hadoop cluster would increase costs. Spark documentation reports faster performance than Hadoop MapReduce for comparable large-scale processing workloads.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!