Opening the paper…
What best describes "big data"? Which of these are implementations of the divide-and-conquer paradigm? A website sees 1 Billion hits every month. The website owner wants to count average hits per customer in the latest month, where a customer is denoted by the IP address of the device from which the customer is accessing the website. The owner has at his disposal a Hadoop cluster of 5 workers and 2 masters each with 1GB of RAM. Which of the following methods is the most likely to finish fastest?