Quiz Space

Intro to Big Data End Term: 1 September 2024, Set QDB1 (May 2024 term)

Question 1

+1 markOne correct option

What best describes "hot potato principle" in the world of big data?

  1. A

    It is a rule-of-thumb where one program does all the work required to “cook” the data fully for the sake of completeness

  2. B

    It is a rule-of-thumb applicable mainly to batch processing of data where fresh data is prioritized for processing ahead of older data

  3. C

    It is a rule-of-thumb applicable only in the world of stream processing where a program does the minimal processing required to produce a logical output which is then handed off via a message store to the next program that is also designed similarly, and so on, until the bigger problem is solved.

  4. D

    It is a principle for network packet routing, and not really applicable in big data.

Question 2

+1 markOne correct option

Which of these are implementations of the divide-and-conquer paradigm?

  1. A

    Spark and Map

  2. B

    Serverless and Message Broker

  3. C

    Hadoop and Spark Streaming

  4. D

    MapReduce and Google Cloud Functions

Question 3

+1 markOne correct option

A website sees 1 Billion hits every month. The website owner wants to count average hits per customer in the latest month, where a customer is denoted by the IP address of the device from which the customer is accessing the website. The owner has at his disposal a Hadoop cluster of 5 workers and 2 masters each with 1GB of RAM. Which of the following methods is the most likely to finish fastest?

  1. A

    Write a MapReduce program where the Map does nothing useful, Combine computes the aggregated hits per customer, Shuffle combined data based on IP address across workers, and in the Reduce, build hash table on each machine with hash key = IP address and hash value = counter, followed by another Reduce that finally computes the avg on top of all hash values.

  2. B

    Write a Spark program that forms a Dataframe as grouping by IP address with count as aggregate, followed by a take into a list in the Spark driver which further computes the average of all the individual counts in the list

  3. C

    Write a Spark program that forms a Dataframe as grouping by IP address with count as aggregate, followed by another stage that computes the avg on top of the Dataframe of the first stage

  4. D

    Write a MapReduce program where 2 pairs of Map Reduce are chained together: 1^(st) pair is where the Map does nothing useful, Shuffle data based on IP address across workers, and in the Reduce, build hash table on each machine with hash key = IP address and hash value = counter, while the 2^(nd) pair is another Map that does nothing useful followed by a Reduce that finally computes avg on top of all hash values.

  5. E

    All will finish in approximately the same time.

Also asked in End Term 28 Apr 2024

27 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 1 Sept 2024 Set QDB1 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 1 Sept 2024, in the May 2024 term, set QDB1: 30 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 1 Sept 2024 Set QDB1 at a glance
TermMay 2024 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions30
Marks50
Duration180 min
MCQ23
MSQ7
Official paperIIT M DEGREE FN EXAM QDB1 01 Sep 2024
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Intro to Big Data