Quiz Space

Introduction to Big Data End Term: 28 April 2024 (January 2024 term)

Question 1

+2 marksOne correct option

What best describes "big data"?

  1. A

    Big data is used to refer to the set of technologies built around Hadoop

  2. B

    Big data refers to the cloud-native design principle

  3. C

    Big data is about All data, Any Time, Any Method

  4. D

    Big data is about volume, velocity, variety

Also asked in Quiz 2 2 Apr 2023, Quiz 2 2 Apr 2023, Quiz 2 6 Aug 2023, Quiz 2 3 Dec 2023

Question 2

+2 marksOne correct option

Which of these are implementations of the divide-and-conquer paradigm?

  1. A

    Spark and Map

  2. B

    Serverless and Message Broker

  3. C

    Hadoop and Spark

  4. D

    MapReduce and Google Cloud Functions

Question 3

+2 marksOne correct option

A website sees 1 Billion hits every month. The website owner wants to count average hits per customer in the latest month, where a customer is denoted by the IP address of the device from which the customer is accessing the website. The owner has at his disposal a Hadoop cluster of 5 workers and 2 masters each with 1GB of RAM. Which of the following methods is the most likely to finish fastest?

  1. A

    Write a MapReduce program where the Map does nothing useful, Combine computes the aggregated hits per customer, Shuffle combined data based on IP address across workers, and in the Reduce, build hash table on each machine with hash key = IP address and hash value = counter, followed by another Reduce that finally computes the avg on top of all hash values.

  2. B

    Write a Spark program that forms a Dataframe as grouping by IP address with count as aggregate, followed by a take into a list in the Spark driver which further computes the average of all the individual counts in the list

  3. C

    Write a Spark program that forms a Dataframe as grouping by IP address with count as aggregate, followed by another stage that computes the avg on top of the Dataframe of the first stage

  4. D

    Write a MapReduce program where 2 pairs of Map Reduce are chained together: 1^(st) pair is where the Map does nothing useful, Shuffle data based on IP address across workers, and in the Reduce, build hash table on each machine with hash key = IP address and hash value = counter, while the 2^(nd) pair is another Map that does nothing useful followed by a Reduce that finally computes avg on top of all hash values.

  5. E

    All will finish in approximately the same time.

Also asked in End Term 1 Sept 2024

27 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 28 Apr 2024 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 28 Apr 2024, in the January 2024 term: 30 questions for 60 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 28 Apr 2024 at a glance
TermJanuary 2024 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions30
Marks60
Duration180 min
MCQ22
MSQ8
Official paperIIT M DEGREE FN EXAM QDB1 28 Apr 2024
Negative markingNo negative marking.
Updated

Same End Term, other subjects

More Intro to Big Data