Quiz Space

Intro to Big Data End Term: 13 April 2025, Set 1-4 (January 2025 term)

Question 1

+2 marksOne correct option

Which of the following is a false characterization of "hot potato” principle?

  1. A

    It is a rule-of-thumb, not a principle, and therefore is not mandatory

  2. B

    It provides the same benefits when incorporated into the design of both batchas well as streaming data applications.

  3. C

    It emphasizes minimal work in each processing step for maximum scalabilityand optimal recoverability in the event of failures.

  4. D

    It is critically dependent on a message store as the via media for itsintermediate outputs.

Also asked in End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025

Question 2

+2 marksOne correct option

A website that started 3 years ago now sees 1 Billion hits every month. The website owner wants to count average hits per customer across all months, where a customer is denoted by the IP address of the device from which the customer is accessing the website. The owner has at his disposal a minimal Hadoop cluster of 2 workers and 1 master each with 1GB of RAM. He would like to run this job every month going forward. Which of the following methods is the most likely to finish every time it is run in the months and years ahead yielding the right result?

  1. A

    Write a MapReduce program where the Map does nothing useful, Combinecomputes the aggregated hits per customer, Shuffle combines data based on IP address across workers, and the Reduce builds a hash table on each machine with hash key = IP address and hash value = running total and sum, with a final Map that emits the avg per customer.

  2. B

    Write a Spark program that forms a Dataframe as grouping by IP address withcount as aggregate, followed by a take into a list in the Spark driver which further computes the average of all the individual counts in the list

  3. C

    Write a Spark program that forms a Dataframe as grouping by IP address withcount as aggregate, followed by another stage that computes the avg on top of the Dataframe of the first stage

  4. D

    Write a MapReduce program where the Map does nothing useful, Combinecomputes the aggregated hits per customer, Shuffle combines data based on IP address across workers, and the Reduce sorts the data on sort key = IP address and then calculates the final avg per customer.

  5. E

    All will finish every time it is run without issues.

Also asked in End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025

Question 3

+2 marksOne correct option

An enterprise software designer wants to leverage the best of Google cloud to minimize the number of administrative overheads associated with her payment processing pipeline while also getting on-demand scalability without sacrificing flexibility. What option should she choose to best serve these needs?

  1. A

    Build the payment processer using VMs – one for Python for the logic, one forinvoking the external payment engine, and one for storing the results

  2. B

    Build the payment processor using Python running on Google Cloud Functionswhere the results are stored on GCS

  3. C

    Build the payment processor using MapReduce with input data and results arestored on HDFS, and deploy both on Dataproc

  4. D

    Build the payment processor using Dataflow on top of data stored on GCS

Also asked in End Term 13 Apr 2025, End Term 13 Apr 2025, End Term 13 Apr 2025

17 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 13 Apr 2025 Set 1-4 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 13 Apr 2025, in the January 2025 term, set 1-4: 20 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 13 Apr 2025 Set 1-4 at a glance
TermJanuary 2025 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions20
Marks50
Duration180 min
MCQ16
MSQ4
Official paperIIT M IMPROVEMENT FN EXAM QIM2 13 Apr
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Intro to Big Data