Quiz Space

Intro to Big Data End Term: 13 April 2025, Set 1-5 (January 2025 term)

Question 1

+2 marksOne correct option

At the onset of every festive season, there is a surge in quick commerce orders (i.e. orders delivered within 15 minutes) on Swiggy. The supply chain head at Swiggy for a city is interested in a real-time view of the inventory of her dark stores (i.e. stores without nameboards where the supplies are kept and used to fulfil app orders). She wants to see this be presented in a monitor mounted in her office wall that refreshes with the latest info on a city map every 1 minute. Along with the info, there also needs to be the current time so that she gets a visual confirmation that this is the latest data. This dashboard allows her to plan for new orders of specific supplies that she is running out of so that no customer is left unsatisfied. What solution option below best solves for the need?

  1. A

    Route a copy of every quick commerce item ordered to a Kafka topic, useSpark Structured Streaming to continuously read from this topic and update the current counts of supplies, and emit using output mode “Update”.

  2. B

    Route a copy of every quick commerce item ordered to a Kafka topic, useSpark Structured Streaming to periodically read from this topic every 1 minute and update the current counts of supplies, and emit all aggregates using the output mode “Complete”.

  3. C

    Route a copy of every quick commerce item ordered to a Kafka topic, useSpark Structured Streaming to periodically read from this topic every 1 minute and count the supplies ordered in that batch, and emit only all aggregates in that batch using the output mode “Append”.

Also asked in End Term 13 Apr 2025

Question 2

+2 marksOne correct option

You are given a Spark program that runs on a Google Dataproc cluster on a daily schedule starting execution at 7AM running typically for 2 hours, to produce as output the total amount spent on purchases made by every customer the previous day. The input data is coming into GCS every minute from a variety of sources as standalone files. Therefore, the business leader now feels that having to wait till 9AM the next day is no longer acceptable and instead ideally wants purchase information for each customer at least every 5 minutes during the day itself. What’s more, she wants to be able to change this time configuration later without involving you.
Which amongst the below represents the best option to achieve the above?

  1. A

    Change the schedule to run every 5 minutes, no other change required.

  2. B

    Change the code to leverage Spark Streaming with streaming window as “5minutes”, & let her manage the execution of the code on Google Dataproc

  3. C

    Change the code to leverage Spark Streaming with streaming window as “5minutes”, go from using Dataproc to Dataflow, & let her manage the execution of the code on Dataflow

  4. D

    Write a Cloud Function to move all incoming per-minute standalone files fromGCS to Pub/Sub, change the code to leverage Spark Streaming with streaming window as “5 mins”, convert from Dataproc to Dataflow, point source to Pub/Sub, & let her manage the execution of the code on Dataflow

Also asked in End Term 13 Apr 2025

Question 3

+2 marksOne correct option

You are asked to build a recommendation engine for an offline store at the checkout counter, using a combination of cloud and point-of-sale (PoS) terminal resources. Consider the following pipeline choices for effecting the same outcome, where Kubernetes is an open-source system for automating deployment, scaling and management of containerized applications, Google Datastore and HBase are both highly-scalable NoSQL database systems for interactive, real-time applications, and Google Vertex AI is a hosted platform that lets you train and deploy ML models
(i) Recommender model training on Python VM on GCP; Data publisher on VM in PoS → Kafka VM on GCP → Spark Streaming on Hadoop VMs on GCP -> Recommender hosted on Python VM on GCP → Relay recommendation to PoS (ii) Recommender model training on Vertex AI; Data publisher on VM in PoS → Kafka VM on GCP → Dataflow -> Recommender hosted on Vertex AI → Relay recommendation to PoS (iii) Recommender model training on Python VM on GCP; Data publisher on VM in PoS → Pub/Sub → Spark Streaming on Google Dataproc → Recommender hosted on Python VM on GCP -> Relay recommendation to PoS (iv) Recommender model training on Python VM on GCP; Data publisher on VM in PoS → Pub/Sub → Dataflow → Recommender hosted on Python VM on GCP → Relay recommendation to PoS (v) Recommender model training on Python VM on GCP; Data publisher on VM in PoS → Pub/Sub → Spark Streaming on Hadoop VMs on GCP → Recommender hosted on Python VM on GCP → Relay recommendation to PoS
Which option below represents the correct order of pipeline options that has the “most IaaS” entry to the left and the “most PaaS” entry to the right?

  1. A

    (i), (v), (iii), (iv), (ii)

  2. B

    (i), (ii), (iii), (iv), (v)

  3. C

    (ii), (iii), (i), (iv), (v)

  4. D

    (v), (iii), (i), (iv), (ii)

  5. E

    All are equally PaaS / IaaS

Also asked in End Term 13 Apr 2025

16 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 13 Apr 2025 Set 1-5 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 13 Apr 2025, in the January 2025 term, set 1-5: 19 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 13 Apr 2025 Set 1-5 at a glance
TermJanuary 2025 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions19
Marks50
Duration180 min
MCQ14
MSQ5
Official paperIIT M IMPROVEMENT FN EXAM QIM2 13 Apr
Negative markingNo negative marking.
Updated

Other sets that day

Same End Term, other subjects

More Intro to Big Data