Quiz Space

Intro to Big Data End Term: 24 December 2023 (September 2023 term)

Question 1

+2 marksOne correct option

What happens when a Spark Structured Streaming pipeline operating with Kafka as source is subject to a failure of a machine in either of the Kafka cluster or the Spark cluster?

  1. A

    Failure of a machine in the Kafka cluster will result in an Exception in the Spark pipeline which will then fail and halt.

  2. B

    The Spark pipeline will not be able to start again from previously committed offset by restarting itself, resulting in at least-once processing semantics

  3. C

    Irrespective of whatever machine fails, Spark will throw an error and halt.

  4. D

    The pipeline will be restarted automatically by Spark which is able to pick up the exact data from Kafka which was being processed at the time of error, resulting in exactly- once semantics.

  5. E

    Data that is being processed will not be processed again, resulting in atmost- once semantics.

Also asked in End Term 3 Sept 2023, End Term 28 Apr 2024, End Term 1 Sept 2024, End Term 1 Sept 2024

Question 2

+2 marksOne correct option

A big data streaming application that uses Kafka as source is observed to be really slow. The Kafka cluster has 2 broker nodes and this application is reading from 1 topic that has 10 partitions. On closer investigation, it was found that Kafka is not scaling to the velocity of input data coming in. What can you first try to do to scale Kafka further while incurring minimal costs?

  1. A

    Add disks to each broker in the cluster, and disks are the cheapest computer component

  2. B

    Add new brokers to the cluster, even though this is more expensive than the other options this is the only foolproof way to scale.

  3. C

    Increase memory in each of the brokers in the cluster. While cost of memory is more than cost of disks, it is still cheaper than adding brokers and helps to scale.

  4. D

    Create more topics and change input application to reroute data to all topics to be able to spread input data better. This is nearly the least expensive since only developer effort is required to change application.

  5. E

    Double the number of partitions for this single topic to be able to spread input data better. This is the least expensive since only administrator effort is required without changing application.

Also asked in End Term 3 Sept 2023, End Term 1 Sept 2024

Question 3

+2 marksOne correct option

A company with headquarters (HQ) in the Middle East operates on a Sunday-Thursday weekday schedule with Friday & Saturday as weekend days. It computes end of week revenue numbers using an ETL pipeline by first computing sales for each day at 1AM local time of the next day, and then summing up the weekly sales every Sunday early morning at 3AM local time. This number gets reported to leadership every Sunday morning 9AM local time, so the ETL pipeline is scheduled to run every Sunday morning at 8AM local time. As a result of management change, it has decided to relocate its HQ to India. Which of the following changes will need to be done to its ETL pipeline to ensure the correct output continues to be produced?

  1. A

    The business time for the final weekly sum operation needs to be changed to that of Monday 3AM India time instead of Sunday 3AM Middle East time.

  2. B

    Since time zone has changed as well as week definition too, the definition of business time has changed. So, the ETL has to be rewritten entirely.

  3. C

    Nothing needs to change since daily sales is available at 1AM Middle East time which is anyway behind India time and so the numbers will be available before leadership comes in at 9AM.

  4. D

    Event time has changed since the event of week ending has changed in definition, and so the ETL needs to be changed to consider the new event in the data.

  5. E

    No change required since neither event time nor business time is changing whereas only the operational time is changing.

Also asked in End Term 3 Sept 2023, End Term 1 Sept 2024

20 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 24 Dec 2023 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 24 Dec 2023, in the September 2023 term: 23 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 24 Dec 2023 at a glance
TermSeptember 2023 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions23
Marks50
Duration180 min
MCQ22
MSQ1
Official paperIIT M DEGREE AN EXAM ADB3 24 Dec 2023
Negative markingNo negative marking.
Updated

Same End Term, other subjects

More Intro to Big Data