Quiz Space

Introduction to Big Data End Term: 3 September 2023 (May 2023 term)

Question 1

+3 marksOne or more correct options

What are some capabilities common to both “streaming processing” & “batch processing” in the context of big data?

Select all that apply.

  1. A

    Batch operates on a set of data elements taken together while streaming can also operate on a set of data elements as determined by the window

  2. B

    Batch operates on data that is static while streaming operates on data that is dynamically changing

  3. C

    Batch processing can assume data as fully specified and complete while streaming cannot make that assumption

  4. D

    Batch processing is typically high latency while streaming processing is necessary for real-time latencies

  5. E

    Both Streaming & Batch processing can operate on massively large data sets

Also asked in End Term 1 Sept 2024

Question 2

+3 marksOne or more correct options

Which of the following are best practices associated with Streaming applications?

Select all that apply.

  1. A

    Use a message store that supports message replay so that no data is lost in processing.

  2. B

    “Hot potato” principle is when the streaming application operates on data from cache (i.e. the hot area of memory) and therefore is able to produce very high throughput

  3. C

    Hadoop is best suited for executing Streaming applications

  4. D

    Use checkpointing when faced with mission-critical workloads that require 100% accuracy.

  5. E

    Handle state pollution by restarting the persistent store software periodically.

Also asked in End Term 24 Dec 2023, End Term 1 Sept 2024

Question 3

+3 marksOne or more correct options

You are given the task of improving the performance of a Spark SQL program. You suspect that the culprit is the main transformation job in the program. When you run EXPLAIN on that SQL, you see that Spark wrongly estimates that there are only 10 values for the key being aggregated, whereas in reality the underlying data has a million values for that key. What actions would you perform from the below to ensure that the right estimates are used?

Select all that apply.

  1. A

    Create all tables as external tables.

  2. B

    Ensure cost based optimizer (CBO) is ON.

  3. C

    Partition all tables on the same key on which the aggregate is happening.

  4. D

    Run ANALYZE on all tables.

  5. E

    Cache the table in a step with actions ahead of the SQL statement that is the culprit.

Also asked in End Term 1 Sept 2024

18 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Intro to Big Data End Term 3 Sept 2023 paper

The IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 3 Sept 2023, in the May 2023 term: 21 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureIntro to Big Data End Term 3 Sept 2023 at a glance
TermMay 2023 term
SubjectIntroduction to Big Data
Course codeBSDA5001
Questions21
Marks50
Duration180 min
MSQ4
MCQ17
Official paperIIT M DEGREE ET1 EXAM QPE1 S2 03 Sep
Negative markingNo negative marking.
Updated

Same End Term, other subjects

More Intro to Big Data