Quiz Space

Introduction to Big Data · Quiz 2 · 3 Dec 2023 · September 2023 term

Intro to Big Data Quiz 2 3 Dec 2023 — Question 23

Question 23

+4 marksOne or more correct options

Which of the following is true?

Select all that apply.

  1. A

    Snapshots of source systems can be created using an event capture tool like CDC and then replaying the events in sequence for that time period.

  2. B

    A data lake is a collection of data to be provided as input for data science algorithms.

  3. C

    Zookeeper is a library for ensuring all services used in big data are monitored, administered and cleaned at appropriate intervals.

  4. D

    A single Spark cluster can have multiple “leader” master nodes.

  5. E

    Spark is optimized for in-memory computation.

  6. F

    Given the RDD underlying a Dataframe, you can recreate the same Dataframe provided you know the schema

Show answer

Correct answers

  • A

    Snapshots of source systems can be created using an event capture tool like CDC and then replaying the events in sequence for that time period.

  • E

    Spark is optimized for in-memory computation.

  • F

    Given the RDD underlying a Dataframe, you can recreate the same Dataframe provided you know the schema

Question 23 of 26 in the IIT Madras BS Introduction to Big Data (Intro to Big Data) Quiz 2 paper sat on 3 Dec 2023, in the September 2023 term (IIT M DEGREE AN2 EXAM QDB2 03 Dec 2023). It carries 4 marks.

This question was also asked in

More questions from this paper

  1. Q1What best describes "big data"?
  2. Q2Which of these represent examples of divide-and-conquer?
  3. Q3Which of the following statements about Spark application architecture is correct?
  4. Q4Consider the problem of sorting a 1 petabyte file of numbers stored in a Hadoop cluster of 10 machines where the size o…
  5. Q5A website sees 1 Billion hits every month. The website owner wants to count average hits per customer in the latest mon…
  6. Q6An enterprise software designer wants to leverage the best of cloud to minimize the number of administrative overheads …
  7. Q7Consider an application that can scale from handling 1000 users to handling 100 million users by simply making copies o…
  8. Q8Linux is an example of an operating system. Which of the following is considered as the "operating system of a cluster …
  9. Q9Figure question
  10. Q10Figure question
  11. Q11What happens behind the scenes when the following code is run?
  12. Q12Consider a file “data.bin” which is formatted as follows: every data record is in the form of pairs of values of the fo…
  13. Q13Consider the program outline as below running on a Spark cluster of 1 driver and 4 worker nodes with 2 executors per wo…
  14. Q14What is the output of the following code?
  15. Q15What is the output of the following code when deployed on a spark cluster?
  16. Q16What is the output of the following code?
  17. Q17What is the purpose of the cache() function of an RDD in PySpark?
  18. Q18What role does metadata play in a Data Lake?
  19. Q19What option(s) best describe the differences between MapReduce and Spark?
  20. Q20Would IRCTC's railway ticket booking application be suitable for a serverless implementation?
  21. Q21You are provided with a Spark program that picks out a list of suspicious transactions based on the amount of the trans…
  22. Q22Which of the following types of data sources can you read successfully without missing data using a program that extrac…
  23. Q24Which of the following statements is/ are true for Google Cloud Functions?
  24. Q25Which of the following is True ?
  25. Q26Which of the following is True about YARN ?