Quiz Space

Introduction to Big Data · End Term · 22 Dec 2024 · September 2024 term · Set QDB4

Question 15: You are given the task of improving the performance of a…

Question 15

+3 marksOne or more correct options

You are given the task of improving the performance of a Spark SQL program. You suspect that the culprit is the main transformation job in the program. When you run EXPLAIN on that SQL, you see that Spark wrongly estimates that there are only 10 values for the key being aggregated, whereas in reality the underlying data has a million values for that key. What actions would you perform from the below to ensure that the right estimates are used?

Select all that apply.

  1. A

    Create all tables as native Spark SQL tables (i.e. available as CatalogTables).

  2. B

    Partition all tables on the same key on which the aggregate is happening.

  3. C

    Ensure cost based optimizer (CBO) is ON.

  4. D

    Run ANALYZE on all tables.

  5. E

    Cache the table in a step with actions ahead of the SQL statement that is the culprit.

Show answer

Correct answers

  • A

    Create all tables as native Spark SQL tables (i.e. available as CatalogTables).

  • C

    Ensure cost based optimizer (CBO) is ON.

  • D

    Run ANALYZE on all tables.

Question 15 of 21 in the IIT Madras BS Introduction to Big Data (Intro to Big Data) End Term paper sat on 22 Dec 2024, in the September 2024 term (IIT M DEGREE AN EXAM QDB4 22 Dec 2024). It carries 3 marks.

More questions from this paper

  1. Q1What happens when a Spark Structured Streaming pipeline operating with Kafka as the source is subject to a machine fail…
  2. Q2A company with headquarters (HQ) in the Middle East operates on a Sunday-Thursday weekday schedule with Friday \& Satur…
  3. Q3In the class, we saw the UDF for mobilenet_v2. Specifically, the predict() function contained the below lines of code
  4. Q4Dhoni is on the crease with a bat in hand that has sensors embedded throughout. The sensors talk to the spider cam ever…
  5. Q5In the class, we saw the UDF for mobilenet_v2. By definition, UDFs are scalar. In Spark, there is another class of user…
  6. Q6You are appointed as a Data Engineer in a company that has a legacy reporting application written in Java which suffers…
  7. Q7You are given a Spark Streaming pipeline that invokes a pre-trained DL model for every image it receives as input and p…
  8. Q8Which of the following code snippets will give a runtime error? (Note: df is a spark dataframe. It has a column called …
  9. Q9Since it is the onset of summer, there is a surge in railway ticket bookings. The business head at IRCTC is interested …
  10. Q10Consider a Structured Streaming application running on Google Dataproc firing up every 10 seconds, consuming any number…
  11. Q11Consider a Kafka system that has two brokers with the exact same specifications. Let’s consider a topic A with a single…
  12. Q12Let us say we are using structured streaming for continuously reading data from Kafka and storing the results back into…
  13. Q13Kubernetes is an open-source system for automating deployment, scaling and management of containerized applications. Go…
  14. Q14A big data streaming application that reads from using Kafka as source is observed to be really slow. The Kafka cluster…
  15. Q16What differentiates “streaming processing” from “batch processing” in the context of big data?
  16. Q17What are the core components for a Big Data Streaming application?
  17. Q18Which of the following statements are true?
  18. Q19Observe the below image and select the options that are true.
  19. Q20A Spark Streaming application is configured to execute once per minute. However, each run takes 10+ minutes consistentl…
  20. Q21You are given a Spark program that runs on a Google Dataproc cluster on a daily schedule from 1AM-12PM to produce as ou…