Question 19
Which of the following is true?
Snapshots of source systems can be created using an event capture tool likeCDC and then replaying the events in sequence for that time period.
A data lake is a collection of data to be provided as input for data sciencealgorithms.
Zookeeper is a library for ensuring all services used in big data are monitored,administered and cleaned at appropriate intervals.
A single Spark cluster can have multiple “leader” master nodes.
Spark is optimized for in-memory computation.
Given the RDD underlying a Dataframe, you can recreate the same Dataframeprovided you know the schema