Question 1
What happens when a Spark Structured Streaming pipeline operating with Kafka as source is subject to a failure of a machine in either of the Kafka cluster or the Spark cluster?
Failure of a machine in the Kafka cluster will result in an Exception in the Spark pipeline which will then fail and halt.
The Spark pipeline will not be able to start again from previously committed offset by restarting itself, resulting in at least-once processing semantics
Irrespective of whatever machine fails, Spark will throw an error and halt.
The pipeline will be restarted automatically by Spark which is able to pick up the exact data from Kafka which was being processed at the time of error, resulting in exactly- once semantics.
Data that is being processed will not be processed again, resulting in atmost- once semantics.