Question 8
What is the primary capability that the Position-wise Feed-Forward Network (FFN) provides to the Transformer architecture?
The ability to attend to other tokens in the sequence and capture dependencies between words
The ability to preserve and utilize positional information about token order
The ability to apply non-linear transformations and learn complex feature interactions within each token's representation
The ability to mix information across different positions in the sequence