Opening the paper…
Large Language Models, End Term
What happens when you use Batch Normalization with a batch size of 1?
What happens when you use Batch Normalization with a batch size of 1? A language model outputs the following logits for the next token: \{cat: 3.2, dog: 2.9, bird: 1.1, fish: 0.5, snake: -0.4\} Before sampling, the decoding pipeline performs: • Temperature scaling with T = 2.0 • Top-K filtering with K = 3 After applying both steps, which tokens remain eligible for sampling? Question text from the original paper, with its maths as pictures