Question 16
Below is a snippet for loading the GPT-2 model configuration:
from transformers import GPT2Config
config = GPT2Config.from_pretrained("gpt2-medium")print(config)the output of the above is
GPT2Config { "activation_function": "gelu_new", "architectures": [ "GPT2LMHeadModel" ], "attn_pdrop": 0.1, "bos_token_id": 50256, "embd_pdrop": 0.1, "eos_token_id": 50256, "initializer_range": 0.02, "layer_norm_epsilon": 1e-05, "model_type": "gpt2", "n_ctx": 1024, "n_embd": 1024, "n_head": 16, "n_inner": null, "n_layer": 24, "n_positions": 1024, "n_special": 0, "predict_special_tokens": true, "reorder_and_upcast_attn": false, "resid_pdrop": 0.1, "scale_attn_by_inverse_layer_idx": false, "scale_attn_weights": true, "summary_activation": null, "summary_first_dropout": 0.1, "summary_proj_to_labels": true, "summary_type": "cls_index", "summary_use_proj": true, "task_specific_params": { "text-generation": { "do_sample": true, "max_length": 50 } }, "transformers_version": "4.47.1", "use_cache": true, "vocab_size": 50257 }Based on the above data, answer the given subquestions.
What does the n_ctx parameter in the GPT-2 configuration represent?
The number of attention heads in the model.
The maximum length of the input sequence in tokens.
The embedding size of each token.
The total number of parameters in the model.