Deep Learning Practice, Quiz 2
Consider the following code using Wav2Vec2Processor to process an audio sample:
import torchfrom transformers import Wav2Vec2Processor
processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-base-960h")
sample_audio = {"array": torch.randn(12000), "sampling_rate": 16000}
inputs = processor(sample_audio["array"], sampling_rate=16000, return_tensors="pt", padding=True)
print(inputs["input_values"].shape)What will the printed shape be?
Consider the following code using Wav2Vec2Processor to process an audio sample: import torch from transformers import Wav2Vec2Processor processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-base-960h") sample_audio = {"array": torch.randn(12000), "sampling_rate": 16000} inputs = processor(sample_audio["array"], sampling_rate=16000, return_tensors="pt", padding=True) print(inputs["input_values"].shape) What will the printed shape be? Consider the following code snippet: from transformers import Wav2Vec2Processor processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-large-960h") input_values = processor("audio.wav", return_tensors="pt", sampling_rate=16000).input_values What is the primary mistake in this code? You have a model that predicts transcriptions for audio clips. You calculate Word Error Rate (WER) using the jiwer package as follows: from jiwer import wer ground_truth = ["hello world", "this is a test"] predictions = ["hello word", "this test"] error = wer(ground_truth, predictions) print(f"WER: {error:.2f}") What will be the output of this code?