uiz Space

January 2025 term · Big Data and Biological Networks · BSBT4002

Big Data and Biological Networks End Term: 13 April 2025 (January 2025 term)

The IIT Madras BS Big Data and Biological Networks (Big Data and Biological Networks) End Term paper sat on 13 Apr 2025, in the January 2025 term: 44 questions for 100 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.

Questions
44
Marks
100
Duration
180 min
MCQ
22
MSQ
16
Written
2
Numerical
4

Updated

Official paper: IIT M IMPROVEMENT AN EXAM QIM3 13 Apr 2025 · No negative marking.

Question 1

+1 markOne correct option

The bond between two amino acids is called:

  1. A

    Peptide bond

  2. B

    Glycosidic bond

  3. C

    Phosphodiester bond

  4. D

    Amine bond

Show answer

Correct answer

  • A

    Peptide bond

Question 2

+1 markOne correct option

How many codons in the genetic code translate to amino acids?

  1. A

    16

  2. B

    38

  3. C

    61

  4. D

    26

Show answer

Correct answer

  • C

    61

Question 3

+1 markOne correct option

Which of the following is a key characteristic of post-translational modifications?

  1. A

    They alter the DNA sequence itself.

  2. B

    They are heritable changes in gene expression that do not involve alterations to the DNA sequence.

  3. C

    For a given gene, they only affect protein structure, not its expression.

  4. D

    They are limited to changes in mRNA splicing.

Show answer

Correct answer

  • C

    For a given gene, they only affect protein structure, not its expression.

Question 4

+1 markOne correct option

Which of the following best describes the fundamental principle behind a Genome-Wide Association Study (GWAS)?

  1. A

    Sequencing the entire genome of individuals with and without a specific trait to identify all causal variants

  2. B

    Comparing the frequency of genetic variants across the entire genome between individuals with a specific trait and a control group to identify variants statistically associated with the trait.

  3. C

    Studying the inheritance patterns of specific candidate genes within families affected by a particular disease.

  4. D

    Identifying rare, high-impact mutations that are directly responsible for causing a specific trait or disease.

Show answer

Correct answer

  • B

    Comparing the frequency of genetic variants across the entire genome between individuals with a specific trait and a control group to identify variants statistically associated with the trait.

Question 5

+1 markOne correct option

Enumerate the total number of strings that can be generated with 20 words?

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • C

Question 6

+1 markOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • A

Question 7

+1 markOne correct option
  1. A

    To learn embeddings for individual nodes

  2. B

    To map a set of node embeddings to a graph-level embedding

  3. C

    To optimize the message-passing mechanism between nodes

  4. D

    To improve the adjacency matrix representation

Show answer

Correct answer

  • B

    To map a set of node embeddings to a graph-level embedding

Question 8

+2 marksOne or more correct options

Which of the following is false?

Select all that apply.

  1. A

    Transcription uses two strands of DNA

  2. B

    Polyadenylation does not happen in prokaryotic cells

  3. C

    Proteins are produced in the nuclei

  4. D

    Intronic genes are only present in eukaryotic cells

Show answer

Correct answers

  • A

    Transcription uses two strands of DNA

  • C

    Proteins are produced in the nuclei

Question 9

+2 marksOne or more correct options

Applications of transcriptomics include:

Select all that apply.

  1. A

    Understanding functional impact of mutations

  2. B

    Identifying genetic drivers of disease

  3. C

    Studying response of organisms to environmental changes

  4. D

    Studying the effect of chromatin modifications

Show answer

Correct answers

  • A

    Understanding functional impact of mutations

  • C

    Studying response of organisms to environmental changes

Question 10

+2 marksOne or more correct options

Which of the following mutations could not affect gene expression?

Select all that apply.

  1. A

    Mutations in promoters

  2. B

    Mutations in the coding region

  3. C

    Mutations in RNA polymerases

  4. D

    Mutations in ribosomal proteins

Show answer

Correct answers

  • B

    Mutations in the coding region

  • D

    Mutations in ribosomal proteins

Question 11

+2 marksOne or more correct options

Which of the following distinguish Next-Generation Sequencing from traditional Sanger sequencing?

Select all that apply.

  1. A

    Ability to sequence millions of bases simultaneously

  2. B

    Significantly higher throughput and speed

  3. C

    Lower cost per base of sequencing

  4. D

    Reliance on chain-terminating dideoxynucleotides.

Show answer

Correct answers

  • A

    Ability to sequence millions of bases simultaneously

  • B

    Significantly higher throughput and speed

  • C

    Lower cost per base of sequencing

Question 12

+2 marksOne or more correct options

Which of the following biological processes involve the DNA:

Select all that apply.

  1. A

    Replication

  2. B

    Transcription

  3. C

    Translation

  4. D

    Epigenetic modification

Show answer

Correct answers

  • A

    Replication

  • B

    Transcription

  • D

    Epigenetic modification

Question 13

+2 marksOne or more correct options

The number of stuck/blocked reactions in the microbe ‘A’ when grown alone is 42. The number of stuck reactions is reduced to 18 when ‘A’ is grown together with ‘B’. Also, when ‘A’ is grown together with ‘C’, the number of stuck reactions in ‘A’ is 25 . Which of the following is/are true?

Select all that apply.

  1. A

    MSI(A;A U B) = 0.57

  2. B

    MSI(A;A U B) = 0.59

  3. C

    MSI(A;A U C) = 0.40

  4. D

    MSI(A;A U C) = 0.39

Show answer

Correct answers

  • A

    MSI(A;A U B) = 0.57

  • C

    MSI(A;A U C) = 0.40

Question 14

+2 marksOne or more correct options

Select all that apply.

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answers

  • B
  • D

Question 15

+2 marksOne or more correct options

Scientists are studying genetic interactions in yeast by creating knockout strains. They observe the following results:
● Single deletion of gene A: viable
● Single deletion of gene B: viable
● Single deletion of gene C: viable
● Single deletion of gene D: viable
● Double deletion of genes A and B: lethal
● Double deletion of genes B and C: viable
● Double deletion of genes C and D: lethal
● Double deletion of genes A and D: viable
Which of the following statements correctly describe synthetic lethal interactions in this system?

Select all that apply.

  1. A

    Genes A and B demonstrate synthetic lethality, suggesting they function in parallel pathways that compensate for each other.

  2. B

    Genes B and C show synthetic lethality, indicating they likely function in the same linear pathway.

  3. C

    Genes C and D show synthetic lethality, suggesting their functions might compensate for each other.

  4. D

    Genes A and D do not show synthetic lethality, proving they have no functional relationship whatsoever.

Show answer

Correct answers

  • A

    Genes A and B demonstrate synthetic lethality, suggesting they function in parallel pathways that compensate for each other.

  • C

    Genes C and D show synthetic lethality, suggesting their functions might compensate for each other.

Question 16

+2 marksOne or more correct options

Researchers conducted a network motif analysis on two biological networks (Network α and Network β) and obtained the following results for a specific subgraph pattern:

NetworkNetworkCount in Original NetworkMean Count in Random NetworksStandard Deviation in Random Networks
αbi-fan105106100
βbi-fan24016020

Which of the following statements are correct?

Select all that apply.

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answers

  • A
  • D

Question 17

+3 marksWritten answer

NOTE: Enter the exact answer without any extra space in the beginning or at the end.

Show answer

Correct answer: Met-Asn-Stop-Tyr

Question 18

+3 marksWritten answer

NOTE: Enter the exact answer without any extra space in the beginning or at the end.

Show answer

Correct answer: CCTGTATCTATTGT

Question 19

+3 marksOne correct option

In a Watts-Strogatz small-world network model with n nodes, k initial nearest neighbors per node (k ≪ n), and rewiring probability p, what happens to the average path length L(p) as the rewiring probability p increases from 0 to 1?

  1. A

    L(p) increases linearly with p

  2. B

    L(p) decreases rapidly at small p and then saturates

  3. C

    L(p) increases exponentially with p

  4. D

    L(p) remains constant regardless of p

Show answer

Correct answer

  • B

    L(p) decreases rapidly at small p and then saturates

Question 20

+3 marksOne correct option

A data science team is analyzing social network interactions to compare user engagement patterns between weekday and weekend activity. They have calculated differential network centrality metrics and obtained the following results for users U1-U4:

ProteinDegree CentralityBetweenness CentralityClustering Coefficient
U1+0.12+0.39-0.41
U2+0.72+0.68-0.49
U3-0.04+0.01+0.03
U4-0.38-0.47+0.58

(Values represent differences: weekend minus weekday)

Based on differential network analysis, which user most likely became a new "bridge" in the weekend network that connects previously separated community groups?

  1. A

    User U1, because it shows moderate increases in degree and betweenness centrality.

  2. B

    User U2, because it shows substantial increases in both degree and betweenness centrality coupled with decreased clustering, suggesting it has gained connections to diverse community clusters.

  3. C

    User U3, because its minimal changes across all metrics indicate stability, which is characteristic of core community members.

  4. D

    User U4, because the decrease in centrality metrics with increased clustering suggests it has become part of a tightly-knit community subgroup.

Show answer

Correct answer

  • B

    User U2, because it shows substantial increases in both degree and betweenness centrality coupled with decreased clustering, suggesting it has gained connections to diverse community clusters.

Question 21

+3 marksOne correct option

What is the average degree (regardless of the direction) of the 4-mer overlap graph constructed from “CAGATGTAT”?

  1. A

    2.22

  2. B

    2.45

  3. C

    3.12

  4. D

    1.74

Show answer

Correct answer

  • D

    1.74

Question 22

+3 marksOne correct option

What is the best global alignment score for the sequences "GTACTG" and "TAGTCTG" with the scoring scheme: matches = +2, mismatches = -1 and a constant gap penalty of -10.

  1. A

    -3

  2. B

    4

  3. C

    -12

  4. D

    6

Show answer

Correct answer

  • C

    -12

Question 23

+3 marksOne correct option

Elucidate the final sequence that is obtained from the kmers: TATT, TGAT, GATC, ATGT, GTAT, CATG, ATCA, TCAT, TGTA.

  1. A

    ATCTGTACAACTC

  2. B

    ATCGGTACAACTC

  3. C

    ATCGGGACAACTC

  4. D

    ATCGGTATAACTC

Show answer

Correct answer

  • B

    ATCGGTACAACTC

Question 24

+3 marksOne correct option

Consider the following statements:
Assertion (A): Dynamic Programming based approach is required to solve the sequence alignment problem.
Reason (R): DNA sequences are very long containing thousands of bases.
Which of the following options is true:

  1. A

    Both A and R are true and R is the correct explanation for A

  2. B

    Both A and R are true and R is not the correct explanation for A

  3. C

    A is true but R is false

  4. D

    Both the statements are false

Show answer

Correct answer

  • B

    Both A and R are true and R is not the correct explanation for A

Question 25

+3 marksOne correct option

Choose the correct option:

TaskCategory
Predicting whether a user is a bot in a social networkRelation Prediction
Property prediction based on molecular graph structuresMulti-label Graph Classification
Content recommendation in online platformsNode Classification
Predicting whether given social networks belong to sports, politics and technologyGraph Classification
  1. A

    All pairs are correct

  2. B

    Only the first three pairs are correct

  3. C

    Only the last two pairs are correct

  4. D

    None of the pairs are correct

Show answer

Correct answer

  • D

    None of the pairs are correct

Question 26

+3 marksOne correct option

Choose the most appropriate option:

Embedding MethodLearning ParadigmFeature Matrix
(1) DeepWalk
(2) GAT(A) Supervised(i) X=IX = I
(3) GraphSAGE(B) Unsupervised(ii) X≠IX \neq I
(4) Laplacian Eigenmaps

XX is the feature matrix and II is the identity matrix

  1. A

    (1), (A), (i); (2), (B), (ii); (3), (A), (ii); (4), (B), (ii);

  2. B

    (1), (B), (i); (2), (B), (i); (3), (A), (ii); (4), (A), (ii);

  3. C

    (1), (A), (ii); (2), (A), (i); (3), (B), (i); (4), (A), (i);

  4. D

    (1), (B), (i); (2), (A), (ii); (3), (A), (ii); (4), (B), (i);

Show answer

Correct answer

  • D

    (1), (B), (i); (2), (A), (ii); (3), (A), (ii); (4), (B), (i);

Question 27

+3 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • D

Question 28

+3 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • D

Question 29

+3 marksOne correct option

Consider a two-layer neural network with the following architecture:

h1=W1x+b1h_1 = W_1 x + b_1

z1=ReLU(h1)z_1 = \text{ReLU}(h_1)

h2=W2z1+b2h_2 = W_2 z_1 + b_2

y^=σ(h2)\hat{y} = \sigma(h_2)

ℓ=−ylog⁡(y^)−(1−y)log⁡(1−y^)\ell = -y \log(\hat{y}) - (1 - y) \log(1 - \hat{y})

Given the following numerical values:

W1=[0.90.20.10.8]W_1 = \begin{bmatrix} 0.9 & 0.2 \\ 0.1 & 0.8 \end{bmatrix}

b1=[0.10.2]b_1 = \begin{bmatrix} 0.1 \\ 0.2 \end{bmatrix}

W2=[0.80.10.50.9]W_2 = \begin{bmatrix} 0.8 & 0.1 \\ 0.5 & 0.9 \end{bmatrix}

b2=[0.780.45]b_2 = \begin{bmatrix} 0.78 \\ 0.45 \end{bmatrix}

x=[12]x = \begin{bmatrix} 1 \\ 2 \end{bmatrix}

What is the value of ∂h2∂b2\frac{\partial h_2}{\partial b_2} , ∂z1∂h1\frac{\partial z_1}{\partial h_1} and ∂h1∂x\frac{\partial h_1}{\partial x}? Note: InI_n is n×nn \times n Identity Matrix

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • D

Question 30

+3 marksOne correct option

A two-layer neural network with the 2 features, x1, x2x_1,\ x_2 is defined below:

h1=W1x+b1z1=ReLU(h1)h2=W2z1+b2y^=σ(h2)ℓ=−y.log(y^)−(1−y).log(1−y^)\begin{aligned} h_1 &= W_1 x + b_1 \\ z_1 &= ReLU(h_1) \\ h_2 &= W_2 z_1 + b_2 \\ \hat{y} &= \sigma(h_2) \\ \ell &= -y.log(\hat{y}) - (1 - y).log(1 - \hat{y}) \end{aligned}

Given the following numerical values:

W1=[0.90.20.10.8]W_1 = \begin{bmatrix} 0.9 & 0.2 \\ 0.1 & 0.8 \end{bmatrix}

b1=[0.10.2]b_1 = \begin{bmatrix} 0.1 \\ 0.2 \end{bmatrix}

W2=[0.80.10.50.9]W_2 = \begin{bmatrix} 0.8 & 0.1 \\ 0.5 & 0.9 \end{bmatrix}

b2=[0.780.45]b_2 = \begin{bmatrix} 0.78 \\ 0.45 \end{bmatrix}

x=[12]x = \begin{bmatrix} 1 \\ 2 \end{bmatrix}

y=1y = 1

Compute the values of ∂ℓ∂y^\frac{\partial \ell}{\partial \hat{y}} and ∂y^∂h2\frac{\partial \hat{y}}{\partial h_2}.

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • D

Question 31

+2 marksNumerical answer

A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:

P1P2P3P4
P10101
P21010
P30101
P41010

Based on the above data, answer the given subquestions.

When a decision task is initiated, neuron 3 shows elevated firing rates, which researchers believe coordinates the decision-making process. After analyzing the network topology, the degree centrality of protein 3 is ______________

Show answer

Correct answer: 0.67

Question 32

+1 markNumerical answer

A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:

P1P2P3P4
P10101
P21010
P30101
P41010

Based on the above data, answer the given subquestions.

The clustering coefficient of protein 3 in this network is ____________

Show answer

Correct answer: 0

Question 33

+1 markNumerical answer

A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:

P1P2P3P4
P10101
P21010
P30101
P41010

Based on the above data, answer the given subquestions.

The betweenness centrality of protein 3 is _____________

Show answer

Correct answer: 0.5

Question 34

+2 marksNumerical answer

In a social network analysis study, researchers are examining friendship connections between users on a new platform. Initially, they observe that each user is connected to exactly k = 49 nearby users in a neighborhood-based pattern, forming a structured network with n = 100 users. However, engagement data suggest that friend recommendations introduce a small-world topology, where some connections span across distant user groups, leading to a mix of local clustering and long-range connections. To better understand this connection pattern, the researchers hypothesize that the network might instead resemble a random graph of the form G(n,p), where p is the probability of a friendship forming between any two users. Estimate the probability p that would yield the same total number of connections as in the original structured network. Report 1000p

Show answer

Correct answer: 495

Question 35

+3 marksOne or more correct options

Which of the following concepts are fundamental to understanding and performing sequence alignment?

Select all that apply.

  1. A

    Gap penalties

  2. B

    Scoring matrices

  3. C

    Dynamic Programming

  4. D

    Random Number Generation

Show answer

Correct answers

  • A

    Gap penalties

  • B

    Scoring matrices

  • C

    Dynamic Programming

Question 36

+3 marksOne or more correct options

Which of the following is true about genome assembly?

Select all that apply.

  1. A

    It can be done only for prokaryotic genomes

  2. B

    Assembly comparing fragments obtained with one another

  3. C

    Assembling the genome using a De-Bruijn graph is easier than using an overlap graph

  4. D

    Success of assembly depends on how accurately we can sequence the fragments of DNA

Show answer

Correct answers

  • B

    Assembly comparing fragments obtained with one another

  • C

    Assembling the genome using a De-Bruijn graph is easier than using an overlap graph

  • D

    Success of assembly depends on how accurately we can sequence the fragments of DNA

Question 37

+3 marksOne or more correct options

Which of the following statements about GraphSAGE are true?

Select all that apply.

  1. A

    GraphSAGE is an inductive learning method.

  2. B

    GraphSAGE is a edge-level representation learning algorithm, while GAT is an node-level representation learning algorithm

  3. C

    GraphSAGE can generate embeddings for previously unseen nodes.

  4. D

    GraphSAGE uses a fixed-size neighborhood for aggregation

Show answer

Correct answers

  • A

    GraphSAGE is an inductive learning method.

  • C

    GraphSAGE can generate embeddings for previously unseen nodes.

  • D

    GraphSAGE uses a fixed-size neighborhood for aggregation

Question 38

+3 marksOne or more correct options

Task-specific labels are scarce during the training phase. The pretraining helps to develop a general understanding of related enumerate with abundant unlabelled data. Then, fine-tuning can be performed for a downstream task of interest. For the molecular property prediction task, which of the following are valid pretraining strategies.

Select all that apply.

  1. A

    Reconstructing the adjacency matrix for each graph using an encoder-decoder approach

  2. B

    Mask nodes and predict the features of masked nodes during training

  3. C

    Predict the embedding similarity between different graphs

  4. D

    Use subgraphs to predict surrounding graph structures

Show answer

Correct answers

  • A

    Reconstructing the adjacency matrix for each graph using an encoder-decoder approach

  • B

    Mask nodes and predict the features of masked nodes during training

  • D

    Use subgraphs to predict surrounding graph structures

Question 39

+3 marksOne or more correct options

Select all that apply.

  1. A
  2. B
  3. C
  4. D
Show answer

Correct answers

  • A
  • C

Question 40

+3 marksOne or more correct options

What are the limitations of shallow embeddings?

Select all that apply.

  1. A

    Shallow embedding approaches do not leverage node features in the encoder

  2. B

    Shallow embedding methods share parameters between nodes in the encoder

  3. C

    Shallow embedding methods can only generate embeddings for nodes that were present during the training phase

  4. D

    Shallow embedding methods directly optimizes a unique embedding vector for each node

Show answer

Correct answers

  • A

    Shallow embedding approaches do not leverage node features in the encoder

  • C

    Shallow embedding methods can only generate embeddings for nodes that were present during the training phase

  • D

    Shallow embedding methods directly optimizes a unique embedding vector for each node

Question 41

+3 marksOne or more correct options

Which of the following are true for graph filters:

Select all that apply.

  1. A

    Cheby-Filter is a spatial-based graph filter

  2. B

    Spectral-based graph filters utilize spectral graph theory to design the filtering operation in the spectral domain

  3. C

    Spatial-based graph filters explicitly leverage the graph structure

  4. D

    GAT-Filter is a spatial-based graph filter

Show answer

Correct answers

  • B

    Spectral-based graph filters utilize spectral graph theory to design the filtering operation in the spectral domain

  • C

    Spatial-based graph filters explicitly leverage the graph structure

  • D

    GAT-Filter is a spatial-based graph filter

Question 42

+2 marksOne correct option
  1. A
  2. B
  3. C
  4. D
Show answer

Correct answer

  • C

Question 43

+2 marksOne correct option

What is the primary challenge in applying classical deep neural networks to graph structured data?

  1. A

    The data is too large.

  2. B

    The graph structure is not a regular grid.

  3. C

    The data lacks features.

  4. D

    The models are too complex.

Show answer

Correct answer

  • B

    The graph structure is not a regular grid.

Question 44

+2 marksOne correct option

What is the primary purpose of introducing a normalization step in a neural network

  1. A

    Enhances model regularization

  2. B

    Ensures differentiability of the function

  3. C

    Reduces variance in the distribution across mini-batches

  4. D

    All of these

Show answer

Correct answer

  • C

    Reduces variance in the distribution across mini-batches