Question 1
The bond between two amino acids is called:
Peptide bond
Glycosidic bond
Phosphodiester bond
Amine bond

The IIT Madras BS Big Data and Biological Networks (Big Data and Biological Networks) End Term paper sat on 13 Apr 2025, in the January 2025 term: 44 questions for 100 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
The bond between two amino acids is called:
Peptide bond
Glycosidic bond
Phosphodiester bond
Amine bond
Correct answer
Peptide bond
How many codons in the genetic code translate to amino acids?
16
38
61
26
Correct answer
61
Which of the following is a key characteristic of post-translational modifications?
They alter the DNA sequence itself.
They are heritable changes in gene expression that do not involve alterations to the DNA sequence.
For a given gene, they only affect protein structure, not its expression.
They are limited to changes in mRNA splicing.
Correct answer
For a given gene, they only affect protein structure, not its expression.
Which of the following best describes the fundamental principle behind a Genome-Wide Association Study (GWAS)?
Sequencing the entire genome of individuals with and without a specific trait to identify all causal variants
Comparing the frequency of genetic variants across the entire genome between individuals with a specific trait and a control group to identify variants statistically associated with the trait.
Studying the inheritance patterns of specific candidate genes within families affected by a particular disease.
Identifying rare, high-impact mutations that are directly responsible for causing a specific trait or disease.
Correct answer
Comparing the frequency of genetic variants across the entire genome between individuals with a specific trait and a control group to identify variants statistically associated with the trait.
Enumerate the total number of strings that can be generated with 20 words?
Correct answer
Correct answer
To learn embeddings for individual nodes
To map a set of node embeddings to a graph-level embedding
To optimize the message-passing mechanism between nodes
To improve the adjacency matrix representation
Correct answer
To map a set of node embeddings to a graph-level embedding
Which of the following is false?
Transcription uses two strands of DNA
Polyadenylation does not happen in prokaryotic cells
Proteins are produced in the nuclei
Intronic genes are only present in eukaryotic cells
Correct answers
Transcription uses two strands of DNA
Proteins are produced in the nuclei
Applications of transcriptomics include:
Understanding functional impact of mutations
Identifying genetic drivers of disease
Studying response of organisms to environmental changes
Studying the effect of chromatin modifications
Correct answers
Understanding functional impact of mutations
Studying response of organisms to environmental changes
Which of the following mutations could not affect gene expression?
Mutations in promoters
Mutations in the coding region
Mutations in RNA polymerases
Mutations in ribosomal proteins
Correct answers
Mutations in the coding region
Mutations in ribosomal proteins
Which of the following distinguish Next-Generation Sequencing from traditional Sanger sequencing?
Ability to sequence millions of bases simultaneously
Significantly higher throughput and speed
Lower cost per base of sequencing
Reliance on chain-terminating dideoxynucleotides.
Correct answers
Ability to sequence millions of bases simultaneously
Significantly higher throughput and speed
Lower cost per base of sequencing
Which of the following biological processes involve the DNA:
Replication
Transcription
Translation
Epigenetic modification
Correct answers
Replication
Transcription
Epigenetic modification
The number of stuck/blocked reactions in the microbe ‘A’ when grown alone is 42. The number of stuck reactions is reduced to 18 when ‘A’ is grown together with ‘B’. Also, when ‘A’ is grown together with ‘C’, the number of stuck reactions in ‘A’ is 25 . Which of the following is/are true?
MSI(A;A U B) = 0.57
MSI(A;A U B) = 0.59
MSI(A;A U C) = 0.40
MSI(A;A U C) = 0.39
Correct answers
MSI(A;A U B) = 0.57
MSI(A;A U C) = 0.40
Correct answers
Scientists are studying genetic interactions in yeast by creating knockout strains. They observe the following results:
● Single deletion of gene A: viable
● Single deletion of gene B: viable
● Single deletion of gene C: viable
● Single deletion of gene D: viable
● Double deletion of genes A and B: lethal
● Double deletion of genes B and C: viable
● Double deletion of genes C and D: lethal
● Double deletion of genes A and D: viable
Which of the following statements correctly describe synthetic lethal interactions in this system?
Genes A and B demonstrate synthetic lethality, suggesting they function in parallel pathways that compensate for each other.
Genes B and C show synthetic lethality, indicating they likely function in the same linear pathway.
Genes C and D show synthetic lethality, suggesting their functions might compensate for each other.
Genes A and D do not show synthetic lethality, proving they have no functional relationship whatsoever.
Correct answers
Genes A and B demonstrate synthetic lethality, suggesting they function in parallel pathways that compensate for each other.
Genes C and D show synthetic lethality, suggesting their functions might compensate for each other.
Researchers conducted a network motif analysis on two biological networks (Network α and Network β) and obtained the following results for a specific subgraph pattern:
| Network | Network | Count in Original Network | Mean Count in Random Networks | Standard Deviation in Random Networks |
|---|---|---|---|---|
| α | bi-fan | 105 | 106 | 100 |
| β | bi-fan | 240 | 160 | 20 |
Which of the following statements are correct?
Correct answers
NOTE: Enter the exact answer without any extra space in the beginning or at the end.
Correct answer: Met-Asn-Stop-Tyr
NOTE: Enter the exact answer without any extra space in the beginning or at the end.
Correct answer: CCTGTATCTATTGT
In a Watts-Strogatz small-world network model with n nodes, k initial nearest neighbors per node (k ≪ n), and rewiring probability p, what happens to the average path length L(p) as the rewiring probability p increases from 0 to 1?
L(p) increases linearly with p
L(p) decreases rapidly at small p and then saturates
L(p) increases exponentially with p
L(p) remains constant regardless of p
Correct answer
L(p) decreases rapidly at small p and then saturates
A data science team is analyzing social network interactions to compare user engagement patterns between weekday and weekend activity. They have calculated differential network centrality metrics and obtained the following results for users U1-U4:
| Protein | Degree Centrality | Betweenness Centrality | Clustering Coefficient |
|---|---|---|---|
| U1 | +0.12 | +0.39 | -0.41 |
| U2 | +0.72 | +0.68 | -0.49 |
| U3 | -0.04 | +0.01 | +0.03 |
| U4 | -0.38 | -0.47 | +0.58 |
(Values represent differences: weekend minus weekday)
Based on differential network analysis, which user most likely became a new "bridge" in the weekend network that connects previously separated community groups?
User U1, because it shows moderate increases in degree and betweenness centrality.
User U2, because it shows substantial increases in both degree and betweenness centrality coupled with decreased clustering, suggesting it has gained connections to diverse community clusters.
User U3, because its minimal changes across all metrics indicate stability, which is characteristic of core community members.
User U4, because the decrease in centrality metrics with increased clustering suggests it has become part of a tightly-knit community subgroup.
Correct answer
User U2, because it shows substantial increases in both degree and betweenness centrality coupled with decreased clustering, suggesting it has gained connections to diverse community clusters.
What is the average degree (regardless of the direction) of the 4-mer overlap graph constructed from “CAGATGTAT”?
2.22
2.45
3.12
1.74
Correct answer
1.74
What is the best global alignment score for the sequences "GTACTG" and "TAGTCTG" with the scoring scheme: matches = +2, mismatches = -1 and a constant gap penalty of -10.
-3
4
-12
6
Correct answer
-12
Elucidate the final sequence that is obtained from the kmers: TATT, TGAT, GATC, ATGT, GTAT, CATG, ATCA, TCAT, TGTA.
ATCTGTACAACTC
ATCGGTACAACTC
ATCGGGACAACTC
ATCGGTATAACTC
Correct answer
ATCGGTACAACTC
Consider the following statements:
Assertion (A): Dynamic Programming based approach is required to solve the sequence alignment problem.
Reason (R): DNA sequences are very long containing thousands of bases.
Which of the following options is true:
Both A and R are true and R is the correct explanation for A
Both A and R are true and R is not the correct explanation for A
A is true but R is false
Both the statements are false
Correct answer
Both A and R are true and R is not the correct explanation for A
Choose the correct option:
| Task | Category |
|---|---|
| Predicting whether a user is a bot in a social network | Relation Prediction |
| Property prediction based on molecular graph structures | Multi-label Graph Classification |
| Content recommendation in online platforms | Node Classification |
| Predicting whether given social networks belong to sports, politics and technology | Graph Classification |
All pairs are correct
Only the first three pairs are correct
Only the last two pairs are correct
None of the pairs are correct
Correct answer
None of the pairs are correct
Choose the most appropriate option:
| Embedding Method | Learning Paradigm | Feature Matrix |
|---|---|---|
| (1) DeepWalk | ||
| (2) GAT | (A) Supervised | (i) |
| (3) GraphSAGE | (B) Unsupervised | (ii) |
| (4) Laplacian Eigenmaps |
is the feature matrix and is the identity matrix
(1), (A), (i); (2), (B), (ii); (3), (A), (ii); (4), (B), (ii);
(1), (B), (i); (2), (B), (i); (3), (A), (ii); (4), (A), (ii);
(1), (A), (ii); (2), (A), (i); (3), (B), (i); (4), (A), (i);
(1), (B), (i); (2), (A), (ii); (3), (A), (ii); (4), (B), (i);
Correct answer
(1), (B), (i); (2), (A), (ii); (3), (A), (ii); (4), (B), (i);
Correct answer
Correct answer
Consider a two-layer neural network with the following architecture:
Given the following numerical values:
What is the value of , and ? Note: is Identity Matrix
Correct answer
A two-layer neural network with the 2 features, is defined below:
Given the following numerical values:
Compute the values of and .
Correct answer
A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:
| P1 | P2 | P3 | P4 | |
|---|---|---|---|---|
| P1 | 0 | 1 | 0 | 1 |
| P2 | 1 | 0 | 1 | 0 |
| P3 | 0 | 1 | 0 | 1 |
| P4 | 1 | 0 | 1 | 0 |
Based on the above data, answer the given subquestions.
When a decision task is initiated, neuron 3 shows elevated firing rates, which researchers believe coordinates the decision-making process. After analyzing the network topology, the degree centrality of protein 3 is ______________
Correct answer: 0.67
A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:
| P1 | P2 | P3 | P4 | |
|---|---|---|---|---|
| P1 | 0 | 1 | 0 | 1 |
| P2 | 1 | 0 | 1 | 0 |
| P3 | 0 | 1 | 0 | 1 |
| P4 | 1 | 0 | 1 | 0 |
Based on the above data, answer the given subquestions.
The clustering coefficient of protein 3 in this network is ____________
Correct answer: 0
A research team studying neural connectivity in the prefrontal cortex has mapped interactions between key neurons involved in decision-making processes. They modeled the network using a modified Watts-Strogatz approach with N=4 nodes (neurons) and the following adjacency matrix:
| P1 | P2 | P3 | P4 | |
|---|---|---|---|---|
| P1 | 0 | 1 | 0 | 1 |
| P2 | 1 | 0 | 1 | 0 |
| P3 | 0 | 1 | 0 | 1 |
| P4 | 1 | 0 | 1 | 0 |
Based on the above data, answer the given subquestions.
The betweenness centrality of protein 3 is _____________
Correct answer: 0.5
In a social network analysis study, researchers are examining friendship connections between users on a new platform. Initially, they observe that each user is connected to exactly k = 49 nearby users in a neighborhood-based pattern, forming a structured network with n = 100 users. However, engagement data suggest that friend recommendations introduce a small-world topology, where some connections span across distant user groups, leading to a mix of local clustering and long-range connections. To better understand this connection pattern, the researchers hypothesize that the network might instead resemble a random graph of the form G(n,p), where p is the probability of a friendship forming between any two users. Estimate the probability p that would yield the same total number of connections as in the original structured network. Report 1000p
Correct answer: 495
Which of the following concepts are fundamental to understanding and performing sequence alignment?
Gap penalties
Scoring matrices
Dynamic Programming
Random Number Generation
Correct answers
Gap penalties
Scoring matrices
Dynamic Programming
Which of the following is true about genome assembly?
It can be done only for prokaryotic genomes
Assembly comparing fragments obtained with one another
Assembling the genome using a De-Bruijn graph is easier than using an overlap graph
Success of assembly depends on how accurately we can sequence the fragments of DNA
Correct answers
Assembly comparing fragments obtained with one another
Assembling the genome using a De-Bruijn graph is easier than using an overlap graph
Success of assembly depends on how accurately we can sequence the fragments of DNA
Which of the following statements about GraphSAGE are true?
GraphSAGE is an inductive learning method.
GraphSAGE is a edge-level representation learning algorithm, while GAT is an node-level representation learning algorithm
GraphSAGE can generate embeddings for previously unseen nodes.
GraphSAGE uses a fixed-size neighborhood for aggregation
Correct answers
GraphSAGE is an inductive learning method.
GraphSAGE can generate embeddings for previously unseen nodes.
GraphSAGE uses a fixed-size neighborhood for aggregation
Task-specific labels are scarce during the training phase. The pretraining helps to develop a general understanding of related enumerate with abundant unlabelled data. Then, fine-tuning can be performed for a downstream task of interest. For the molecular property prediction task, which of the following are valid pretraining strategies.
Reconstructing the adjacency matrix for each graph using an encoder-decoder approach
Mask nodes and predict the features of masked nodes during training
Predict the embedding similarity between different graphs
Use subgraphs to predict surrounding graph structures
Correct answers
Reconstructing the adjacency matrix for each graph using an encoder-decoder approach
Mask nodes and predict the features of masked nodes during training
Use subgraphs to predict surrounding graph structures
Correct answers
What are the limitations of shallow embeddings?
Shallow embedding approaches do not leverage node features in the encoder
Shallow embedding methods share parameters between nodes in the encoder
Shallow embedding methods can only generate embeddings for nodes that were present during the training phase
Shallow embedding methods directly optimizes a unique embedding vector for each node
Correct answers
Shallow embedding approaches do not leverage node features in the encoder
Shallow embedding methods can only generate embeddings for nodes that were present during the training phase
Shallow embedding methods directly optimizes a unique embedding vector for each node
Which of the following are true for graph filters:
Cheby-Filter is a spatial-based graph filter
Spectral-based graph filters utilize spectral graph theory to design the filtering operation in the spectral domain
Spatial-based graph filters explicitly leverage the graph structure
GAT-Filter is a spatial-based graph filter
Correct answers
Spectral-based graph filters utilize spectral graph theory to design the filtering operation in the spectral domain
Spatial-based graph filters explicitly leverage the graph structure
GAT-Filter is a spatial-based graph filter
Correct answer
What is the primary challenge in applying classical deep neural networks to graph structured data?
The data is too large.
The graph structure is not a regular grid.
The data lacks features.
The models are too complex.
Correct answer
The graph structure is not a regular grid.
What is the primary purpose of introducing a normalization step in a neural network
Enhances model regularization
Ensures differentiability of the function
Reduces variance in the distribution across mini-batches
All of these
Correct answer
Reduces variance in the distribution across mini-batches