Question 1
Which of the following distributions is/are not symmetric in nature (choose all that are applicable)?
Standard Normal distribution
Standard Binomial distribution
Uniform distribution
Poisson distribution

The IIT Madras BS Business Analytics (Business Analytics) End Term paper sat on 13 Apr 2025, in the January 2025 term, set QDD1: 33 questions for 45 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
Which of the following distributions is/are not symmetric in nature (choose all that are applicable)?
Standard Normal distribution
Standard Binomial distribution
Uniform distribution
Poisson distribution
Correct answer
Poisson distribution
Which of the following is/are required to build an empirical distribution? (choose all that are applicable)
PDF or PMF
Sample data
Summary Statistics
None of these
Correct answers
PDF or PMF
Sample data
Summary Statistics
Given the Chisquare table in Figure-2, what is the conclusion from the test at a 95% significance level? (choose all that may be applicable)
| .995 | .990 | .975 | .950 | .900 | .500 | .100 | .050 | .025 | .010 | .005 | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | .00+ | .00+ | .00+ | .00+ | .02 | .45 | 2.71 | 3.84 | 5.02 | 6.63 | 7.88 |
| 2 | .01 | .02 | .05 | .10 | .21 | 1.39 | 4.61 | 5.99 | 7.38 | 9.21 | 10.60 |
| 3 | .07 | .11 | .22 | .35 | .58 | 2.37 | 6.25 | 7.81 | 9.35 | 11.34 | 12.84 |
| 4 | .21 | .30 | .48 | .71 | 1.06 | 3.36 | 7.78 | 9.49 | 11.14 | 13.28 | 14.86 |
| 5 | .41 | .55 | .83 | 1.15 | 1.61 | 4.35 | 9.24 | 11.07 | 12.83 | 15.09 | 16.75 |
| 6 | .68 | .87 | 1.24 | 1.64 | 2.20 | 5.35 | 10.65 | 12.59 | 14.45 | 16.81 | 18.55 |
| 7 | .99 | 1.24 | 1.69 | 2.17 | 2.83 | 6.35 | 12.02 | 14.07 | 16.01 | 18.48 | 20.28 |
| 8 | 1.34 | 1.65 | 2.18 | 2.73 | 3.49 | 7.34 | 13.36 | 15.51 | 17.53 | 20.09 | 21.96 |
| 9 | 1.73 | 2.09 | 2.70 | 3.33 | 4.17 | 8.34 | 14.68 | 16.92 | 19.02 | 21.67 | 23.59 |
| 10 | 2.16 | 2.56 | 3.25 | 3.94 | 4.87 | 9.34 | 15.99 | 18.31 | 20.48 | 23.21 | 25.19 |
| 11 | 2.60 | 3.05 | 3.82 | 4.57 | 5.58 | 10.34 | 17.28 | 19.68 | 21.92 | 24.72 | 26.76 |
| 12 | 3.07 | 3.57 | 4.40 | 5.23 | 6.30 | 11.34 | 18.55 | 21.03 | 23.34 | 26.22 | 28.30 |
| 13 | 3.57 | 4.11 | 5.01 | 5.89 | 7.04 | 12.34 | 19.81 | 22.36 | 24.74 | 27.69 | 29.82 |
| 14 | 4.07 | 4.66 | 5.63 | 6.57 | 7.79 | 13.34 | 21.06 | 23.68 | 26.12 | 29.14 | 31.32 |
| 15 | 4.60 | 5.23 | 6.27 | 7.26 | 8.55 | 14.34 | 22.31 | 25.00 | 27.49 | 30.58 | 32.80 |
Figure-2
REJECT the NULL HYPOTHESIS and conclude that the number of defects FOLLOWS a poisson distribution
DO NOT REJECT the NULL HYPOTHESIS and conclude that the number of defects FOLLOWS a poisson distribution
REJECT the NULL HYPOTHESIS and conclude that the number of defects DOES NOT FOLLOW a poisson distribution
DO NOT REJECT the NULL HYPOTHESIS and conclude that the number of defects DOES NOT FOLLOW a poisson distribution
REJECT the ALTERNATIVE HYPOTHESIS and conclude that the number of defects FOLLOWS a poisson distribution
DO NOT REJECT the ALTERNATIVE HYPOTHESIS and conclude that the number of defects FOLLOWS a poisson distribution
REJECT the ALTERNATIVE HYPOTHESIS and conclude that the number of defects DOES NOT FOLLOW a poisson distribution
DO NOT REJECT the ALTERNATIVE HYPOTHESIS and conclude that the number of defects DOES NOT FOLLOW a poisson distribution
Correct answer
REJECT the NULL HYPOTHESIS and conclude that the number of defects DOES NOT FOLLOW a poisson distribution
If the attribute values in the conjoint analysis is a continuous variable and the data is collected in a pairwise order, then what approach can be used (choose all that is/are applicable)
Optimization approach
Regression approach
Statistical approach
None of these
Correct answer
Optimization approach
The part worth can be defined as (choose all that may be applicable)
Level utilities
The utility for that level of attribute
Utility for separate parts of the products
None of these
Correct answers
Level utilities
The utility for that level of attribute
Utility for separate parts of the products
In Figure-3 the customer wants to decide between the products O1 & O2, and x denotes the coordinates of the ideal product. Which of the following is/are true?
Customers will prefer O1 when d2>d1
Customers will prefer O2 when d1<d2
Both Customers will prefer O1 when d2>d1 & Customers will prefer O2 when d1<d2
None of these
Correct answer
Customers will prefer O1 when d2>d1
There are 7 business units and you are using the DEA to compare them. You solve the LP for business unit 5. You find from the constraint expression that business unit 1 has obtained an efficiency of 1 and business unit 2 has obtained an efficiency of 1 with the optimal weights of business unit 5. Which of the following statements is correct? (choose all that is/are applicable)
Business unit 5 is inefficient
Business unit 1 is efficient
Business unit 5 is efficient
Business unit 2 is efficient
Correct answers
Business unit 1 is efficient
Business unit 2 is efficient
In DEA, when can the Linear Programming model be used for calculating the weights of efficiency (weighted outputs/weighted inputs)? (choose all that is/are applicable)
After converting the ratio into the linear objective function
After normalizing the denominator
By setting a constraint on the efficiency of all DMUs to be lesser than or equal to 1
None of these
Correct answers
After converting the ratio into the linear objective function
After normalizing the denominator
By setting a constraint on the efficiency of all DMUs to be lesser than or equal to 1
(The following is a purely imaginary scenario)
A demand response curve is modelled using linear regression. The partial regression output is given in Figure-1 below. Given this information, answer the subquestions.
ANOVA
| df | SS | |
|---|---|---|
| Regression | 7076613 | |
| Residual | ||
| Total | 9 | 7184891 |
| Coefficients | Standard Error | |
|---|---|---|
| Intercept | 39942 | |
| X Variable 1 | -29.5 |
Figure-1
What is the elasticity of the demand response curve at a price of Rs. 50? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 0.035 (accepted within ±0.015)
(The following is a purely imaginary scenario)
A demand response curve is modelled using linear regression. The partial regression output is given in Figure-1 below. Given this information, answer the subquestions.
ANOVA
| df | SS | |
|---|---|---|
| Regression | 7076613 | |
| Residual | ||
| Total | 9 | 7184891 |
| Coefficients | Standard Error | |
|---|---|---|
| Intercept | 39942 | |
| X Variable 1 | -29.5 |
Figure-1
What is the market size? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 39942
(The following is a purely imaginary scenario)
A demand response curve is modelled using linear regression. The partial regression output is given in Figure-1 below. Given this information, answer the subquestions.
ANOVA
| df | SS | |
|---|---|---|
| Regression | 7076613 | |
| Residual | ||
| Total | 9 | 7184891 |
| Coefficients | Standard Error | |
|---|---|---|
| Intercept | 39942 | |
| X Variable 1 | -29.5 |
Figure-1
What is the satiating price? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 1353.5 (accepted within ±1.5)
(The following is a purely imaginary scenario)
A demand response curve is modelled using linear regression. The partial regression output is given in Figure-1 below. Given this information, answer the subquestions.
ANOVA
| df | SS | |
|---|---|---|
| Regression | 7076613 | |
| Residual | ||
| Total | 9 | 7184891 |
| Coefficients | Standard Error | |
|---|---|---|
| Intercept | 39942 | |
| X Variable 1 | -29.5 |
Figure-1
Based on the elasticity, which of the following statements are TRUE
The demand is elastic
The demand is inelastic
The price is elastic
The price is inelastic
Correct answer
The demand is inelastic
(The following is a purely imaginary scenario)
A demand response curve is modelled using linear regression. The partial regression output is given in Figure-1 below. Given this information, answer the subquestions.
ANOVA
| df | SS | |
|---|---|---|
| Regression | 7076613 | |
| Residual | ||
| Total | 9 | 7184891 |
| Coefficients | Standard Error | |
|---|---|---|
| Intercept | 39942 | |
| X Variable 1 | -29.5 |
Figure-1
What percentage of the total linear variability in demand is captured by this model (given in figure-1)? (Note: Enter the answer in “percentage rounded to two decimal places without the percentage sign. For example, if the answer is “1.234%”, then enter it as “1.23”)
Correct answer: 98.5 (accepted within ±0.5)
(The following is a purely imaginary scenario)
A survey was conducted among 100 students who participated in all events at IITMs recent “Saarang” event. The students were asked to rate various aspects of their experience on a scale of “-5 to +5”, where “-5” indicates a very negative experience and “+5” indicates a very positive experience. The average student rating for each event was taken across four key parameters: “Event Ambience”, “Fairness of Event Judges”, “Event Conduct”, “Event Prize Money”. The target variable was “Event Performance” which measures the overall performance of each event. Based on the collected data, a correlation matrix as specified in Table-2 is obtained for the independent variables. With this information, answer the given sub-questions
| Pairwise Correlation Values | ||||
|---|---|---|---|---|
| Event Ambience | Fairness of Event Judges | Event Conduct | Event Prize Money | |
| Event Ambience | 1 | -0.04 | 0.49 | 0.55 |
| Fairness of Event Judges | -0.04 | 1 | -0.011 | 0.022 |
| Event Conduct | 0.49 | -0.011 | 1 | -0.51 |
| Event Prize Money | 0.55 | 0.022 | -0.51 | 1 |
Table-2
What would be the R-Square value if a model is built with “Event Conduct” as the independent variable and “Event Prize Money” as the dependent variable? (Note: Enter the answer in percentage rounded to two decimal places without the “%” symbol. For example, if the answer is “1.234%”, then enter it as “1.23”)
Correct answer: 26 (accepted within ±1)
(The following is a purely imaginary scenario)
A survey was conducted among 100 students who participated in all events at IITMs recent “Saarang” event. The students were asked to rate various aspects of their experience on a scale of “-5 to +5”, where “-5” indicates a very negative experience and “+5” indicates a very positive experience. The average student rating for each event was taken across four key parameters: “Event Ambience”, “Fairness of Event Judges”, “Event Conduct”, “Event Prize Money”. The target variable was “Event Performance” which measures the overall performance of each event. Based on the collected data, a correlation matrix as specified in Table-2 is obtained for the independent variables. With this information, answer the given sub-questions
| Pairwise Correlation Values | ||||
|---|---|---|---|---|
| Event Ambience | Fairness of Event Judges | Event Conduct | Event Prize Money | |
| Event Ambience | 1 | -0.04 | 0.49 | 0.55 |
| Fairness of Event Judges | -0.04 | 1 | -0.011 | 0.022 |
| Event Conduct | 0.49 | -0.011 | 1 | -0.51 |
| Event Prize Money | 0.55 | 0.022 | -0.51 | 1 |
Table-2
What would be the Adjusted R-Square value if a model is built between “Fairness of Event Judges” as the independent variable and “Event Ambience” as the dependent variable? (Note: Enter the answer in percentage rounded to two decimal places without the “%” symbol. For example, if the answer is “1.234%”, then enter it as “1.23”)
Correct answer: 15.5 (accepted within ±0.5)
(The following is a purely imaginary scenario)
A survey was conducted among 100 students who participated in all events at IITMs recent “Saarang” event. The students were asked to rate various aspects of their experience on a scale of “-5 to +5”, where “-5” indicates a very negative experience and “+5” indicates a very positive experience. The average student rating for each event was taken across four key parameters: “Event Ambience”, “Fairness of Event Judges”, “Event Conduct”, “Event Prize Money”. The target variable was “Event Performance” which measures the overall performance of each event. Based on the collected data, a correlation matrix as specified in Table-2 is obtained for the independent variables. With this information, answer the given sub-questions
| Pairwise Correlation Values | ||||
|---|---|---|---|---|
| Event Ambience | Fairness of Event Judges | Event Conduct | Event Prize Money | |
| Event Ambience | 1 | -0.04 | 0.49 | 0.55 |
| Fairness of Event Judges | -0.04 | 1 | -0.011 | 0.022 |
| Event Conduct | 0.49 | -0.011 | 1 | -0.51 |
| Event Prize Money | 0.55 | 0.022 | -0.51 | 1 |
Table-2
What will be the percentage increase in the standard error corresponding to the beta value of “Event Ambience”, if a multiple linear regression model where both “Event Ambience” and “Event Price Money” are used as explanatory variables to predict “Event Performance”? (Note: Enter the answer in percentage rounded to two decimal places. For example, if the answer is “1.234%”, then enter it as “1.23”)
Correct answer: 19 (accepted within ±1)
There are six business units. There are two outputs and one input under consideration. You are solving the optimization problem for business unit 3, and you find that the efficiency is 0.8. You see that the dual variables corresponding to the constraints of business units 2 and 5 are non- zero, and the dual variables corresponding to the constraints of other units are zero. The dual variables corresponding to the constraints of business units 2 and 5 are 0.5 and 0.3, respectively. Based on Table 6, answers the given sub-questions
How much will the Output 1 in HCU 3?
Correct answer: 8062.5 (accepted within ±1.5)
There are six business units. There are two outputs and one input under consideration. You are solving the optimization problem for business unit 3, and you find that the efficiency is 0.8. You see that the dual variables corresponding to the constraints of business units 2 and 5 are non- zero, and the dual variables corresponding to the constraints of other units are zero. The dual variables corresponding to the constraints of business units 2 and 5 are 0.5 and 0.3, respectively. Based on Table 6, answers the given sub-questions
How much will the Output 2 in HCU 3?
Correct answer: 10.75 (accepted within ±0.25)
(The following is a purely imaginary scenario)
Ms. Teddy, is the owner of a toy manufacturing company. The company produces its iconic “Teddy Bear” toys in a facility that operates for 8 hours a day, and 20 days in month. Ms. Teddy has recently completed the BA course and wants to see if her manufacturing facility is producing toys where the defects per ship follows a Poisson distribution. Accordingly, she collected data for the past month indicating the number of defects produced in a shift. This is provided in Table-1. Using this information answer the given sub-questions
| Production Day in the Month | Number of Defective Teddy Bears Produced |
|---|---|
| Day-1 | 5 |
| Day-2 | 0 |
| Day-3 | 4 |
| Day-4 | 0 |
| Day-5 | 4 |
| Day-6 | 4 |
| Day-7 | 2 |
| Day-8 | 5 |
| Day-9 | 3 |
| Day-10 | 3 |
| Day-11 | 1 |
| Day-12 | 5 |
| Day-13 | 2 |
| Day-14 | 1 |
| Day-15 | 5 |
| Day-16 | 5 |
| Day-17 | 2 |
| Day-18 | 1 |
| Day-19 | 1 |
| Day-20 | 1 |
Table-1
How many bins will be present in the frequency table for the statistical test to be conducted? (Note: Enter an INTEGER answer*)*
Correct answer: 6
(The following is a purely imaginary scenario)
Ms. Teddy, is the owner of a toy manufacturing company. The company produces its iconic “Teddy Bear” toys in a facility that operates for 8 hours a day, and 20 days in month. Ms. Teddy has recently completed the BA course and wants to see if her manufacturing facility is producing toys where the defects per ship follows a Poisson distribution. Accordingly, she collected data for the past month indicating the number of defects produced in a shift. This is provided in Table-1. Using this information answer the given sub-questions
| Production Day in the Month | Number of Defective Teddy Bears Produced |
|---|---|
| Day-1 | 5 |
| Day-2 | 0 |
| Day-3 | 4 |
| Day-4 | 0 |
| Day-5 | 4 |
| Day-6 | 4 |
| Day-7 | 2 |
| Day-8 | 5 |
| Day-9 | 3 |
| Day-10 | 3 |
| Day-11 | 1 |
| Day-12 | 5 |
| Day-13 | 2 |
| Day-14 | 1 |
| Day-15 | 5 |
| Day-16 | 5 |
| Day-17 | 2 |
| Day-18 | 1 |
| Day-19 | 1 |
| Day-20 | 1 |
Table-1
What is the value of the computed test statistic for the statistical test that is to be performed by Ms. Teddy to very her claim? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 590.5 (accepted within ±1.5)
(The following is a purely imaginary scenario)
Ms. Teddy, is the owner of a toy manufacturing company. The company produces its iconic “Teddy Bear” toys in a facility that operates for 8 hours a day, and 20 days in month. Ms. Teddy has recently completed the BA course and wants to see if her manufacturing facility is producing toys where the defects per ship follows a Poisson distribution. Accordingly, she collected data for the past month indicating the number of defects produced in a shift. This is provided in Table-1. Using this information answer the given sub-questions
| Production Day in the Month | Number of Defective Teddy Bears Produced |
|---|---|
| Day-1 | 5 |
| Day-2 | 0 |
| Day-3 | 4 |
| Day-4 | 0 |
| Day-5 | 4 |
| Day-6 | 4 |
| Day-7 | 2 |
| Day-8 | 5 |
| Day-9 | 3 |
| Day-10 | 3 |
| Day-11 | 1 |
| Day-12 | 5 |
| Day-13 | 2 |
| Day-14 | 1 |
| Day-15 | 5 |
| Day-16 | 5 |
| Day-17 | 2 |
| Day-18 | 1 |
| Day-19 | 1 |
| Day-20 | 1 |
Table-1
How many degrees of freedom is present for the test statistic that is to be conducted by Ms. Teddy? (Note: Enter an INTEGER answer*)*
Correct answer: 4
(The following is a purely imaginary scenario)
The relationship between Demand “D” and Selling Price “P” is given by the equation D(p) = 780 – 9*P. Then answer the given sub-questions.
If the intention is to maximize the profit, then what is the optimal selling price if the item is going to be made at Rs. 30 per unit? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 58.5 (accepted within ±0.5)
(The following is a purely imaginary scenario)
The relationship between Demand “D” and Selling Price “P” is given by the equation D(p) = 780 – 9*P. Then answer the given sub-questions.
What is the maximum profit that can be generated? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 20424.5 (accepted within ±0.5)
As a Business Analyst, we are trying to find the DMUs that are efficient, where there are 2 outputs (Number of Leads & Sales) and one constant input. Based on the graph given in Figure-4 below, answer the given sub-questions
Which DMUs are in the economic frontier?
(3,2,1)
(3,2,5)
(4,5,2)
(3,4,1)
Correct answer
(3,4,1)
As a Business Analyst, we are trying to find the DMUs that are efficient, where there are 2 outputs (Number of Leads & Sales) and one constant input. Based on the graph given in Figure-4 below, answer the given sub-questions
Which DMUs are the reference units for DMU 2?
(5,4)
(4,1)
(3,5)
(3,1)
Correct answer
(4,1)
As a Business Analyst, we are trying to find the DMUs that are efficient, where there are 2 outputs (Number of Leads & Sales) and one constant input. Based on the graph given in Figure-4 below, answer the given sub-questions
Which DMUs are the reference units for DMU 5?
(1,4)
(1,3)
(3,4)
None of these
Correct answer
(3,4)
(The following is a purely imaginary scenario)
An insurance company believes that people can be divided into two classes: “Class-1: Those who are accident prone” and “Class-2: Those who are not accident prone”. The company’s statistics show that an accident-prone person will have an accident at sometime within a faxed 1-year period with probability 0.4, whereas this probability decreases to 0.2 for a person who is not accident prone. It is assumed that 30 percent of the human population is accident prone. Given this information, answer the sub-questions
What is the probability that a new policyholder will have an accident within a year of purchasing a policy? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 0.26 (accepted within ±0.01)
(The following is a purely imaginary scenario)
An insurance company believes that people can be divided into two classes: “Class-1: Those who are accident prone” and “Class-2: Those who are not accident prone”. The company’s statistics show that an accident-prone person will have an accident at sometime within a faxed 1-year period with probability 0.4, whereas this probability decreases to 0.2 for a person who is not accident prone. It is assumed that 30 percent of the human population is accident prone. Given this information, answer the sub-questions
Suppose a new policyholder has an accident within a year of purchasing a policy. Then, what is the probability that the person is accident prone? (Note: Enter the answer rounded to two decimal places. For example, if the answer is “1.234”, then enter it as “1.23”)
Correct answer: 0.46 (accepted within ±0.02)
A fintech company wants to assess the performance of its classification model, which predicts loan approval (positive class) or rejection (negative class). Table 3 presents the data used for model development, while Table 4 provides the estimated coefficients from the logistic regression model. To validate the model, the actual loan decisions made by the fintech company are shown in Table 5.
| Cust_id | Salary | Age |
|---|---|---|
| 1 | 60 | 30 |
| 2 | 9 | 35 |
| 3 | 8 | 40 |
| 4 | 16 | 25 |
| 5 | 15 | 50 |
| 6 | 50 | 21 |
Table 3: Data used for Building a Logistic Regression Model
| Threshold | 0.35 |
|---|---|
| b0 | 0.03 |
| b1 | 0.04 |
| b2 | -0.03 |
Table 4: Estimated Coefficients from Logistic Regression
| Y_act |
|---|
| 1 |
| 1 |
| 0 |
| 0 |
| 0 |
| 0 |
Table 5: The actual loan decisions
Based on the provided data, answer the given sub-questions
What is the precision of class 1?
(Note: Enter the answer in percentage rounded to two decimal places. For example, if the answer is “10.235%”, then enter it as “10.24”)
Correct answer: 33.33 (accepted within ±0.01)
A fintech company wants to assess the performance of its classification model, which predicts loan approval (positive class) or rejection (negative class). Table 3 presents the data used for model development, while Table 4 provides the estimated coefficients from the logistic regression model. To validate the model, the actual loan decisions made by the fintech company are shown in Table 5.
| Cust_id | Salary | Age |
|---|---|---|
| 1 | 60 | 30 |
| 2 | 9 | 35 |
| 3 | 8 | 40 |
| 4 | 16 | 25 |
| 5 | 15 | 50 |
| 6 | 50 | 21 |
Table 3: Data used for Building a Logistic Regression Model
| Threshold | 0.35 |
|---|---|
| b0 | 0.03 |
| b1 | 0.04 |
| b2 | -0.03 |
Table 4: Estimated Coefficients from Logistic Regression
| Y_act |
|---|
| 1 |
| 1 |
| 0 |
| 0 |
| 0 |
| 0 |
Table 5: The actual loan decisions
Based on the provided data, answer the given sub-questions
What is the sensitivity of the model?
(Note: Enter the answer in percentage rounded to two decimal places. For example, if the answer is “10.235%”, then enter it as “10.24”)
Correct answer: 50 (accepted within ±0.05)
A fintech company wants to assess the performance of its classification model, which predicts loan approval (positive class) or rejection (negative class). Table 3 presents the data used for model development, while Table 4 provides the estimated coefficients from the logistic regression model. To validate the model, the actual loan decisions made by the fintech company are shown in Table 5.
| Cust_id | Salary | Age |
|---|---|---|
| 1 | 60 | 30 |
| 2 | 9 | 35 |
| 3 | 8 | 40 |
| 4 | 16 | 25 |
| 5 | 15 | 50 |
| 6 | 50 | 21 |
Table 3: Data used for Building a Logistic Regression Model
| Threshold | 0.35 |
|---|---|
| b0 | 0.03 |
| b1 | 0.04 |
| b2 | -0.03 |
Table 4: Estimated Coefficients from Logistic Regression
| Y_act |
|---|
| 1 |
| 1 |
| 0 |
| 0 |
| 0 |
| 0 |
Table 5: The actual loan decisions
Based on the provided data, answer the given sub-questions
What is the specificity of the model?
(Note: Enter the answer in percentage rounded to two decimal places. For example, if the answer is “10.235%”, then enter it as “10.24”)
Correct answer: 50 (accepted within ±0.05)
A fintech company wants to assess the performance of its classification model, which predicts loan approval (positive class) or rejection (negative class). Table 3 presents the data used for model development, while Table 4 provides the estimated coefficients from the logistic regression model. To validate the model, the actual loan decisions made by the fintech company are shown in Table 5.
| Cust_id | Salary | Age |
|---|---|---|
| 1 | 60 | 30 |
| 2 | 9 | 35 |
| 3 | 8 | 40 |
| 4 | 16 | 25 |
| 5 | 15 | 50 |
| 6 | 50 | 21 |
Table 3: Data used for Building a Logistic Regression Model
| Threshold | 0.35 |
|---|---|
| b0 | 0.03 |
| b1 | 0.04 |
| b2 | -0.03 |
Table 4: Estimated Coefficients from Logistic Regression
| Y_act |
|---|
| 1 |
| 1 |
| 0 |
| 0 |
| 0 |
| 0 |
Table 5: The actual loan decisions
Based on the provided data, answer the given sub-questions
What is the correct interpretation of the coefficient b1? (Choose that is/are applicable)
If the salary increases by 1 unit, the log of odds of the application acceptance increases by 0.04.
If the salary increases by 1 unit, the odds of the application getting accepted increase by 4% (e^(0.4) = 1.04)
None of these
Correct answers
If the salary increases by 1 unit, the log of odds of the application acceptance increases by 0.04.
If the salary increases by 1 unit, the odds of the application getting accepted increase by 4% (e^(0.4) = 1.04)
A fintech company wants to assess the performance of its classification model, which predicts loan approval (positive class) or rejection (negative class). Table 3 presents the data used for model development, while Table 4 provides the estimated coefficients from the logistic regression model. To validate the model, the actual loan decisions made by the fintech company are shown in Table 5.
| Cust_id | Salary | Age |
|---|---|---|
| 1 | 60 | 30 |
| 2 | 9 | 35 |
| 3 | 8 | 40 |
| 4 | 16 | 25 |
| 5 | 15 | 50 |
| 6 | 50 | 21 |
Table 3: Data used for Building a Logistic Regression Model
| Threshold | 0.35 |
|---|---|
| b0 | 0.03 |
| b1 | 0.04 |
| b2 | -0.03 |
Table 4: Estimated Coefficients from Logistic Regression
| Y_act |
|---|
| 1 |
| 1 |
| 0 |
| 0 |
| 0 |
| 0 |
Table 5: The actual loan decisions
Based on the provided data, answer the given sub-questions
What is the correct interpretation of the coefficient b2? (Choose that is/are applicable)
If the age increases by 1 unit, the log of odds of the application acceptance decreases by 0.03.
If the age increases by 1 unit, the odds of the application getting accepted decrease by 3% (e^(-0.03) = 0.97).
None of these
Correct answers
If the age increases by 1 unit, the log of odds of the application acceptance decreases by 0.03.
If the age increases by 1 unit, the odds of the application getting accepted decrease by 3% (e^(-0.03) = 0.97).