uiz Space

September 2023 term · Linear Statistical Models · BSMA3012

Linear Statistical Models End Term: 24 December 2023 (September 2023 term)

The IIT Madras BS Linear Statistical Models (Linear Statistical Models) End Term paper sat on 24 Dec 2023, in the September 2023 term: 3 questions for 54 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.

Questions
3
Marks
54
Duration
180 min
MCQ
3

Updated

Official paper: IIT M DEGREE FN EXAM FDB1 24 Dec 2023 · No negative marking.

Question 1

+6 marksOne correct option
  1. Consider the Linear Model:

yi=β0+β1xi+ϵi;1≤i≤1000y_i = \beta_0 + \beta_1 x_i + \epsilon_i \quad ; \quad 1 \leq i \leq 1000

If ϵi∼N(0,σ2)\epsilon_i \sim N(0, \sigma^2), then define (do not solve for) the least square estimate of β=[β0β1]\beta = \begin{bmatrix} \beta_0 \\ \beta_1 \end{bmatrix}. [2 Marks]

  1. Let (y‾,Xβ‾,σ2In×n)(\underline{y}, X\underline{\beta}, \sigma^2 I_{n\times n}) be a linear model. Suppose p‾Tβ‾\underline{p}^T\underline{\beta} is estimable. Then derive an expression for the variance of BLUE of p‾Tβ‾\underline{p}^T\underline{\beta} [4 Marks]
  1. A

    I have written answers on the answer sheets

  2. B

    Not applicable

Show answer

Correct answer

  • A

    I have written answers on the answer sheets

Question 2

+22 marksOne correct option

3.Consider the data set scores on a class of 98 students in IIT-M. For each of 98 students, the composite score obtained in the class and the average number of hours studied per week is recorded. The following regression output was obtained using the scores data set

text
Call:
lm(formula = score ~ hours)
Residuals:
Min 1Q Median 3Q Max
-39.680 -14.675 0.215 14.088 54.785
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 7.8805 4.6247 1.704 0.0916 .
hours 7.1852 0.6134 11.713 <2e-16 ***
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
Residual standard error: 20.08 on 96 degrees of freedom
Multiple R-squared: 0.5883, Adjusted R-squared: 0.584
F-statistic: 137.2 on 1 and 96 DF, p-value: < 2.2e-16

For the following questions, please explain clearly which parts of the output are the basis for your answers. Include the units of variables wherever possible. Use the tabulated values in the cheat sheet.

(a) What is the predictor variable? What is the response variable? [2 Marks]

(b) Write the equation for the estimated conditional mean function, using the appropriate numerical values in the output. [2 Marks]

(c) Based on the estimated coefficients, can you give an estimate of E[Y∣X=0]E[Y|X = 0]? If yes, what is it (and show your work); if not, explain why not (missing information, inappropriate assumptions, etc.). [3 Marks]

(d) Give a 95% confidence interval for β1\beta_1, assuming all the model assumptions hold. [2 Marks]

(e) What is σ^2\hat{\sigma}^2, the in-sample mean-squared error? [3 Marks]

(f) From nn, the standard error of β^1\hat{\beta}_1 and σ^2\hat{\sigma}^2, can you find the sample variance in population across cities? If so, what is it? If not, explain. [5 Marks]

(g) Which part (or parts) of the output (if any) tests the assumption that the relationship between the predictor variable and the response variable is linear? [5 Marks]

  1. A

    I have written answers on the answer sheets

  2. B

    Not applicable

Show answer

Correct answer

  • A

    I have written answers on the answer sheets

Question 3

+26 marksOne correct option
  1. Four factories producing soft-drinks are randomly chosen to check if the amount of caffeine in 1L bottles varies throughout the factories. For this the amount of caffeine (in a hundred mg) in three 1L bottles of soft-drinks of each of the selected factories is recorded in the following table:
BrandObs. 1Obs. 2Obs. 3
AA544
BB446
CC555
DD764

Based on the given information, answer the following questions:

(a) Suppose we assume the linear model:

yij=μ+αi+ϵij    ;i=1,2,3,4   and   j=1,2,3y_{ij} = \mu + \alpha_i + \epsilon_{ij} \;\; ; i = 1, 2, 3, 4 \;\text{ and }\; j = 1, 2, 3

where, αi∼N(0,σμ2)\alpha_i \sim N(0, \sigma^2_\mu) and ϵ∼N(0,σ2)\epsilon \sim N(0, \sigma^2)

Identify whether the given model is the fixed effect model or the random effect model. Give reason. [2 Marks]

(b) Compute the treatment means for each of the given treatment, i.e. Y‾i.\overline{Y}_{i.} and the overall mean of the given observations, i.e.Y‾..\overline{Y}_{..}. [2 Marks]

(c) Find the sum of squares due to treatment, i.e., SStreatmentSS_{treatment} and sum of squares due to error, i.e. SSerrorSS_{error}. [4 Marks]

(d) Compute mean sum of squares due to treatment and error, i.e. MStreatmentMS_{treatment} and MSerrorMS_{error}. [2 Marks]

(e) Find the estimate value of σμ2\sigma^2_\mu. [2 Marks]

(f) Draw the ANOVA table for the above given model. [3 Marks]

(g) An analyst wish to test if there is a variability in amount of caffeine across all the 4 factories. Write down the null and alternative hypothesis to be tested for it. [2 Marks]

(h) Perform hypothesis testing for the above defined hypothesis and conclude your results. [4 Marks]

(i) The analyst also wishes to test if the grand mean caffeine content in 1L soft-drinks bottles is 5 (in a hundred mg) or not. Perform hypothesis test for the same. [5 Marks]

  1. A

    I have written answers on the answer sheets

  2. B

    Not applicable

Show answer

Correct answer

  • A

    I have written answers on the answer sheets