ECON E371 Class Recording 10/06
Watch on YouTube →
Overview
Professor Lantis reviews multiple linear regression concepts for an upcoming exam, showing why rescaling a predictor changes coefficient units but not its t-statistic or R², and how adjusted R² penalizes adding predictors. The lecture develops the joint F-test for added variables, interprets its formula and the effects of sample size and model dimensions, works through GPA and COVID-regression examples in Stata, and revisits imperfect multicollinearity and exam preparation.
Key takeaways
- Rescaling a predictor changes the coefficient and standard error into the new units, but leaves their ratio—the t-statistic—and the coefficient’s statistical significance unchanged.
- Ordinary R² cannot decrease when predictors are added; adjusted R² is more useful for comparing models because it penalizes additional predictors.
- A joint F-test evaluates whether a group of added regression coefficients is zero simultaneously, avoiding the use of separate t-tests as a substitute for a group-level test.
- For the joint F-statistic, Q is the number of added restrictions, K is the number of predictors in the unrestricted model, and N is the sample size; residual degrees of freedom are N − K − 1.
- In the GPA example, adding math/science credits and employment status raises R² from 0.1956 to 0.1997; the joint p-value of 0.02 supports rejection at 5% but not at 1%.
- Imperfect multicollinearity, such as the relationship between cigarette taxes and pack prices, can inflate standard errors and weaken individual significance tests even when coefficient estimates are not biased.
Chapters
- Problem set 3 is due Friday; the current quiz closes at 11:59 p.m. on the day of class.
- Quiz answers become available after the deadline, and Professor Lantis recommends reviewing them before the exam next Thursday.
- Thursday and the following Tuesday are planned for regression review and student questions, including practice problems and past quizzes.
- For home sale price regressed on income in dollars, the coefficient represents the predicted price change per $1 of income.
- Expressing income in $10,000 units changes the coefficient’s interpretation and numerical size to a predicted change per $10,000.
- The standard error scales with the coefficient, so their ratio—the t-statistic—does not change; statistical significance therefore remains the same.
- Changing income’s units does not change R² because R² measures the proportion of variation in home prices explained, not dollar amounts.
- Ordinary R² cannot fall when predictors are added, even if they contribute little explanatory value.
- Adjusted R² penalizes model complexity; Professor Lantis illustrates the idea with R² rising from 0.50 to 0.80 as the predictor count rises from two to four.
- A joint hypothesis test asks whether all coefficients on a group of added variables are zero simultaneously.
- The restricted model omits the added predictors; the unrestricted model includes them.
- Separate coefficient tests do not provide the same test of joint significance or control the overall error rate for the group.
- Adding predictors cannot reduce ordinary R², so the unrestricted-minus-restricted R² difference is nonnegative.
- The joint test uses an F-statistic and a right-tail p-value: large F values indicate evidence that the added variables improve the model.
- Under the null, the added variables jointly contribute no explanatory power; the test evaluates that group-level claim.
- The F-statistic compares the fit improvement from the restricted to unrestricted model with residual variation in the unrestricted model.
- In the R² form, the improvement is divided by Q, the number of added variables, and scaled by residual variation over N − K − 1.
- The equivalent residual-sum-of-squares form subtracts unrestricted-model residuals from restricted-model residuals because the restricted model leaves at least as much unexplained variation.
- The F-statistic uses N − K − 1 in its residual degrees of freedom, while the F distribution also depends on Q.
- With more observations, standard errors generally shrink, test statistics rise, and p-values fall when the estimated effect is nonzero.
- Smaller standard errors also narrow confidence intervals, making rejection of a false null more likely.
- Q counts the predictors added to the unrestricted model; holding fit improvement fixed, adding more predictors spreads that improvement across more restrictions and lowers the F-statistic.
- K is the number of predictors in the unrestricted model, and the residual degrees of freedom are N − K − 1.
- A larger starting model generally leaves less unexplained variation for a small group of added predictors to explain, weakening evidence for their joint contribution.
- The F distribution has two degrees of freedom: Q for the restrictions and N − K − 1 for residual variation.
- Rejecting the joint null means the added predictors explain a statistically significant additional amount of variation in the outcome.
- Professor Lantis says the exam will provide a formula sheet and Stata output; students should focus especially on interpreting Q, K, and N.
- The restricted GPA model includes age, high-school GPA, and a gender indicator; its R² is 0.1956.
- The unrestricted model adds math/science credits and whether the student works while attending college, raising R² to 0.1997.
- The joint test gives approximately F = 3.57 and p = 0.02: reject at the 10% and 5% levels, but not at the 1% level.
- The Stata `robust` option adjusts standard errors for heteroskedasticity; the `test` command evaluates the two added variables jointly.
- The example predicts COVID cases using unemployment rate, metro status, vaccination, and other predictors.
- The restricted model has R² = 0.8973; the unrestricted model has R² = 0.8994 after adding three variables.
- For the hand calculation, use Q = 3, N = 100, and K = 6 predictors in the unrestricted model.
- Small differences in rounded R² values can materially change the calculated F-statistic, so the exam will specify rounding precision.
- Increasing sample size improves precision and can make nonzero effects more statistically detectable, but it does not systematically change the expected coefficient estimates.
- The adjusted-R² penalty becomes less severe as N grows: the lecture contrasts a factor such as 2/1 in a tiny sample with 99/98 at N = 100.
- The class quiz question contained multiple incorrect answer choices; Professor Lantis notes that coefficient estimates should not change systematically just because sample size increases.
- Perfect multicollinearity occurs when predictors determine one another exactly; including freshman, sophomore, junior, and senior indicators together is an example.
- Imperfect multicollinearity occurs when predictors are strongly related but not exact—for example, age and years since high school, or cigarette taxes and price per pack.
- Multicollinearity does not itself bias OLS coefficient estimates under the usual assumptions, but it raises standard errors and makes individual effects harder to distinguish.
- Larger standard errors reduce t-statistics and make rejection of the affected coefficient’s null less likely; at least one correlated predictor’s standard error may be affected.
- The exam is next Thursday and will focus primarily on linear regression, with a smaller number of questions reviewing earlier material.
- Most questions will be short-answer rather than multiple-choice, often requiring interpretation of Stata output and follow-up calculations.
- Professor Lantis advises working through the practice problems, quizzes, and problem-set interpretations rather than relying on generated code without checking it.
- Office hours are available during the current and following week for questions about practice problems and exam preparation.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.