ECON 371 Class Recording 9/22
Watch on YouTube →
Overview
Professor Lantis completes the simple linear regression unit by interpreting dummy-variable coefficients, generating predictions in Stata, standardizing variables, solving for implied x-values, and evaluating explanatory power with R-squared. The class also introduces heteroskedasticity: the Breusch–Pagan-style `estat imtest` test treats homoskedastic errors as the null, while Stata's `robust` or `vce(robust)` options adjust standard errors when the null is rejected.
Key takeaways
- Dummy-variable slopes compare the one-coded group with the zero-coded reference group, while a binary dependent variable turns the slope into a change in predicted probability.
- Stata's `_b[variable]` and `_b[_cons]` retrieve coefficients from the most recent regression, enabling precise prediction and algebraic calculations without manual rounding.
- Standardizing income changes the interpretation from dollars to standard deviations, and standardizing both variables produces unitless slope estimates unaffected by whether income was initially measured in dollars or thousands.
- R-squared equals explained sum of squares divided by total sum of squares; a statistically significant coefficient can still have weak predictive value, as shown by COVID vaccination explaining only about 2.4% of case-count variation.
- Heteroskedasticity affects inference rather than necessarily changing coefficient estimates: it can distort standard errors and therefore alter t-statistics, p-values, and null-hypothesis decisions.
- In Stata, `estat imtest` tests the homoskedasticity null, and `robust` or `vce(robust)` adjusts standard errors when changing error variance makes conventional inference unreliable.
Chapters
- The simple linear regression quiz becomes available after class and is due by the end of Friday; Thursday's multiple regression material is excluded.
- The second problem set uses a data file and document linked in the assignment and focuses on simple linear regression.
- Professor Lantis warns that coding questions receive little or no partial credit and that the quiz uses a lockdown browser.
- A dummy-variable coefficient compares the one group with the zero reference group rather than representing a continuous unit change.
- In the GPA example, student athletes are predicted to have GPAs approximately 0.28 points lower than non-student athletes.
- For a gender indicator, the predicted value is the intercept for the zero group and the intercept plus the dummy coefficient for the one group.
- When metro-county status is the dependent variable, a slope coefficient represents a change in the predicted probability of belonging to the one group.
- A 1-percentage-point increase in unemployment is associated with roughly a 0.045 decrease in the probability that a county is metropolitan, equivalent to a 4.5-percentage-point reduction.
- The intercept is the predicted probability for a county with 0% unemployment, although that value may not be economically realistic.
- A very small p-value does not imply that an effect is substantively important; a statistically significant probability change of 0.000000001 may have negligible economic meaning.
- Standardizing variables by subtracting the mean and dividing by the standard deviation converts measurements into unitless standard-deviation changes.
- Professor Lantis connects standardized regression effects to the broader idea of elasticity, while noting that standardization alone is not elasticity.
- The class reuses the Colorado county COVID dataset and the Week 4 Stata do-file.
- An initial regression predicts county vaccination percentage from median household income measured in dollars.
- The income coefficient is small because it describes the effect of a $1 increase; multiplying the coefficient by 1,000 or 3,000 rescales the effect to larger income changes.
- Stata's `display` command performs calculations such as `2 + 2` and can retrieve coefficients from the most recently estimated regression.
- The notation `_b[variable]` retrieves a slope coefficient, while `_b[_cons]` retrieves the intercept.
- Stored coefficients only correspond to the most recent regression, so `_b[median_income]` cannot be used after running a model that omits median income.
- A predicted value follows the regression equation y-hat = intercept + slope multiplied by the hypothetical x-value.
- For median household income of $50,000, Stata evaluates the intercept plus the income coefficient times 50,000.
- The example produces a predicted county vaccination rate of approximately 38.8%.
- Stata's `std()` function creates a variable measured in standard deviations from the mean rather than dollars.
- With standardized income as x, the slope is interpreted as the change in vaccination percentage associated with a one-standard-deviation increase in income.
- The example estimates an increase of approximately 4.828 percentage points in predicted vaccination for a one-standard-deviation income increase.
- When both vaccination percentage and income are standardized, the slope measures the change in vaccination's standard deviations associated with a one-standard-deviation income increase.
- Changing income from dollars to thousands of dollars before standardizing does not change the standardized regression coefficient because standardization removes the original units.
- The intercept in a regression with standardized x equals the predicted y-value when x is zero, which corresponds to the mean of the original income variable.
- To find the x-value associated with a target y-hat, subtract the intercept from the target and divide by the slope.
- Stata can calculate `(50 - _b[_cons]) / _b[median_income]` without manually rounding the regression coefficients.
- The example estimates that median household income of approximately $83,000 corresponds to a predicted vaccination rate of 50%.
- R-squared measures the proportion of total variation in y explained by the regression's predicted values.
- If a model explains 200 units of variation out of 1,000 total units, R-squared is 0.20, or 20%.
- Stata reports explained or model sum of squares and total sum of squares, with R-squared calculated as ESS divided by TSS.
- The vaccination-rate model explains about 2.4% of variation in county COVID cases, while the household-income model explains about 4.5%.
- A larger absolute test statistic indicates a smaller p-value even when Stata rounds multiple p-values to 0.000.
- County population explains roughly 98% of variation in total COVID cases because larger populations mechanically allow more cases, illustrating a strong but potentially uninformative relationship.
- Homoskedasticity means the variance of prediction errors is constant across x-values; heteroskedasticity means the error variance changes with x.
- A widening spread of test-score errors at higher student-teacher ratios is an example of heteroskedasticity.
- Heteroskedasticity does not necessarily change coefficient estimates, but it can distort standard errors, test statistics, p-values, and rejection decisions.
- Stata's `estat imtest` tests homoskedasticity as the null hypothesis after a regression; a p-value of 0.07 rejects the null at alpha = 0.10 but not at 0.05 or 0.01.
- When heteroskedasticity is detected, adding `robust` or `vce(robust)` to the regression produces heteroskedasticity-consistent standard errors.
- The robust option leaves coefficient estimates unchanged but can alter standard errors, test statistics, and p-values; Professor Lantis notes that robust standard errors often increase but can occasionally decrease.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.