E371 Class Recording
Watch on YouTube →
Overview
Professor Lantis develops simple linear regression from least-squares fitting through coefficient interpretation, hypothesis tests, confidence intervals, and Stata output. Examples using marijuana use and income, county COVID data, and student-athlete GPAs show how to interpret continuous and dummy variables, while repeated-sample simulations illustrate how sample size affects precision and how confidence intervals capture population coefficients.
Key takeaways
- Ordinary least squares estimates the regression line by minimizing the sum of squared residuals, and the simple-regression slope equals the sample covariance of x and y divided by the sample variance of x.
- A regression slope always describes the predicted outcome change for a one-unit predictor change; rescaling income from dollars to thousands changes the coefficient’s units and makes its interpretation more useful.
- For a zero-slope null hypothesis, the regression test statistic is the estimated slope divided by its standard error; a confidence interval excluding zero supports rejection at the corresponding confidence level.
- Larger samples make coefficient estimates more precise: repeated samples of 100 counties produce a tighter slope distribution than samples of 30, while smaller samples increase standard errors.
- With a zero-one predictor, the slope is the difference between the predicted averages of the one and zero groups; with a zero-one outcome, predicted changes are probability changes.
- A statistically significant coefficient does not by itself establish a large practical effect: the COVID-income example motivates expressing household income in $1,000 units rather than interpreting a per-dollar change.
Chapters
- The fitted line predicts each observation’s outcome as ŷ = β̂₀ + β̂₁x; the residual is the difference between actual y and predicted ŷ.
- Ordinary least squares chooses the intercept and slope to minimize the sum of squared residuals, penalizing large prediction errors more heavily.
- The estimated slope is the sample statistic used to approximate the unknown population slope; its formula is the covariance of x and y divided by the variance of x.
- The slope β̂₁ predicts the change in y for a one-unit increase in x; the intercept β̂₀ predicts y when x equals zero.
- In the example relating monthly marijuana-use days to income, the intercept predicts income for someone reporting zero use days.
- The estimated slope indicates about $253 lower predicted income for each additional reported marijuana-use day per month; the fitted line slopes downward.
- In the county example, the outcome is COVID cases per 1,000 people and the predictor is the county vaccination percentage.
- The intercept predicts about 50.9 cases per 1,000 people when the vaccination percentage is zero.
- The slope is roughly −0.25 cases per 1,000 for a one-percentage-point increase; a 10-point increase corresponds to about 2.5 fewer cases per 1,000.
- A regression coefficient estimated from one sample can differ from the population coefficient, just as a sample mean can differ from the population mean.
- Across repeated samples, coefficient estimates are treated as a sampling distribution centered around the true population slope.
- A coefficient test statistic uses the estimate minus the hypothesized value, divided by its standard error; Stata calculates the more complex regression standard error.
- For the usual test of whether x and y are related, the null hypothesis sets the slope to zero; the test statistic is therefore the estimated slope divided by its standard error.
- The county vaccination regression reports a test statistic near −11 and a p-value rounded to zero, providing strong evidence against a zero slope.
- Professor Lantis reviews two-sided standard-normal critical values of about 1.65, 1.96, and 2.58 for 90%, 95%, and 99% confidence levels.
- A slope confidence interval is formed by adding and subtracting a critical value times the coefficient’s standard error.
- The income example’s approximate 95% interval runs from −$350 to −$150 per additional marijuana-use day; because it excludes zero, the zero-slope null is rejected.
- If zero falls inside a confidence interval, the corresponding zero-slope null cannot be rejected at that confidence level.
- The hypothesized slope need not be zero: the marijuana example’s interval also excludes a slope of −$100, so that value can be tested and rejected.
- For an oil-well births example, a standard error much larger than the estimated slope produces a test statistic with absolute value below one.
- The reported p-value is about 0.96, and the confidence interval includes zero, so the data do not support rejecting a zero-slope null.
- A confidence interval’s midpoint recovers the coefficient estimate: the example interval’s midpoint is close to the reported estimate of −3.9.
- Professor Lantis notes that quiz questions may omit parts of Stata output and require students to reconstruct results from the remaining information.
- A regression of county COVID cases on population has an intercept near −16, which predicts cases for a county with zero residents—a value outside the observed data range.
- The population coefficient is highly significant, with a test statistic around 250, consistent with larger counties having more cases.
- When household income is measured in dollars, its slope can look numerically small; dividing income by 1,000 expresses the predictor in thousands of dollars.
- The scaled income regression gives a slope around −0.4 per $1,000, an interpretation more useful than the effect per dollar.
- A Stata program repeatedly samples 30 counties from the full county dataset, runs the vaccination regression, and stores 1,000 slope estimates.
- The mean of those 1,000 estimates is close to the population-data slope of roughly −0.25, and their distribution is approximately normal.
- Repeating the exercise with samples of 100 produces a tighter distribution of estimates than samples of 30.
- Professor Lantis uses simulated intervals to illustrate coverage: a nominal 90% interval, built with a critical value of 1.65, contains the population slope in roughly 90% of repeated samples.
- A dummy or indicator predictor takes only the values zero and one; its regression slope is the difference in predicted group averages.
- For the GPA example, zero identifies non-athletes and one identifies student-athletes, so the intercept is the predicted GPA for non-athletes.
- The athlete coefficient indicates a predicted GPA about 28 points lower for student-athletes than for non-athletes, with a p-value near zero.
- The same hypothesis-testing and confidence-interval procedures apply to dummy predictors as to continuous predictors.
- When the dependent variable is a zero-one indicator, changes in its predicted value represent changes in probability.
- In the county example, the outcome indicates whether a county is metropolitan and the predictor is the unemployment rate.
- A slope near −0.045 means a one-percentage-point increase in unemployment is associated with about a 0.045 decrease in predicted metro probability, or 4.5 percentage points.
- For a regression of changes in county COVID cases on a South-region indicator, the p-value is about 0.01: the null is rejected at 10% and 5% levels, but narrowly not at 1%.
- A result significant at 99% confidence would also be significant at lower confidence levels; significance cannot occur only at the highest level.
- Reducing sample size increases the standard error, which lowers the test statistic’s magnitude and generally raises the p-value.
- Professor Lantis postpones the comparison of R-squared and adjusted R-squared until the next class.
- The in-class quiz access code is 1020; Professor Lantis says incomplete questions will be revisited after the next class covers the remaining material.
- The next problem set is posted, and students are advised to begin its practice problems early because the quiz window runs from Tuesday through the following Friday.
- Additional virtual office hours are announced for the next day from 9 to 11, with adjusted Monday office-hour times to be posted.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.