ECON 371 Class Recording 10/1
Watch on YouTube →
Overview
Professor Lantis explains how categorical variables, perfect and imperfect multicollinearity, and omitted-variable bias affect multiple regression estimates and their interpretation. Using Stata examples on NCAA athletic revenue, traffic stops, and wages, the class practices choosing a reference category, interpreting dummy-variable coefficients and p-values, and addressing correlated regressors.
Key takeaways
- A regression with every dummy for a complete category set has perfect multicollinearity; omit one category, and interpret every remaining coefficient relative to that reference group.
- In the NCAA revenue example, conference indicators alone explain about 69% of observed revenue variation, but omitted factors such as facilities, players, and fan bases prevent interpreting that fit as a causal conference effect.
- When a binary outcome such as ticket issuance is regressed on indicators, coefficients are changes in predicted probability; multiple-regression interpretations should also state that other included factors are held constant.
- Imperfect multicollinearity primarily inflates standard errors: in the wage example, adding age raised the experience standard error from about 2.94 to 3.55 and removed its statistical significance.
- Omitted-variable bias depends on both the omitted variable's effect on the outcome and its covariance with the included regressor; a negative effect and positive covariance imply downward bias.
- Transformations such as GDP per capita or experience divided by age can represent related information in one regressor when entering correlated variables separately makes their individual effects difficult to estimate.
Chapters
0:00
Quiz 3, Problem Set 3, and Midterm Schedule
- Professor Lantis says Quiz 3 will cover material through this class and be due by the end of Tuesday, October 6.
- Problem Set 3 is planned for release the next day and due by the end of the following Friday.
- The midterm is scheduled for the following Thursday; exam practice problems will be posted, with office hours available Monday.
3:07
Perfect Multicollinearity in Categorical Dummy Variables
- Perfect multicollinearity occurs when one regressor is an exact linear function of another, such as duplicate variables or two identifiers that always match.
- Creating a dummy for every education category makes the categories mutually exclusive and collectively exhaustive, so one dummy is perfectly determined by the others.
- Stata detects perfect multicollinearity and drops a variable; the omitted category becomes the reference group for interpreting the remaining coefficients.
9:47
Choosing an Education Reference Group for Wage Comparisons
- A high-school dummy coefficient measures the predicted wage difference between high-school-only graduates and the omitted education group.
- Leaving out a bachelor's degree could make the high-school coefficient negative and the master's/PhD coefficient positive, reflecting their comparisons with bachelor's graduates.
- For ordinal categories such as education, choosing the lowest or highest category as the base can make results easier to interpret and present.
14:38
Building NCAA Conference Dummies in Stata
- The NCAA dataset records athletic-program revenue and conference names as text, which cannot be used directly as numeric regression predictors.
- Stata's `keep` command can retain observations meeting a condition, such as having a nonmissing conference; `drop` removes mistakenly created variables.
- Professor Lantis demonstrates a Stata `foreach` loop to create a dummy variable for each conference, using quotation marks for text comparisons.
22:27
Stata `encode` and `i.` Factor Variables for Categories
- Stata's `encode` converts text conference names into numeric codes with associated value labels, allowing categorical data to be handled by regression tools.
- The `i.` factor-variable notation automatically creates category indicators, avoiding manual dummy creation.
- Stata chooses the omitted category automatically when using `i.`; checking the data is necessary to identify the reference group.
- At the 1% significance level, the conference coefficients are compared with the omitted group using the rule p-value < 0.01; the Pac-12 comparison is not significant.
27:00
Traffic-Stop Ticket Model: Driver-Race Indicators
- The traffic-stop data uses coded values for stop outcomes and driver race, so a codebook is needed to interpret what each number represents.
- Professor Lantis creates a ticket indicator and race dummies, then regresses ticket issuance on driver race.
- With a binary ticket outcome, a dummy-variable coefficient describes a change in predicted probability; for example, the white-driver estimate is about 0.035.
- None of the driver-race coefficients is statistically significant at the 10% level, and the corresponding 90% confidence intervals include zero.
31:24
Adding Officer Characteristics and Stop-Reason Controls
- The expanded ticket model includes driver- and officer-race indicators, officer gender, years of service, and stop-reason categories.
- One category must be omitted from each complete set of race dummies to avoid perfect multicollinearity; `i.` can create stop-reason indicators even when the numeric codes are not decoded.
- Holding the other regressors constant, one additional year of officer service is associated with a 0.0008 decrease in the predicted ticket probability, or 0.08 percentage points.
- The model's R-squared is about 0.1, signaling that much of the outcome variation remains unexplained and that omitted-variable concerns remain.
40:55
Interpreting Race Estimates and Limits of Traffic-Stop Results
- In the expanded ticket model, the white-driver indicator is the omitted group; the black-driver coefficient is negative and statistically significant relative to white drivers.
- The Hispanic-driver result is near, but outside, the 1% significance threshold; Professor Lantis cautions against broad conclusions because relevant factors may be omitted.
- When the outcome changes from receiving a ticket to being arrested, R-squared falls by roughly half and many race indicators become insignificant.
- The class briefly notes that multiplying two dummy variables can construct indicators for joint categories, such as combinations of race and gender.
47:00
How Imperfect Multicollinearity Inflates Standard Errors
- Imperfect multicollinearity occurs when regressors are strongly related but are not exact duplicates or perfect linear functions of one another.
- The main consequence emphasized is inflated standard errors for affected coefficients, making the variables' separate effects harder to distinguish.
- Because t-statistics divide coefficient estimates by standard errors, larger standard errors lower the absolute t-statistic, raise the p-value, and reduce the chance of rejecting a zero-effect null.
- Wider confidence intervals can then include zero even when an estimate has a meaningful magnitude.
53:00
When to Keep, Drop, or Address Correlated Regressors
- If age and experience are included only as controls for a coefficient of primary interest, Professor Lantis says their individual significance may matter less than retaining the controls.
- If the research question concerns the separate effect of age or experience, inflated standard errors create a more direct inference problem.
- Increasing the sample size can reduce standard errors and improve precision, but it does not eliminate the underlying correlation between regressors.
- Perfectly redundant categorical indicators require dropping one category rather than relying on a larger sample.
56:00
Transforming Correlated Variables: GDP per Capita and Experience-to-Age
- GDP and population tend to move together, so GDP per capita can combine them into one interpretable measure rather than entering both as separate regressors.
- For the wage example, Professor Lantis constructs experience divided by age as a transformed variable.
- The transformed model retains information related to age and experience while avoiding their separate, strongly correlated entries.
- This approach is presented as useful when age and experience are controls and the main interest lies in another coefficient, such as IQ.
1:03:13
Omitted-Variable Bias in the Officer-Gender Estimate
- A regression of ticket issuance on an officer-male dummy alone gives an estimated probability difference of about -0.034, or -3.4 percentage points.
- Years of service may be omitted: Professor Lantis reasons that it lowers ticket probability and may be positively correlated with being male in the older traffic-stop data.
- The bias sign is the omitted variable's effect multiplied by its covariance with the included regressor, divided by the included regressor's variance.
- A negative service effect times a positive male–service covariance yields downward bias in the estimated male coefficient; the class distinguishes population β from sample estimate β-hat.
1:08:58
Attendance Quiz and Next-Class Preparation
- Professor Lantis creates a brief in-class quiz with the numeric response 22 and publishes it through the Canvas Quizzes tab.
- Students are encouraged to begin Problem Set 3 early and review the practice problems for Quiz 3.
- The planned next class will continue with examples and material leading into midterm review.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.