ECON E371 Class Recording 9/10
Watch on YouTube →
Overview
Professor Lantis closes the statistics review with Stata demonstrations of grouped means, indicator variables, Welch two-sample t-tests, and interpreting p-values, then introduces simple linear regression as a way to estimate relationships and test whether slopes differ from zero. Using election-county data and shale-well birth data, the class connects covariance and variance to the OLS slope, explains squared prediction errors and confidence intervals, and reviews upcoming assignments, attendance, and the canceled Tuesday class.
Key takeaways
- For two independent groups, a difference-in-means test evaluates the observed mean gap against a hypothesized gap of zero, using the standard error of the difference; Stata can calculate the Welch test when unequal variances are assumed.
- A two-sided test can be evaluated either by comparing the absolute test statistic with critical values or by comparing its p-value with α; a statistic of 1.21 and p-value of 0.22 fails to reject at the 90%, 95%, and 99% levels.
- In simple OLS regression, the estimated slope is covariance(X,Y) divided by variance(X), and the intercept makes the fitted line pass through the sample means of X and Y.
- OLS minimizes the sum of squared residuals, which prevents positive and negative errors from canceling and gives unusually large prediction errors extra weight.
- For the shale-well and county-birth example, the reported slope test statistic of about 0.03 and p-value of 0.97 provide no statistical evidence that well counts predict births in the analyzed sample.
- A regression confidence interval wider at 99% than at 95% reflects the greater range needed for higher confidence; if the interval excludes zero, the zero-slope null is rejected at the corresponding level.
Chapters
- The first problem set and weekly quiz are due by the end of the following day; the quiz covers material through confidence intervals.
- Professor Lantis offers virtual office hours from 9 a.m. to noon for last-minute questions.
- Tuesday class is canceled for a medical commitment; students should expect a Canvas announcement and likely a replacement recording to watch before Thursday.
- The course plans six problem sets and six weekly quizzes, generally covering about two weeks of material; quizzes test concepts and Stata-output interpretation, not code production.
- To compare two groups, the null hypothesis is typically that their population means are equal, so the hypothesized difference is zero.
- The test statistic uses the difference between sample means minus the hypothesized difference, divided by the standard error of that difference.
- For independent groups, the variance of the difference combines each group’s sample variance divided by its sample size.
- Professor Lantis connects the test to an example comparing starting salaries for economics and anthropology majors.
- Sample covariance measures whether observations tend to lie on the same or opposite sides of the means of X and Y; sample calculations divide by n−1.
- Correlation standardizes covariance by both variables’ standard deviations, removing units such as inch-pounds and producing a unitless relationship measure.
- An Excel spreadsheet must be brought into Stata with an import command rather than opened like a native .dta dataset; the first spreadsheet row can be specified as variable names.
- The global folder path must match the local computer, and saving over an existing Stata dataset requires the replace option.
- Stata’s egen command calculates the mean unemployment rate, which can serve as a threshold for classifying counties as high- or low-unemployment.
- A generated indicator variable records one when a county meets the chosen condition and zero otherwise.
- The bysort prefix lets Stata summarize or calculate a mean separately for each value of a grouping indicator.
- A group-mean variable assigns the Trump-vote proportion for the relevant high- or low-unemployment county group.
- Indicator thresholds can use a selected value, such as unemployment above 10%, rather than only the sample mean.
- Stata’s vertical bar represents OR, while an ampersand represents AND when combining logical conditions.
- Professor Lantis creates county indicators for high income and a high white population share, then combines them to identify counties meeting either condition.
- The example compares Trump-vote proportions across those combined county categories before asking whether the difference is statistically significant.
- Stata’s ttest command calculates the two group means, their difference, its standard error, a test statistic, and p-values for alternative hypotheses.
- The class uses unequal variances because two real-world groups should not generally be assumed to have identical variances.
- For a two-sided test, the large-sample critical values are approximately 1.65, 1.96, and 2.58 at the 90%, 95%, and 99% confidence levels.
- A statistic with absolute value 7.9 and a p-value rounded to zero strongly rejects equal Trump-vote proportions between high- and low-unemployment counties.
- A statistically significant difference between groups does not by itself establish that the grouping characteristic caused the outcome.
- Commute time, for example, may be correlated with urban residence or other characteristics that also relate to voting.
- The semester’s broader goal is to move from measuring correlations toward identifying causal relationships.
- Simple linear regression models an outcome Y as a function of one explanatory variable X, with intercept β₀ and slope β₁.
- In the exam example, exam score is the dependent variable and study hours are the independent variable.
- A positive slope indicates an upward relationship, a negative slope a downward relationship, and a flat line no linear relationship.
- The class initially focuses on one X variable; multiple predictors and nonlinear forms such as quadratic or logarithmic relationships come later.
- The fitted line gives a predicted Y for each X value, while individual observations may have actual outcomes above or below that prediction.
- The regression error is defined as actual Y minus predicted Y and is denoted ε in Professor Lantis’s notation.
- People with the same study hours can earn different exam scores, so the fitted value represents an average prediction rather than a guarantee for each person.
- A useful regression line should make predictions close to actual outcomes overall, even though some individual errors may remain large.
- Ordinary least squares selects the line that minimizes the sum of squared residuals across observations.
- Squaring makes positive and negative errors contribute positively instead of canceling each other out.
- Squaring also penalizes large prediction errors disproportionately: an error of 2 contributes 4, compared with 1 for an error of 1.
- The objective is to improve fit across the full dataset, not to eliminate the error for every individual observation.
- Minimizing the squared-error expression with respect to β₀ and β₁ produces two equations for the two regression coefficients.
- The estimated slope equals the sample covariance between X and Y divided by the sample variance of X.
- The estimated intercept is the mean of Y minus the estimated slope multiplied by the mean of X.
- Professor Lantis skips the full calculus and algebra derivation, noting that students will use Stata rather than calculate the coefficients by hand.
- The class opens a county-level birth dataset in Stata and plots births against the number of shale wells.
- The motivating concern is that fracking wells might harm reproductive health and reduce births in nearby areas.
- Many counties have zero wells, making the full scatterplot difficult to interpret visually.
- A regression can be restricted to counties with a nonzero well count using a condition such as wells != 0.
- In Stata’s regression command, the outcome variable comes first and the explanatory variable second; the constant term is the intercept β₀.
- The null hypothesis of no linear effect is β₁ = 0, and the estimated slope is evaluated with a test statistic and p-value.
- For the restricted shale-well sample, the reported test statistic is about 0.03 and the p-value about 0.97, so the class finds no evidence of a relationship with births.
- A small test statistic fails to exceed the two-sided critical values 1.65, 1.96, or 2.58, illustrating the difference between a test statistic and a p-value.
- For the full dataset, the covariance of births and shale wells divided by the wells variance is about −3.92, matching the regression slope when the sample restriction is removed.
- Stata reports a 95% confidence interval by default; specifying a 99% level produces a wider interval around the slope estimate.
- A slope confidence interval is centered on the estimated slope, and whether it includes zero corresponds to whether the zero-slope null can be rejected at that confidence level.
- For example, an interval from 6 to 14 excludes zero, while a proposed interval from 6 to 16 is not centered on an estimate of 10 and would be inconsistent with the usual symmetric construction.
- Students must enter the in-class quiz code 23.81 and submit a response to every question before leaving to receive attendance credit.
- Professor Lantis says initially incorrect responses will be regraded for full credit when students have submitted answers and attended class.
- Problem-set submissions may use an annotated Stata do-file alone or a do-file plus a Word document, but the code used must be included.
- Students should check Canvas for the Tuesday-class replacement plan and watch the expected recording before Thursday.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.