ECON E371 Class Recording 8/27
Watch on YouTube →
Overview
Professor Lantis reviews how means, variances, medians, skewness, and kurtosis describe data, then introduces Stata as the course tool for analyzing real datasets. The class works through launching Stata via IU Anywhere and Citrix, organizing OneDrive files, writing and running saved do-files, importing CSVs, managing paths and .dta files, and using commands to inspect, filter, and summarize data.
Key takeaways
- For the values 2, 4, and 6, the mean is 4 and the population variance is 8/3; squaring deviations makes distance from the mean count positively and disproportionately.
- A mean greater than the median signals positive skew because unusually high observations pull the mean upward more than the median.
- A saved Stata do-file preserves commands, but dataset changes require saving the data separately as a `.dta` file.
- A global path and a consistent OneDrive folder reduce broken file references when students rerun Stata commands across course assignments.
- Stata’s `sum, detail` output connects descriptive statistics to distribution shape by reporting the median, skewness, and kurtosis alongside basic summary measures.
- The `keep if year == 2000` pattern filters a panel dataset to one year, while `egen` can create variables containing summary statistics for later analysis.
Chapters
0:00
Course Logistics, In-Class Quizzes, and the Statistics Review
- The first in-class quizzes begin the following week and run in Canvas for about 5–10 minutes with an access code.
- Students should bring a device that can access Canvas; Professor Lantis says quiz questions will often be reviewed together.
- The first practice-problem set is available on Canvas, and quiz questions assess Stata output rather than code-writing.
3:00
Calculating a Mean and Variance from 2, 4, and 6
- For values 2, 4, and 6, the mean is (2 + 4 + 6) / 3 = 4.
- Variance is calculated by squaring each deviation from the mean: 4, 0, and 4; dividing their sum by 3 gives 8/3.
- Squaring makes positive and negative deviations contribute positively and gives greater weight to observations farther from the mean.
- Standard deviation is the square root of variance.
7:00
Using the Median and Skewness to Describe Distributions
- The median is found by removing the lowest and highest values in turn until the center observation remains.
- An extremely high value, such as Mark Zuckerberg’s income in an income dataset, can pull the mean rightward while leaving the median relatively unchanged.
- When the mean exceeds the median, the distribution has positive or right skew; when the mean is below the median, it has negative or left skew.
- Skewness relates to the direction of deviations; kurtosis describes tail thickness, but Professor Lantis says neither measure will be a major focus.
12:00
Launching Stata Through IU Anywhere and Citrix
- Students authorize OneDrive through IU’s cloud-storage setup before opening Stata from IU Anywhere’s Apps list.
- Citrix Workspace must be installed to launch IU Anywhere applications, and users may need to permit the application to access their device.
- The class encounters slow or failed connections in the room; Professor Lantis suggests trying IU Secure Wi-Fi and following up on technical problems after class or during office hours.
15:00
Saving Stata Commands in a Do-File
- Stata’s Results window displays command output, while the Do-file Editor stores code that can be saved and run again.
- Typing commands only in the command box is inefficient because those commands are not saved when Stata closes.
- A Stata do-file is a text file with the .do extension; students can open one through File > Open and reuse its commands.
18:00
Importing the State Cigarette-Tax CSV
- The class imports a Canvas CSV containing state names, years, cigarette pack prices, and taxes per pack.
- File > Import provides a preview of the data before Stata converts it into the working dataset.
- The import interface generates Stata code that can be copied into a do-file and executed later.
- Professor Lantis recommends keeping course datasets together in one OneDrive folder to simplify file locations.
23:00
Troubleshooting Citrix, OneDrive, and File Access
- A failed Stata launch can indicate a Citrix Workspace setup problem; the exact fix may vary by device.
- Do-files and datasets must be accessible through OneDrive for the IU Anywhere environment to open them.
- Students should distinguish importing data from opening a do-file: CSVs are imported, while saved do-files are opened through File > Open.
- When a permission prompt appears, students may need to choose an option such as always allowing access.
29:00
Fixing File Paths with a Global Directory
- Imported-file commands include a path specific to the computer and folder where the dataset is stored.
- Moving a dataset to a different folder breaks commands that still point to its old location.
- Professor Lantis demonstrates defining a global path once and referring to it in later commands, reducing repeated path edits.
- Keeping files in one course folder makes the global path stable across weekly datasets.
37:00
Opening Saved Data and Understanding Stata’s Data Windows
- The Data Browser displays variable names as columns and observations as rows; the Data Editor can change values and is not recommended for this workflow.
- The cigarette-tax file initially appears to contain multiple states in one year, but scrolling reveals observations across many years.
- Multiple states observed over multiple years form panel data, rather than a single-year cross-section.
- Stata marks text variables in red and numeric variables in black; numeric variables can be used for calculations such as means.
43:00
Renaming Variables Without Changing the Original Data
- The `rename` command makes opaque or overly long variable names easier to recognize, such as changing a state-description field to a clearer name.
- Stata variable names cannot contain spaces; an underscore can separate words, as in `state_name`.
- Renaming changes the working dataset, so the updated data must be saved separately if the change should persist.
49:00
Saving Stata Data as a .dta File
- Saving the do-file preserves commands, but it does not save changes made to the dataset.
- Use File > Save As in Stata’s main program to save the working data as a `.dta` file.
- A saved `.dta` file can be reopened through File > Open and retains changes such as renamed variables.
- The interface displays the corresponding save command, which can be copied into a do-file.
54:00
Replacing Files, Adding Comments, and Filtering Observations
- When saving over an existing Stata dataset through code, add the `replace` option; the interface can generate it after a confirmation prompt.
- Two backslashes mark comments in a do-file, allowing students to document commands and write interpretations alongside executable code.
- The `keep if` command can isolate observations meeting a condition, such as retaining only records with year equal to 2000.
- Stata uses `==` for an equality condition; greater-than and less-than signs are also available for filters.
1:02:00
Summarizing State Cigarette Taxes and Interpreting Skew
- Stata’s `sum` command reports the mean, standard deviation, minimum, maximum, and observation count for a variable such as state cigarette tax.
- The Statistics menu can produce descriptive statistics, while the command-line version helps students learn reusable syntax.
- Adding `detail` to the summary command reports additional distribution information, including the median, skewness, and kurtosis.
- A mean above the median and a positive skewness measure both indicate a right-skewed distribution.
1:10:00
Creating Mean, Median, and Standard-Deviation Variables with egen
- The `egen` command creates a new variable using a built-in function, such as storing the mean of the state-tax variable for every observation.
- The class also generates variables for the median and standard deviation, then checks that their values match the descriptive-statistics output.
- Generated variables can support later comparisons, such as calculating average taxes for states in different regions.
- Professor Lantis recommends checking Stata syntax with reliable searches because AI-generated Stata code can contain serious errors.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.