STAT 638 (Fall 2026), Lecture 1
Watch on YouTube →
Overview
Samiran Sinha introduces STAT 638’s Bayesian analysis curriculum and its hybrid 600 on-campus/700 online format, then explains assessment rules and a substantial group project worth 45% of the course grade. The lecture develops the Bayesian framework—combining a data model and prior to obtain a posterior—and teaches Bayes’ rule through a COVID-19 testing example before reviewing discrete and continuous random variables.
Key takeaways
- Bayesian inference combines a probability model for observed data with a prior distribution on unknown parameters, then uses the resulting posterior to summarize estimates and uncertainty.
- With 10% disease prevalence, 90% test sensitivity, and a 5% false-positive rate, only 66.67% of positive results correspond to infection; the positive-test probability is 13.5%.
- STAT 638’s project is worth 45% of the grade and requires a reproducible R Markdown/HTML report, a meaningful Bayesian extension, comparison with another method, and an individual Q&A component.
- Quizzes and the exam are timed, open-book, not open-internet, and unproctored; each is available in Canvas for 24 hours, but its timer begins when the student starts.
- For both discrete and continuous random variables, expectations and cumulative probabilities follow parallel definitions: sums over probability masses for discrete variables and integrals over densities for continuous variables.
Chapters
0:00
STAT 638 Sections, Class Logistics, and Bayesian Course Goals
- STAT 638 has a 600 on-campus section and a 700 online section, with identical syllabi and recorded lectures.
- Class meets Monday, Wednesday, and Friday from 10:20 to 11:10; Samiran Sinha offers office hours Monday and Wednesday from 11:30 to 12:30, in person and on Zoom.
- The course covers Bayesian models, prior and posterior distributions, uncertainty quantification, and computational methods; the stated prerequisite is STAT 630 or equivalent.
- The optional textbook is Peter Hoff’s A First Course in Bayesian Statistical Methods.
4:00
Quizzes, Exams, and a Grading Plan Adapted to Hybrid Learning
- Grades are based on four quizzes, one exam, and a project; homework and posted solutions are for practice and are not collected or graded.
- Sinha describes moving away from graded homework because students can readily access generative AI solutions.
- Quizzes and the exam are timed, open-book, not open-internet, and unproctored; the course includes no proctoring for these assessments.
- The exam is scheduled for November 6 and covers material through November 3; project presentations are tentatively scheduled for November 23, November 30, December 2, and December 3.
8:00
24-Hour Assessment Windows and the 45% Bayesian Project
- Quizzes and the exam are available in Canvas for 24 hours, from 6 a.m. on the assessment day to 6 a.m. the following day.
- The timer starts when a student enters the assessment password, and Canvas rejects work after time expires; starting near the end of the window can prevent receiving the full allotted time.
- The project accounts for 45% of the course grade: 25% for the written report, 10% for the presentation, and 10% for individual question-and-answer performance.
- Project groups will contain three or four students, with assignments planned by the fourth class day; detailed instructions are on Canvas.
11:00
Project Report Requirements: Reproducible R Markdown and a 2,500-Word Limit
- Each group must formulate a scientific question, conduct and evaluate a Bayesian analysis, add a meaningful extension, and compare it with an appropriate alternative method.
- The written deliverable is an R Markdown file with HTML output; the main report is limited to 2,500 words, excluding references, figures, tables, captions, supplementary material, and the AI-use statement.
- Analysis code must be included for reproducibility, while routine code should be hidden, collapsible, or placed in an appendix outside the word limit.
- Groups should communicate results with clear prose, tables, and figures rather than raw software output.
15:00
Generative AI Rules, Zoom Presentations, and Cross-Section Grouping
- Generative AI may assist with coding, debugging, statistical reasoning, and writing, but each group remains responsible for the accuracy of its submission.
- Sinha asks students to revise AI-generated writing in their own voice and avoid generic ChatGPT phrasing.
- Each group receives eight minutes to present, followed by roughly five to ten minutes of questions; students are questioned individually.
- Students seeking cross-section teammates should post their full name and preferred contact email in the Canvas discussion by August 28; Sinha plans 12 distinct project topics.
20:00
Canvas Resources, Recorded Lectures, and Practice Materials
- Class lectures are recorded and posted to YouTube, where captions are available.
- Canvas includes lecture materials, old exams from 2024 and 2025, and homework solutions.
- Homework is ungraded practice, with tentative dates provided to help students keep pace with the course.
- Sinha shows an HTML report generated from an R Markdown file as a model for project presentation.
23:30
Bayesian and Frequentist Inference: Priors, Parameters, and Uncertainty
- Statistical analysis uses a sample as a snapshot of a population to make inferences about population characteristics.
- Frequentist inference treats unknown parameters as fixed and bases inference on observed data and repeated-sampling behavior; Bayesian inference models unknown parameters as random variables.
- Bayesian analysis combines a data-generating probability model with a prior distribution to produce a posterior distribution.
- Posterior means, medians, and percentiles summarize parameter values and uncertainty; computation can be challenging for complex models.
29:00
Bayes’ Rule from Sample Spaces to Partitioned Events
- A probability is assigned to an event in an experiment’s sample space and lies between zero and one.
- A partition consists of mutually exclusive events whose union is the full sample space; a coin toss split into heads and tails is a simple example.
- For a partition A₁,…,Aₘ and event B, Bayes’ rule gives P(Aⱼ | B) = P(B | Aⱼ)P(Aⱼ) / P(B).
- The denominator is the marginal probability of B, computed by summing P(B | Aᵢ)P(Aᵢ) across the partition.
33:30
COVID-19 Test Example: Why a Positive Result Does Not Mean Certainty
- The example assumes 10% COVID-19 prevalence, 90% sensitivity, and a 5% false-positive rate among people without infection.
- The probability of a positive result is 0.90 × 0.10 + 0.05 × 0.90 = 0.135, or 13.5%.
- Bayes’ rule gives P(infected | positive) = 0.09 / 0.135 = 66.67%, despite the test’s 90% sensitivity.
- A tree diagram separates infected and healthy cases, then positive and negative results, making the mutually exclusive paths and conditional probabilities explicit.
45:30
Discrete and Continuous Random Variables, Distributions, and Expectations
- A random variable maps outcomes in a sample space to values; a discrete variable is described by a probability mass function whose probabilities sum to one.
- For a discrete variable, the cumulative distribution function sums probabilities up to a chosen value, and expectation is the probability-weighted sum of possible values.
- A continuous variable has a probability density that integrates to one; the density at a point is not itself a probability, while integrating over an interval gives its probability.
- Continuous CDFs and expectations use integrals where discrete versions use sums, and the probability of any one exact value is zero.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Samiran Sinha.