STAT 638 (Fall 2026), Lecture 17
Watch on YouTube →
Overview
Samiran Sinha develops a Bayesian hierarchical normal model for comparing mean pain thresholds across four hair-color groups, emphasizing that assigning prior distributions to hyperparameters enables information sharing between groups. He derives Gibbs sampling conditionals for group means, population mean, within-group variance, and between-group variance, then demonstrates a 50,000-iteration R implementation using sufficient statistics and autocorrelation and trace plots for convergence diagnostics.
Key takeaways
- A Bayesian hierarchical model shares information across groups only when hyperparameters such as mu and tau squared are assigned prior distributions rather than fixed numerical values.
- The normal hierarchical model has conjugate Gibbs updates: inverse-gamma conditionals for tau squared and sigma squared, and normal conditionals for mu and every theta_i.
- The conditional update for theta_i combines the group sample mean with the population mean mu, producing shrinkage toward the overall mean with precision contributions n_i divided by sigma squared and 1 divided by tau squared.
- Group sample sizes, sample means, and sample variances are sufficient statistics, so the posterior computation does not require retaining individual observations.
- Sinha’s R implementation uses 50,000 Gibbs iterations and evaluates dependence with autocorrelation plots and stability with trace plots before using draws for inference.
- The identity separating within-group residual variation from variation of group means is essential for deriving the sigma squared conditional and mirrors an ANOVA sum-of-squares decomposition.
Chapters
- Samiran Sinha introduces Chapter 8 on hierarchical models and their central purpose: sharing information between groups.
- The motivating example compares mean pain thresholds among four hair-color groups.
- The inferential goals are estimating each group’s population mean and testing whether the group means are equal.
- Observations in group i follow a normal distribution with mean theta_i and common variance sigma squared.
- Group means theta_1 through theta_M follow a normal distribution with mean mu and variance tau squared.
- The hyperparameters mu, tau squared, and sigma squared receive prior distributions, including normal and inverse-gamma priors.
- Fixing the higher-level parameters numerically would eliminate hierarchical information sharing; they must remain random under priors.
- The unknown quantities are the M group means, mu, tau squared, and sigma squared, while prior hyperparameters are treated as known.
- The joint likelihood combines the data likelihood with normal priors for theta_i, an inverse-gamma prior for tau squared, an inverse-gamma prior for sigma squared, and a normal prior for mu.
- The posterior is proportional to this joint likelihood, but direct computation of the full joint posterior is difficult because of the required integrations.
- Gibbs sampling updates one parameter at a time conditional on the remaining parameters and all observed data.
- The conditional distributions are tractable: inverse-gamma for tau squared and sigma squared, and normal for mu and each theta_i.
- Normal random variables can be sampled directly with standard statistical software; inverse-gamma draws can be generated by taking the reciprocal of a gamma draw.
- Sinha stresses that AI-generated likelihoods and conditionals should be checked because they can be incorrect in some cases.
- Collecting the tau squared terms from the joint posterior produces an inverse-gamma kernel.
- The conditional shape parameter is based on the number of groups M and the prior parameter eta_0, specifically (M + eta_0) divided by 2.
- The scale term combines the squared deviations of theta_i from mu with the prior contribution involving eta_0 and tau_0 squared.
- For computation, Sinha distinguishes inverse-gamma scale from the gamma rate used before taking the reciprocal.
- The sigma squared conditional is obtained by collecting the sigma squared terms in the likelihood and its inverse-gamma prior.
- The residual sum of squares is decomposed into within-group variation and variation of each group mean around theta_i.
- For group i, the identity separates the data into (n_i - 1)s_i squared plus n_i times the squared difference between the sample mean and theta_i.
- The resulting sigma squared conditional is inverse-gamma with a shape determined by the total sample size and the prior parameter nu_0.
- Terms involving mu come from the normal distribution of the theta_i values and the normal prior centered at mu_0.
- Writing theta_i as (theta_i minus theta-bar) plus (theta-bar minus mu) eliminates the cross-product because the centered deviations sum to zero.
- Completing the square yields a normal conditional distribution for mu.
- The conditional variance is the inverse of M divided by tau squared plus 1 divided by tau_0 squared; the mean is the corresponding precision-weighted combination of theta-bar and mu_0.
- For each group, only the terms involving theta_i are retained: the group contribution n_i times (y-bar_i minus theta_i) squared and the hierarchical term (theta_i minus mu) squared.
- Completing the square gives a normal conditional for theta_i with precision n_i divided by sigma squared plus 1 divided by tau squared.
- The conditional mean combines y-bar_i and mu, weighted by their respective precisions.
- The update demonstrates hierarchical shrinkage: group estimates are pulled toward the common population mean mu.
- All Gibbs conditionals depend on each group’s sample size n_i, sample mean y-bar_i, and sample variance s_i squared rather than on individual observations.
- Sinha identifies y-bar_i and s_i squared as sufficient statistics for the normal hierarchical analysis.
- When the example provides sample standard deviations, their squares are used as the group sample variances.
- The coding setup specifies prior values such as mu_0 = 50, tau_0 squared = 50 squared, nu_0 = 1, sigma_0 squared = 10, eta_0 = 1, and tau_0 squared = 10.
- Sinha initializes theta, sigma squared, tau squared, and mu to 1 and runs 50,000 Gibbs iterations, storing each draw.
- Each loop samples theta_i, sigma squared, tau squared, and mu from their derived full conditional distributions.
- Autocorrelation is noticeable at lags 1 and 2, especially for tau squared and sigma squared, but declines at larger lags.
- Trace plots fluctuate around stable horizontal levels without visible trends, supporting convergence; skewness is apparent for sigma squared and tau squared, so diagnostics should precede statistical inference.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Samiran Sinha.