🎯 Galton Board S1
Watch balls cascade through a grid of pegs — each bounce is an independent 50/50 left-or-right decision. Yet collectively they build a perfect bell curve: the Binomial B(n, 0.5) converging to the Normal distribution.
Expected: \(\mu = np = 6,\ \sigma = \sqrt{np(1-p)} \approx 1.73\). The red curve is the theoretical \(N(6,\,1.73^2)\) — watch the histogram converge to it!
🔔 Normal Distribution Explorer S1
The x-axis window is fixed, so you can truly see what \(\mu\) and \(\sigma\) do: \(\mu\) slides the bell left and right, \(\sigma\) makes it taller and narrower or shorter and wider. Shaded bands show the Empirical Rule (~68%, ~95%, ~99.7%).
Tip: Set \(\sigma = 0.2\) and watch the spike; set \(\sigma = 5\) and watch it flatten. The total area is always 1!
📊 Distributions Comparison S1 FS
Compare Binomial, Poisson and Normal side by side. Watch how Poisson approximates Binomial when n is large & p is small, and how Normal approximates both when \(np\ge 5\) and \(n(1-p)\ge 5\).
Approximation rules: Poisson \(\approx\) Binomial when \(n\ge 20\) and \(p\le 0.05\) (\(\lambda=np\)). Normal \(\approx\) Binomial when \(np\ge 5\) and \(n(1-p)\ge 5\).
🎲 Law of Large Numbers S1
Roll dice repeatedly and watch the sample mean converge to the expected value. The running average bounces wildly at first, then settles near E(X) — that is the Law of Large Numbers in action.
For a fair die \(E(X)=3.5\), so for \(n\) dice the average of sums converges to \(3.5n\). The lower histogram shows the frequency distribution approaching the true shape.
🧠 Central Limit Theorem FS
Sample from any distribution — uniform, exponential, even wildly skewed — and the distribution of sample means becomes Normal. Increase the sample size n and watch the green theoretical curve \(N(\mu,\,\sigma^2/n)\) hug the histogram.
Even the highly skewed Exponential produces a bell-shaped sampling distribution for n = 30+. Try n = 1 (no averaging) vs n = 50!
📏 Confidence Intervals FS
Generate many 95% confidence intervals from a population with known mean \(\mu = 50\). Each new sample gives a different interval — and about 95% of them should contain \(\mu\). Try raising n to see the intervals shrink!
Formula: \(CI=\bar{X}\pm z\dfrac{\sigma}{\sqrt{n}}\)
Green = contains \(\mu\). Red = misses \(\mu\). Coverage rate should approach the confidence level. Key: \(\mu\) is fixed — it is the interval that changes with each sample!
🌿 Stem-and-Leaf Plot S1
A quick way to display and order raw data while keeping every value. The stem holds the leading digits, the leaf is the final digit. Great for finding the median and quartiles by eye.
Tip: the shape of the leaves tells you the shape of the distribution — check Skewed or Bimodal!
📦 Box-and-Whisker Plot S1
The box shows the middle 50% of the data (Q1 to Q3), the line inside is the median, and the whiskers reach the smallest / largest values within \(1.5\times\text{IQR}\). Anything beyond is an outlier — plotted as a dot.
Reading shape: if the right whisker is longer than the left, the data is skewed right; a box off-centre tells the same story.
📊 Histogram — Frequency Density S1
In a histogram the area of each bar is proportional to its frequency, so bar height must be frequency density = frequency ÷ class width. Toggle between equal and unequal class widths to see why!
Check: add up every (density × width) — it must equal the total frequency.
📈 Cumulative Frequency Graph S1
Plot cumulative frequency against the upper class boundary to get the S-shaped ogive. Draw a horizontal line at k% of n to read the k-th percentile — the median is the 50th percentile.
Median = 50th percentile, Q1 = 25th, Q3 = 75th. The graph reading uses interpolation inside each class, so it may differ slightly from the exact value — exactly as in the exam!
🎲 Geometric Distribution S1
X = number of trials needed until the first success, with \(P(\text{success})=p\) on every trial. The PMF is a never-ending decreasing staircase: \(P(X=x)=q^{x-1}p\).
What is k? k is the specific number of trials you want to examine. The orange bar shows \(P(X=k)\) — the chance that the first success happens exactly on trial k. The green dot shows \(P(X\le k)\) — the chance of success within k tries. Left axis = PMF, right axis = CDF.
Play: run 1000 experiments — the histogram of "trials to first success" should hug the theoretical bars. The longest bar is always at x = 1!
🧪 Hypothesis Testing (Normal) S1
Test \(H_0:\ \mu=\mu_0\) against an alternative using the sample mean. If the observed \(\bar{x}\) lands in the rejection region (red), we reject \(H_0\) — but we could be wrong with probability \(\alpha\) (Type I error).
↓ Standardised onto N(0,1): same test, \(z=(\bar{x}-\mu_0)/SE\) ↓
Reject \(H_0\) if \(|z| > z^*\) (or if p-value \(< \alpha\)). \(\alpha\) = P(rejecting \(H_0\) when it is true) = Type I error.
📐 t-Distribution Hypothesis Test FS
When \(\sigma\) is unknown and estimated from the sample, the test statistic follows a t-distribution with \(\text{df} = n-1\). Set up the test just like the normal test — but now you supply the sample standard deviation \(s\), and the heavier tails of \(t\) give a larger critical value than \(z\).
t vs z: because \(s\) is only an estimate of \(\sigma\), the t-distribution has heavier tails — so \(t^* > z^*\) and you need stronger evidence to reject. As \(n \to \infty\), \(t \to N(0,1)\). Toggle "Show N(0,1)" to see them merge!
🧩 Pooled-Variance Two-Sample t-Test FS
Compare two independent sample means assuming the two populations have a common variance. The pooled variance combines both samples: \(s_p^2=\dfrac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2}\)
Try: set both true means equal (e.g. 50 and 50) and run 1000 tests — about \(\alpha = 5\%\) should be rejected (Type I error). Then move \(\mu_2\) apart and watch the power grow.
Why is \(s_p^2=\dfrac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2}\) the right way to estimate the common variance?
1. The assumption. The two samples come from Normal populations with a common variance \(\sigma^2\) (and are drawn independently). We never observe \(\sigma^2\) directly — only the two sample variances \(s_1^2\) and \(s_2^2\), which are both noisy estimates of the same number.
2. Each sample's information is measured in degrees of freedom. For a Normal sample, \(\dfrac{(n-1)s^2}{\sigma^2}\sim\chi^2(n-1)\). A larger sample gives a more reliable variance estimate, and its weight is exactly its d.f. So to combine two estimates of the same \(\sigma^2\), we weight each by its own d.f.: \[(n_1-1)s_1^2+(n_2-1)s_2^2\]
3. Chi-square additivity. The samples are independent, so the two chi-square pieces add up: \[\frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{\sigma^2}\sim\chi^2(n_1+n_2-2).\] Dividing by \(n_1+n_2-2\) therefore gives an unbiased estimator of \(\sigma^2\) that uses every observation from both samples — exactly as if the two samples were merged into one big sample of size \(n_1+n_2\).
4. Intuition. It is a weighted average of \(s_1^2\) and \(s_2^2\) in which the bigger sample dominates. In the balanced case \(n_1=n_2=n\) it simplifies to \((s_1^2+s_2^2)/2\) with d.f. \(2n-2\).
5. What if \(\sigma_1^2\ne\sigma_2^2\)? Then pooling is invalid — the correct tool is Welch's t-test (no pooling, fractional d.f.). In this demo, "Simulate datasets" draws both groups from the same true \(\sigma\), so you can check that the empirical rejection rate stays at \(\alpha\) when \(H_0\) is true — evidence that the pooled formula does its job.
🏆 Wilcoxon & Sign Tests FS
Three non-parametric tests in one module. Rank-Sum for two independent samples; Sign Test and Signed-Rank for paired data. No Normality assumption needed!
Simulation: under \(H_0\) (both groups from the same population) the empirical rejection rate should be near \(\alpha = 5\%\). Now set \(\delta = 1\) and watch it jump!
🔢 Chi-Squared Goodness-of-Fit S1 FS
Compare observed frequencies with expected frequencies. The test statistic \(\chi^2=\sum\dfrac{(O-E)^2}{E}\) is large when the data disagree with the model. With \(df=k-1\), reject the model if \(\chi^2\gt\chi^{2*}(\alpha)\).
| Category | Observed O | Expected E | \((O-E)^2/E\) |
|---|
Fun: run 1000 tests on a fair die — about \(\alpha = 5\%\) get rejected (Type I error). Switch to the biased die and see the rejection rate climb!