Presentation slides and interactive tools for study design and everyday data analysis — sample size & power, general statistics, survival analysis, and time series.
Slide decks from workshops and training sessions on biostatistics, epidemiology, and study design.
Introductory slides on getting started with R for statistical analysis.
DownloadBasic z‑approximation formulas for common study designs. For complex or adaptive designs, consult a biostatistician before finalising your protocol.
n = z² · p(1−p) / d² (with finite population correction if N is provided)
Two‑sided z‑test for two independent proportions (normal approximation).
Two‑sided z‑test for the difference between two independent means.
n = (z · SD / E)²
These calculators use standard normal (z) approximations intended for planning and teaching purposes — always confirm with exact methods / statistical software before finalising a protocol.
Paste a list of numbers to get descriptive statistics, run a hypothesis test (parametric or non‑parametric), a chi‑square test, or build a confidence interval.
Assumes the sample is approximately normally distributed. If it isn't (skewed, small n, outliers), use the Wilcoxon Signed‑Rank test under Non‑Parametric instead.
Welch's t‑test — assumes both groups are approximately normally distributed. If not, use Mann‑Whitney U under Non‑Parametric instead.
One‑sample / paired‑differences test — no normality assumption. Tests whether the median differs from a hypothesised value (for paired data, enter the differences and set μ₀ = 0).
Normal approximation with continuity correction and tie correction — suitable for n ≳ 10; very small samples should use exact tables.
Two independent samples, no normality assumption — tests whether one group tends to have larger values than the other.
Normal approximation with continuity correction and tie correction — suitable for moderate/larger n; very small samples should use exact tables.
Enter a 2×2 contingency table (e.g. exposed/unexposed vs. disease/no disease).
CI = x̄ ± t(α/2, df) · SE, using Student's t distribution (df = n − 1).
CI = x̄ ± t(α/2, df) · (SD/√n), using Student's t distribution (df = n − 1).
Wald: p̂ ± z·√(p̂(1−p̂)/n). Wilson score and Agresti–Coull generally perform better for small n or p̂ near 0 or 1.
CI = (x̄₁−x̄₂) ± t(α/2, df) · SE, Welch's method (no equal‑variance assumption).
Wald: (p̂₁−p̂₂) ± z·SE. Wilson/Agresti–Coull use the Newcombe/MOVER method, combining each group's individual interval.
Enter a 2×2 table (e.g. exposed/unexposed vs. disease/no disease) — same layout as the Chi‑square tool.
CI = exp[ln(OR) ± z(α/2) · √(1/a + 1/b + 1/c + 1/d)]. A 0.5 continuity correction is applied to all cells automatically if any cell is 0.
Enter a 2×2 table (e.g. exposed/unexposed vs. disease/no disease) — same layout as the Chi‑square tool.
CI = exp[ln(RR) ± z(α/2) · √(1/a − 1/(a+b) + 1/c − 1/(c+d))]. A 0.5 continuity correction is applied to all cells automatically if any cell is 0.
Educational tool for quick exploratory analysis — not a substitute for full statistical software (R, STATA, SAS) for publication‑grade results.
Estimate a survival curve from time‑to‑event data with censoring, or compare two groups' survival with the log‑rank test.
One row per subject: time,status — status = 1 for an event (e.g. death, relapse), 0 for censored. Comma or newline separated.
S(t) = ∏tᵢ≤t (1 − dᵢ/nᵢ) — the Kaplan–Meier product‑limit estimator. Greenwood's formula gives the CI at each step.
Same time,status format, one dataset per group — tests whether the two groups' survival curves differ.
χ² = (O₁−E₁)² / V, df = 1, compared across every distinct event time in the combined dataset.
Educational tool for quick exploratory survival analysis — not a substitute for full statistical software (R's survival package, STATA, SAS) for publication‑grade results.
Fit a linear trend to a sequence of measurements, or smooth it with a moving average.
Enter values in time order (equally spaced) — e.g. monthly case counts. Fits y = a + b·t and tests whether the slope b differs from zero.
Ordinary least‑squares regression of the series on its time index; the slope's significance is tested with a t‑test (df = n − 2).
Smooths short‑term fluctuations to reveal the underlying pattern.
Simple (centred‑lag) moving average: each smoothed point averages the current value and the previous (window − 1) values.
Educational tool for quick exploratory analysis — for seasonality, autocorrelation, or forecasting, use dedicated time‑series software (R, Python statsmodels).