A significant one-way ANOVA tells you that not all group means are equal. It doesn't tell you which ones differ. With four groups there are six possible pairs, and the F-test is silent on all of them. Post hoc tests fill that gap by comparing pairs of groups while keeping the overall false-positive rate under control.
There are many post hoc procedures. In practice, four cover nearly every situation: Tukey's HSD, Bonferroni, Games-Howell, and Dunnett's test. This article explains what each one does, when to use it, and how they compare on the same data.
Why not just run t-tests on every pair?
Each t-test at α = 0.05 has a 5% chance of a false positive. Run several and the chances compound. The familywise error rate, the probability of at least one false positive across all comparisons, is approximately 1 − (1 − α)^m for m independent comparisons:
Groups Pairwise comparisons Familywise error rate at α = 0.05
3 3 14%
4 6 26%
5 10 40%
6 15 54%
With six groups, uncorrected t-tests produce at least one false positive more often than not. Post hoc tests adjust the comparisons so the familywise error rate stays at α.
The example
Four suppliers deliver the same machined pin. Six pins from each are measured for hardness (HRC):
Supplier n Mean SD
A 6 48.97 0.85
B 6 51.97 0.72
C 6 49.20 0.84
D 6 51.07 0.80
The ANOVA gives F(3, 20) = 19.58, p < 0.0001, with MSW = 0.649. Hardness differs across suppliers. The question is which ones.
The six pairwise differences are:
Pair Difference
B − A 3.00
B − C 2.77
D − A 2.10
D − C 1.87
B − D 0.90
C − A 0.23
Tukey's HSD
Tukey's honestly significant difference test is the standard choice for comparing all pairs when group variances are similar.
It uses the studentized range distribution, which describes the spread between the largest and smallest of k sample means. The critical difference for equal group sizes is:
HSD = q(α; k, df_W) × √(MSW / n)
For this example, q(0.05; 4, 20) = 3.958, so:
HSD = 3.958 × √(0.649 / 6) = 3.958 × 0.329 = 1.30 HRC
Any pair differing by more than 1.30 is significant:
Pair Difference Tukey p Significant?
B − A 3.00 < 0.001 Yes
B − C 2.77 < 0.001 Yes
D − A 2.10 0.001 Yes
D − C 1.87 0.004 Yes
B − D 0.90 0.246 No
C − A 0.23 0.958 No
Conclusion: the suppliers fall into two clusters. B and D are harder than A and C. Within each cluster, the suppliers can't be distinguished.
For unequal group sizes, the Tukey-Kramer variant replaces √(MSW/n) with √(MSW/2 × (1/nᵢ + 1/nⱼ)) for each pair. Most software applies this automatically.
Use Tukey when: you want all pairwise comparisons, variances are roughly equal, and group sizes are equal or not too different.
Bonferroni
The Bonferroni correction divides α by the number of comparisons and runs ordinary t-tests at the stricter level. With six comparisons, each is tested at 0.05 / 6 = 0.0083.
Equivalently, multiply each unadjusted p-value by the number of comparisons (capping at 1).
Using MSW as the pooled variance and df_W = 20, the critical t is t(0.0083/2, 20) = 2.927, and the critical difference is:
2.927 × √(2 × 0.649 / 6) = 2.927 × 0.465 = 1.36 HRC
That's slightly wider than Tukey's 1.30. Here it leads to the same conclusions: the same four pairs are significant, B − D (Bonferroni p = 0.40) and C − A (p = 1.00) are not.
The trade-off. Bonferroni is simple and works for any set of comparisons, not just all pairs. But for all-pairwise comparisons it's more conservative than Tukey, and the gap widens as the number of groups grows. With six groups (15 comparisons), Bonferroni's critical difference is noticeably wider and it will miss differences Tukey detects.
Use Bonferroni when: you have a small number of specific, pre-planned comparisons rather than all pairs, or when you're comparing a subset of groups where Tukey doesn't apply.
Games-Howell
Tukey and Bonferroni both use MSW, a single pooled variance, as the error term for every comparison. That's only appropriate if the groups have similar variances. When they don't, and especially when group sizes also differ, both can be misleading.
Games-Howell doesn't pool. For each pair it uses only the two groups' own variances and sample sizes, computes Welch-style degrees of freedom for that pair, and refers the result to the studentized range distribution. It's the post hoc counterpart to Welch's ANOVA.
For the hardness data, the variances are similar (SDs from 0.72 to 0.85), so Games-Howell agrees with Tukey:
Pair Games-Howell p
B − A < 0.001
B − C < 0.001
D − A 0.006
D − C 0.013
B − D 0.236
C − A 0.962
The p-values are a little larger because each comparison uses only about 10 degrees of freedom rather than the pooled 20. That's the cost of not assuming equal variances. When variances really are unequal, it's a cost worth paying, because Tukey's p-values would then be wrong.
Use Games-Howell when: group variances differ (a ratio of largest to smallest SD above about 2), particularly with unequal group sizes, or whenever you ran Welch's ANOVA.
Dunnett's test
Sometimes you don't want every pair. You want to compare each group against a single reference: a control, the current supplier, the existing process. Dunnett's test is designed for exactly this. With k groups there are only k − 1 comparisons instead of k(k − 1)/2, so it's more powerful than Tukey for that question.
If supplier A is the incumbent:
Comparison Dunnett p
B vs A < 0.001
C vs A 0.923
D vs A 0.001
B and D are harder than the incumbent; C is not distinguishable from it.
Use Dunnett when: the question is "which groups differ from the control?" and comparisons among the non-control groups don't matter.
Choosing a test
Situation Test
All pairs, similar variances Tukey's HSD (Tukey-Kramer for unequal n)
All pairs, unequal variances Games-Howell
Each group vs one control Dunnett
A few specific pre-planned comparisons Bonferroni (or planned contrasts)
Choose the test before looking at the results. Picking whichever one makes your preferred difference significant defeats the purpose of the correction.
Do you need a significant ANOVA first?
Traditionally, post hoc tests are run only after a significant omnibus F. For Tukey and Games-Howell, this gatekeeping isn't strictly necessary, since both control the familywise error rate on their own. Running them only after a significant F makes them slightly more conservative. Dunnett's test and planned contrasts are often run without the omnibus test at all, since they answer a more specific question.
In practice, following the convention (ANOVA first, then post hoc if significant) is fine and expected by most readers.
Reporting
Report the post hoc test by name, the critical difference or adjusted p-values, and the conclusion in terms of the means:
"Hardness differed across suppliers, F(3, 20) = 19.58, p < 0.001. Tukey's HSD (critical difference 1.30 HRC) showed that suppliers B (M = 51.97) and D (M = 51.07) were harder than A (M = 48.97) and C (M = 49.20); B and D did not differ from each other, nor did A and C."
A compact display often used in tables is letter grouping: groups sharing a letter don't differ significantly. Here, B and D would be "a," A and C "b."
Running the analysis
The ANOVA table supplies everything the post hoc tests need: MSW, df_W, and the group means and sizes. A one-way ANOVA calculator that reports MSW and group statistics gives you the inputs for computing Tukey's HSD by hand, and is a quick way to confirm the omnibus result before moving on to pairwise comparisons.
Summary
A significant ANOVA says a difference exists; post hoc tests find it while controlling the familywise error rate. Use Tukey's HSD for all pairwise comparisons with similar variances, Games-Howell when variances differ, Dunnett when comparing each group against a control, and Bonferroni for a small set of planned comparisons. Choose before you see the results, and report the conclusion in terms of which means differ and by how much.