Before a capability study means anything, the measurement system that generated the data has to be trustworthy. A process can have genuinely tight variation and still fail a capability requirement — not because the process is bad, but because the gage measuring it is adding so much noise that the data no longer reflects reality. That's what Gage R&R (Repeatability and Reproducibility) exists to catch. ## Two Sources of Measurement Variation Repeatability is the variation you get from the same operator measuring the same part with the same gage, repeatedly. It reflects the gage itself — its resolution, its mechanical consistency, how sensitive it is to exactly how the operator positions the part. Reproducibility is the variation between different operators measuring the same parts with the same gage. It reflects technique differences — how firmly someone applies a caliper, how they read an analog scale, how consistently they align a part against a fixture. A study needs multiple operators, multiple parts spanning the expected range of variation, and multiple repeat measurements per part per operator — the standard AIAG design is 2-3 operators, 10 parts, 2-3 trials each — to separate these two sources from each other and from actual part-to-part variation. ## Reading the %GRR Result The output that matters most is %GRR — the percentage of total observed variation that's attributable to the measurement system rather than to real differences between parts. AIAG's standard thresholds: - Under 10%: the gage is acceptable. - 10–30%: acceptable depending on the application, cost of the gage, cost of repair, and criticality of the measurement — this is a judgment call, not an automatic pass. - Over 30%: unacceptable. The measurement system needs improvement before it can be trusted for process control or capability work. A %GRR over 30% doesn't necessarily mean the gage itself is broken. It can mean the parts sampled didn't span enough real variation (which inflates %GRR by shrinking the denominator), or that operator technique differs enough to need retraining, or that the gage's resolution is too coarse relative to the tolerance being measured. ## The ANOVA vs. Average and Range Method There are two calculation methods in circulation. The older Average and Range (X-bar & R) method is simpler to hand-calculate but can't separate the operator-by-part interaction from pure reproducibility — it lumps them together. The ANOVA method decomposes variation into part, operator, operator-by-part interaction, and repeatability separately, which is more accurate and is what most current AIAG guidance recommends as the default. The interaction term matters more than people give it credit for. If Operator A consistently measures certain parts higher while Operator B measures those same parts lower — a real interaction, not just noise — the Average and Range method can't detect that pattern at all. ANOVA can, and that pattern often points to something specific: a fixture that behaves differently depending on part orientation, or an instruction that different operators are interpreting differently. ## Number of Distinct Categories (ndc) %GRR isn't the only number worth checking. Number of Distinct Categories estimates how many truly distinguishable groups the measurement system can reliably separate across the part-to-part variation observed. AIAG's guidance calls for ndc of 5 or greater. A gage can pass on %GRR and still return an ndc of 2 or 3 — which means it functionally can't tell parts apart with any resolution, even though the percentage-based metric looks acceptable. Both numbers need checking; neither one alone tells the full story. ## Where Gage R&R Studies Go Wrong The most common design mistake is picking parts that don't represent real production variation — grabbing 10 parts off the line at random instead of deliberately selecting parts that span the low end to the high end of the expected range. That artificially shrinks part-to-part variation in the denominator and inflates %GRR, making a perfectly good gage look unacceptable. The second common mistake is running the study with operators who know which part they're re-measuring, which lets memory substitute for genuine repeat measurement and produces artificially tight repeatability that doesn't hold up in daily use. ## Running the Study Correctly Separating repeatability from reproducibility with the ANOVA method, checking both %GRR and ndc, and structuring the study design correctly from the start is where manual Gage R&R work most commonly breaks down. [SigmaDesk's Gage R&R calculator](https://sigmadesk.app/gage-rr-calculator/) runs the full ANOVA-based study with AIAG-standard reporting, free in the browser. It's built alongside the rest of the [SigmaDesk](https://sigmadesk.app/) SPC platform — control charts, process capability, and attribute charts under the same toolkit. Run the Gage R&R before trusting a capability study, not after the capability numbers come back and don't make sense.