Suppose ten parts are measured by three operators, three times each. The resulting 90 values do not vary only because the parts are different. They also contain equipment variation, technique variation, positioning effects, environmental effects, and possibly operator-specific responses to particular part features.
A Gage R&R study tries to estimate those sources separately. The value of the study is not the label printed at the top of the report. The value is knowing whether the measurement system can reveal the process signal without adding too much noise—and knowing what to improve when it cannot.
What Gage R&R is actually measuring
Every observed reading combines a value attributable to the part with error introduced by the measurement process. In a simplified crossed study, the measurement error is divided into repeatability and reproducibility. Part-to-part variation is desirable in the study because it provides the signal the gage must distinguish.
The word gage can be misleading because the study evaluates more than the instrument. The measurement system includes the device, fixture, software, method, operator, part presentation, environment, and the way results are recorded. A calibrated micrometer can still be part of an unsuitable measurement system if the datum is interpreted differently by each operator or the part cannot be located consistently.
The study design comes before the percentage
A crossed design means every operator measures every selected part in every trial. That creates repeated observations for the same operator–part combination and lets the analysis compare within-cell variation, operator differences, part differences, and operator × part interaction.
This balanced structure is the sample used in the product graphs below. It illustrates the logic of a crossed study; the correct design for a real study must follow the applicable customer, industry, and internal requirements.
The parts should represent the process range that matters for the decision. Selecting ten nearly identical parts can make a capable gage appear unable to distinguish part-to-part variation. Selecting only extreme parts can make discrimination look artificially strong. The study should also use the real fixture, work instruction, environmental conditions, and operators who normally perform the measurement.
Randomization matters. If an operator measures Part 1 three times consecutively, they can remember the previous value and unintentionally reproduce it. Blinding part identity where practical and mixing the measurement order reduce memory and sequence effects. The study is meant to reproduce the operating measurement process—not a cleaner laboratory version that will never exist in production.
Repeatability and reproducibility answer different questions
Repeatability
Can the same operator obtain similar values on the same part?
Repeatability is the within-operator, within-part variation under the study conditions. It is often called equipment variation, but the observed effect can also include short-term positioning, contact, resolution, fixture, and technique variation.
- instrument resolution or noise
- fixture movement or inconsistent seating
- contact force or probe alignment
- short-term environmental variation
- unclear repeated-measurement technique
Reproducibility
Does the measurement response change across operators?
Reproducibility captures differences between appraisers. In an ANOVA study it can include an overall operator effect and an operator × part interaction, depending on the selected model and treatment of the interaction.
- different datum or edge interpretation
- systematic force or alignment differences
- operator-specific setup or zeroing
- inconsistent work-instruction interpretation
- part features that affect operators differently
The distinction determines the improvement route. Replacing a gage is unlikely to correct an operator offset created by inconsistent datum selection. Retraining operators is unlikely to correct inadequate instrument resolution. The total GRR value tells you the size of the combined issue; the components tell you where to look.
ANOVA and Average & Range are not interchangeable labels
Two common calculation approaches are the Average & Range method and the analysis of variance method. Both estimate repeatability, reproducibility, Gage R&R, and part-to-part variation, but they do not use the same information or assumptions.
| Method | What it uses | Practical implication |
|---|---|---|
| Average & Range | Cell ranges, operator averages, part averages, and study constants linked to the design. | Transparent for hand calculation and teaching, but it provides a more compressed view of the sources. |
| Crossed ANOVA | Mean squares from the part, operator, operator × part, and repeatability terms. | Can estimate the interaction explicitly and provides significance tests for model terms. |
ANOVA is especially useful when the same part may be measured differently by different operators. That pattern is not just a global operator offset; it is an interaction. One operator might agree on simple cylindrical parts but respond differently when a chamfer, surface, or locating feature complicates the measurement.
Method note: MSA Studio’s ANOVA report expresses study variation as 6 × standard deviation. Some worksheets, legacy templates, or calculation guides may use different constants or spread conventions. Compare the method, model, and convention before comparing two reported percentages digit for digit.
The percentage column changes the question
A report can show %Contribution, %Study Variation, and %Tolerance for the same source. These numbers are not alternative spellings of the same result. They use different denominators and should not be compared as if they were interchangeable.
| Reported metric | Sample Total GRR | Question answered |
|---|---|---|
| %ContributionBased on variance | 0.48% | What share of the estimated total variance is attributed to the measurement system? |
| %Study VariationBased on standard-deviation spread | 6.93% | How large is the GRR study variation relative to the total variation represented in this study? |
| %ToleranceStudy variation ÷ specification width | 7.99% | How much of the engineering tolerance band is consumed by measurement variation? |
Because variance is squared, %Contribution is usually numerically much smaller than %Study Variation. Reporting “GRR is 0.48%” without naming the column would make this sample sound far better than the more decision-relevant 6.93% of study variation or 7.99% of tolerance. Always report the metric name with the value.
%Study Variation depends on the part variation captured by the sample. %Tolerance depends on the entered specification width. They can therefore support different conclusions. A broad part range can reduce %Study Variation while leaving %Tolerance unchanged. A wide tolerance can make %Tolerance look small even if the system is not discriminating finely enough for process improvement. The intended use of the measurement is what decides which comparison matters most.
ndc describes discrimination—not accuracy
The number of distinct categories estimates how many non-overlapping groups of part variation the measurement system can distinguish within the study. It is derived from the relationship between part-to-part variation and Gage R&R variation, commonly using a factor of 1.41 before truncation to an integer.
The sample can distinguish many part levels
This supports the conclusion that measurement noise is small relative to the selected part range. It does not prove that measurements are unbiased, calibrated, stable over time, or correct across the full operating range.
A low ndc may indicate excessive measurement variation, insufficient part spread, or both. That is why “select wider parts” is not automatically a corrective action: it may improve ndc without improving the measurement system. First decide whether the original parts represented the range in which the gage must make useful distinctions.
How the six Gage R&R graphs tell the story
The numerical table ranks the variation components. The graphs show the mechanism behind those values. The panel below uses the same 3-operator × 10-part × 3-trial dataset and the same six diagnostic views generated by Mechatrovich MSA Studio v1.3.1. Hover over points or bars to inspect the underlying values.
Gage R&R Diagnostic Graphs — Diameter
Crossed ANOVA sample · Digital micrometer · 0.001 mm resolution
Components of Variation
Measurements by Parts
R Chart by Operators
Measurements by Operators
Xbar Chart by Operators
Parts × Operators Interaction
These graphs are diagnostic views, not six separate acceptance tests. Read them together with the study design, variance-component table, ANOVA terms, characteristic risk, and intended measurement decision.
Start with the relative sizes. In a useful study, part-to-part variation should dominate the measurement components. Compare the same named percentage across sources; do not compare %Contribution for one bar with %Tolerance for another.
Look for clear separation among part centers with relatively tight clusters of repeated readings. Heavy overlap can reflect measurement noise, a narrow part range, or both.
This is the first repeatability check. A point above the upper control limit identifies an operator–part cell whose repeated measurements are unusually inconsistent. Investigate that cell before relying on the pooled repeatability estimate.
Compare the centers and spreads. A consistent shift between operator distributions suggests a systematic technique or setup difference. Similar centers with different spreads point more toward repeatability differences.
Unlike an ordinary process-control chart, many points outside the limits can be desirable here: the limits reflect measurement-system repeatability, so separated part means show that real part differences are visible above the measurement noise.
Parallel operator traces suggest the operator effect is reasonably consistent across parts. Crossing or changing gaps indicate that particular part features may interact with operator technique. The graph helps localize what an interaction p-value only summarizes.
Acceptance bands are a screen, not the engineering decision
Commonly taught guidance treats measurement variation below 10% as generally acceptable, 10% to 30% as potentially acceptable depending on application, and above 30% as generally unacceptable. These bands are useful for screening, but they do not remove the need to understand risk and purpose.
A characteristic used for safety, regulatory acceptance, tight process adjustment, or automated feedback can need a more demanding system than a broad screening measurement. A 9% headline result also does not cancel evidence of an unstable range, a strong operator offset, a problematic interaction, or poor study execution.
Confirm the parts, operators, fixture, method, resolution, environment, and measurement order reflect real use.
Check part IDs, operator IDs, trial sequence, missing readings, rounding, transcription, and any excluded observations.
Investigate out-of-control R-chart cells and unusually large within-cell ranges before accepting the pooled result.
Review operator centers, spreads, work methods, and any operator × part pattern.
Use study variation, tolerance, or a historical process variation intentionally—not simply whichever gives the smallest percentage.
Read ndc together with part selection and the measurement resolution needed for control or classification.
Document the characteristic importance, customer requirements, safeguards, limitations, and follow-up plan.
What a Gage R&R study does—and does not—prove
It helps estimate
Short-term repeatability, between-operator reproducibility, operator × part interaction in an ANOVA model, part-to-part variation within the selected sample, and measurement discrimination relative to that sample.
It does not automatically prove
Calibration, traceability, bias, linearity, long-term stability, correct specification limits, representative sampling, measurement uncertainty for every use, or suitability for every product and operating range.
Bias asks whether the measurement system is centered on an accepted reference. Linearity asks whether bias changes across the operating range. Stability asks whether the system changes over time. Gage R&R addresses precision components under the study conditions; these other properties need their own evidence.
Turn a poor result into a targeted investigation
“Improve the measurement system” is too broad to be actionable. Use the component and graph pattern to choose the first investigation.
Observe repeated measurements of the worst operator–part cells. Check resolution, seating, contact force, fixture movement, zeroing, surface condition, and environmental sensitivity.
Compare setup, datum selection, alignment, force, reading convention, and work-instruction interpretation. Use reference parts during alignment and retraining.
Inspect the particular parts where operator lines separate or cross. Look for geometry, access, burrs, flexibility, surface finish, or fixturing that changes the technique.
Determine whether measurement noise is excessive or the chosen parts cover too little of the decision range. Do not widen the part sample merely to improve a metric.
Resolve the special-cause cell before treating the overall repeatability estimate as representative. Record the observation and repeat the study after a defined correction.
After the method is changed, repeat the study under the intended operating conditions. A corrected work instruction is not evidence by itself; the repeated study shows whether the improvement reduced the relevant variation component without creating a new problem elsewhere.
Continue from concept to evidence
Use the complete worked resource to see how the same sample moves from 90 measurements through ANOVA, the Gage Evaluation table, all six graphs, and an engineering decision. Use the separate calculation guide when you specifically need the Average & Range formula trail.
For a quick study review, open the free MSA preview. When the work requires crossed ANOVA, Average & Range analysis, the diagnostic report shown above, and printable output in one workflow, continue to Mechatrovich MSA Studio.
References and method sources
- AIAG, Measurement Systems Analysis (MSA), 4th Edition — the automotive-industry reference for evaluating and improving measurement systems.
- NIST/SEMATECH e-Handbook, Gauge R&R Studies — design considerations and production-gage characterization.
- NIST, Analysis of Variability — repeatability, reproducibility, and longer-term variation concepts.
- Minitab Support, Interpret the key results for Crossed Gage R&R Study — definitions and graph-reading guidance for %Study Variation, %Tolerance, and diagnostic charts.