Measurement System Analysis · Engineering Guide

Gage R&R Explained: Repeatability, Reproducibility, %GRR and ndc

Gage R&R is not one pass-or-fail percentage. It is a structured way to separate measurement noise from real part differences, find where that noise comes from, and decide whether the data is trustworthy enough for its intended use.

UPDATED 2026-08-01ABOUT 17 MINUTESACTUAL MSA STUDIO GRAPHSCROSSED STUDIES

Suppose ten parts are measured by three operators, three times each. The resulting 90 values do not vary only because the parts are different. They also contain equipment variation, technique variation, positioning effects, environmental effects, and possibly operator-specific responses to particular part features.

A Gage R&R study tries to estimate those sources separately. The value of the study is not the label printed at the top of the report. The value is knowing whether the measurement system can reveal the process signal without adding too much noise—and knowing what to improve when it cannot.

What Gage R&R is actually measuring

Every observed reading combines a value attributable to the part with error introduced by the measurement process. In a simplified crossed study, the measurement error is divided into repeatability and reproducibility. Part-to-part variation is desirable in the study because it provides the signal the gage must distinguish.

WHAT YOU RECORDObserved variationThe spread in all measurements collected during the study.
=
PROCESS SIGNALPart-to-part variationReal differences among the selected parts.
+
MEASUREMENT NOISEGage R&RRepeatability plus reproducibility under the study conditions.

The word gage can be misleading because the study evaluates more than the instrument. The measurement system includes the device, fixture, software, method, operator, part presentation, environment, and the way results are recorded. A calibrated micrometer can still be part of an unsuitable measurement system if the datum is interpreted differently by each operator or the part cannot be located consistently.

The study design comes before the percentage

A crossed design means every operator measures every selected part in every trial. That creates repeated observations for the same operator–part combination and lets the analysis compare within-cell variation, operator differences, part differences, and operator × part interaction.

OPERATORS3
PARTS10
TRIALS3
TOTAL READINGS90

This balanced structure is the sample used in the product graphs below. It illustrates the logic of a crossed study; the correct design for a real study must follow the applicable customer, industry, and internal requirements.

The parts should represent the process range that matters for the decision. Selecting ten nearly identical parts can make a capable gage appear unable to distinguish part-to-part variation. Selecting only extreme parts can make discrimination look artificially strong. The study should also use the real fixture, work instruction, environmental conditions, and operators who normally perform the measurement.

Randomization matters. If an operator measures Part 1 three times consecutively, they can remember the previous value and unintentionally reproduce it. Blinding part identity where practical and mixing the measurement order reduce memory and sequence effects. The study is meant to reproduce the operating measurement process—not a cleaner laboratory version that will never exist in production.

Repeatability and reproducibility answer different questions

Repeatability

Can the same operator obtain similar values on the same part?

Repeatability is the within-operator, within-part variation under the study conditions. It is often called equipment variation, but the observed effect can also include short-term positioning, contact, resolution, fixture, and technique variation.

  • instrument resolution or noise
  • fixture movement or inconsistent seating
  • contact force or probe alignment
  • short-term environmental variation
  • unclear repeated-measurement technique

Reproducibility

Does the measurement response change across operators?

Reproducibility captures differences between appraisers. In an ANOVA study it can include an overall operator effect and an operator × part interaction, depending on the selected model and treatment of the interaction.

  • different datum or edge interpretation
  • systematic force or alignment differences
  • operator-specific setup or zeroing
  • inconsistent work-instruction interpretation
  • part features that affect operators differently

The distinction determines the improvement route. Replacing a gage is unlikely to correct an operator offset created by inconsistent datum selection. Retraining operators is unlikely to correct inadequate instrument resolution. The total GRR value tells you the size of the combined issue; the components tell you where to look.

ANOVA and Average & Range are not interchangeable labels

Two common calculation approaches are the Average & Range method and the analysis of variance method. Both estimate repeatability, reproducibility, Gage R&R, and part-to-part variation, but they do not use the same information or assumptions.

MethodWhat it usesPractical implication
Average & RangeCell ranges, operator averages, part averages, and study constants linked to the design.Transparent for hand calculation and teaching, but it provides a more compressed view of the sources.
Crossed ANOVAMean squares from the part, operator, operator × part, and repeatability terms.Can estimate the interaction explicitly and provides significance tests for model terms.

ANOVA is especially useful when the same part may be measured differently by different operators. That pattern is not just a global operator offset; it is an interaction. One operator might agree on simple cylindrical parts but respond differently when a chamfer, surface, or locating feature complicates the measurement.

Total Gage R&R
σGRR = √(σrepeatability2 + σreproducibility2)
Independent variance components combine by adding variances, not standard deviations.
Total variation
σtotal = √(σGRR2 + σpart2)
This separates measurement noise from the part signal represented by the study.

Method note: MSA Studio’s ANOVA report expresses study variation as 6 × standard deviation. Some worksheets, legacy templates, or calculation guides may use different constants or spread conventions. Compare the method, model, and convention before comparing two reported percentages digit for digit.

The percentage column changes the question

A report can show %Contribution, %Study Variation, and %Tolerance for the same source. These numbers are not alternative spellings of the same result. They use different denominators and should not be compared as if they were interchangeable.

Reported metricSample Total GRRQuestion answered
%ContributionBased on variance0.48%What share of the estimated total variance is attributed to the measurement system?
%Study VariationBased on standard-deviation spread6.93%How large is the GRR study variation relative to the total variation represented in this study?
%ToleranceStudy variation ÷ specification width7.99%How much of the engineering tolerance band is consumed by measurement variation?

Because variance is squared, %Contribution is usually numerically much smaller than %Study Variation. Reporting “GRR is 0.48%” without naming the column would make this sample sound far better than the more decision-relevant 6.93% of study variation or 7.99% of tolerance. Always report the metric name with the value.

TOTAL GRR · %SV6.93%Combined measurement variation
REPEATABILITY · %SV4.15%Within operator–part cells
REPRODUCIBILITY · %SV5.55%Operator and interaction sources
PART-TO-PART · %SV99.76%Dominant study signal

%Study Variation depends on the part variation captured by the sample. %Tolerance depends on the entered specification width. They can therefore support different conclusions. A broad part range can reduce %Study Variation while leaving %Tolerance unchanged. A wide tolerance can make %Tolerance look small even if the system is not discriminating finely enough for process improvement. The intended use of the measurement is what decides which comparison matters most.

ndc describes discrimination—not accuracy

The number of distinct categories estimates how many non-overlapping groups of part variation the measurement system can distinguish within the study. It is derived from the relationship between part-to-part variation and Gage R&R variation, commonly using a factor of 1.41 before truncation to an integer.

Number of distinct categories
ndc = floor(1.41 × σpart ÷ σGRR)
A larger part signal relative to measurement noise produces more distinguishable categories.
20SAMPLE NDC

The sample can distinguish many part levels

This supports the conclusion that measurement noise is small relative to the selected part range. It does not prove that measurements are unbiased, calibrated, stable over time, or correct across the full operating range.

A low ndc may indicate excessive measurement variation, insufficient part spread, or both. That is why “select wider parts” is not automatically a corrective action: it may improve ndc without improving the measurement system. First decide whether the original parts represented the range in which the gage must make useful distinctions.

How the six Gage R&R graphs tell the story

The numerical table ranks the variation components. The graphs show the mechanism behind those values. The panel below uses the same 3-operator × 10-part × 3-trial dataset and the same six diagnostic views generated by Mechatrovich MSA Studio v1.3.1. Hover over points or bars to inspect the underlying values.

Gage R&R Diagnostic Graphs — Diameter

Crossed ANOVA sample · Digital micrometer · 0.001 mm resolution

PRODUCT OUTPUT

Components of Variation

Measurements by Parts

R Chart by Operators

Measurements by Operators

Xbar Chart by Operators

Parts × Operators Interaction

These graphs are diagnostic views, not six separate acceptance tests. Read them together with the study design, variance-component table, ANOVA terms, characteristic risk, and intended measurement decision.

Components of Variation

Start with the relative sizes. In a useful study, part-to-part variation should dominate the measurement components. Compare the same named percentage across sources; do not compare %Contribution for one bar with %Tolerance for another.

Measurements by Parts

Look for clear separation among part centers with relatively tight clusters of repeated readings. Heavy overlap can reflect measurement noise, a narrow part range, or both.

R Chart by Operators

This is the first repeatability check. A point above the upper control limit identifies an operator–part cell whose repeated measurements are unusually inconsistent. Investigate that cell before relying on the pooled repeatability estimate.

Measurements by Operators

Compare the centers and spreads. A consistent shift between operator distributions suggests a systematic technique or setup difference. Similar centers with different spreads point more toward repeatability differences.

Xbar Chart by Operators

Unlike an ordinary process-control chart, many points outside the limits can be desirable here: the limits reflect measurement-system repeatability, so separated part means show that real part differences are visible above the measurement noise.

Parts × Operators Interaction

Parallel operator traces suggest the operator effect is reasonably consistent across parts. Crossing or changing gaps indicate that particular part features may interact with operator technique. The graph helps localize what an interaction p-value only summarizes.

Acceptance bands are a screen, not the engineering decision

Commonly taught guidance treats measurement variation below 10% as generally acceptable, 10% to 30% as potentially acceptable depending on application, and above 30% as generally unacceptable. These bands are useful for screening, but they do not remove the need to understand risk and purpose.

< 10%Generally considered acceptable
10–30%May be acceptable for the application
> 30%Generally considered unacceptable

A characteristic used for safety, regulatory acceptance, tight process adjustment, or automated feedback can need a more demanding system than a broad screening measurement. A 9% headline result also does not cancel evidence of an unstable range, a strong operator offset, a problematic interaction, or poor study execution.

01
Was the study representative?

Confirm the parts, operators, fixture, method, resolution, environment, and measurement order reflect real use.

02
Were the data collected correctly?

Check part IDs, operator IDs, trial sequence, missing readings, rounding, transcription, and any excluded observations.

03
Is repeatability stable?

Investigate out-of-control R-chart cells and unusually large within-cell ranges before accepting the pooled result.

04
Is reproducibility understandable?

Review operator centers, spreads, work methods, and any operator × part pattern.

05
Does the comparison match the decision?

Use study variation, tolerance, or a historical process variation intentionally—not simply whichever gives the smallest percentage.

06
Is discrimination sufficient?

Read ndc together with part selection and the measurement resolution needed for control or classification.

07
Does the risk justify approval?

Document the characteristic importance, customer requirements, safeguards, limitations, and follow-up plan.

What a Gage R&R study does—and does not—prove

It helps estimate

Short-term repeatability, between-operator reproducibility, operator × part interaction in an ANOVA model, part-to-part variation within the selected sample, and measurement discrimination relative to that sample.

It does not automatically prove

Calibration, traceability, bias, linearity, long-term stability, correct specification limits, representative sampling, measurement uncertainty for every use, or suitability for every product and operating range.

Bias asks whether the measurement system is centered on an accepted reference. Linearity asks whether bias changes across the operating range. Stability asks whether the system changes over time. Gage R&R addresses precision components under the study conditions; these other properties need their own evidence.

Turn a poor result into a targeted investigation

“Improve the measurement system” is too broad to be actionable. Use the component and graph pattern to choose the first investigation.

High repeatability

Observe repeated measurements of the worst operator–part cells. Check resolution, seating, contact force, fixture movement, zeroing, surface condition, and environmental sensitivity.

Operator offset

Compare setup, datum selection, alignment, force, reading convention, and work-instruction interpretation. Use reference parts during alignment and retraining.

Operator × part

Inspect the particular parts where operator lines separate or cross. Look for geometry, access, burrs, flexibility, surface finish, or fixturing that changes the technique.

Low ndc only

Determine whether measurement noise is excessive or the chosen parts cover too little of the decision range. Do not widen the part sample merely to improve a metric.

Unstable R chart

Resolve the special-cause cell before treating the overall repeatability estimate as representative. Record the observation and repeat the study after a defined correction.

After the method is changed, repeat the study under the intended operating conditions. A corrected work instruction is not evidence by itself; the repeated study shows whether the improvement reduced the relevant variation component without creating a new problem elsewhere.

Continue from concept to evidence

Use the complete worked resource to see how the same sample moves from 90 measurements through ANOVA, the Gage Evaluation table, all six graphs, and an engineering decision. Use the separate calculation guide when you specifically need the Average & Range formula trail.

For a quick study review, open the free MSA preview. When the work requires crossed ANOVA, Average & Range analysis, the diagnostic report shown above, and printable output in one workflow, continue to Mechatrovich MSA Studio.

References and method sources

  1. AIAG, Measurement Systems Analysis (MSA), 4th Edition — the automotive-industry reference for evaluating and improving measurement systems.
  2. NIST/SEMATECH e-Handbook, Gauge R&R Studies — design considerations and production-gage characterization.
  3. NIST, Analysis of Variability — repeatability, reproducibility, and longer-term variation concepts.
  4. Minitab Support, Interpret the key results for Crossed Gage R&R Study — definitions and graph-reading guidance for %Study Variation, %Tolerance, and diagnostic charts.

Need the complete Gage R&R workflow?

MSA Studio keeps study setup, measurement entry, ANOVA, Average & Range, diagnostic charts, and reporting together.

View MSA Studio →