Skip to content

guide 9 min read

Can You Trust Your Packaging Measurements? A Practical Gage R&R Guide

Plan a 10-part, three-operator Gage R&R study, interpret a computed packaging example, and decide whether dimensional readings support acceptance or capability analysis.

On this page

In the synthetic teaching example used throughout this guide, one inspector records a reusable tray width of 120.36 mm; another records 120.52 mm on the same tray. The upper specification limit is 120.50 mm. Their readings support different acceptance decisions, even though they used the same instrument. All dimensions, measurements and calculated results in this example are illustrative, not industrial evidence.

Before asking the supplier to correct its process, establish whether your inspection process can distinguish a dimensional problem from a measurement problem. A crossed Gage R&R study separates variation within an operator’s repeated readings from differences between operators and differences between parts. It also gives the investigation a direction: change the measurement method, investigate the instrument, or proceed with evidence that supports the intended decision.

Find out what is changing before choosing a remedy

Suppose an inspector removes and repositions the tray, then measures it again. The scatter in those readings describes repeatability under those conditions. It includes the effects of handling and positioning, not just the instrument’s electronics. When different inspectors measure the same trays, their differences contribute to reproducibility. A consistently higher reading from one inspector suggests a different problem from erratic readings by everyone. Meanwhile, dimensional differences that persist across inspectors are the part-to-part variation the measurement process needs to distinguish. These are central questions in NIST’s guidance on gauge-study issues.

Agreement does not establish correctness. If all three inspectors read high relative to an accepted reference value, the process has bias that agreement alone cannot reveal. Calibration compares the instrument with reference standards under specified conditions; it does not demonstrate that inspectors seat a tray correctly or contact the intended surfaces. NIST’s discussion of measurement bias specifically distinguishes calibration conditions from the working environment.

Stability adds the time dimension. A method that agrees this morning may drift as equipment wears or environmental conditions change. Track a suitable stable check item over time rather than treating one short R&R study as evidence of continuing stability. The study below examines a defined measurement process over a short interval; it cannot establish long-term performance or absolute accuracy.

Define one measurement and run a balanced crossed study

For this example, measure the external width between two opposed machined contact lands on a rigid, reusable aluminum transport tray. The specified width is 120.00 ± 0.50 mm, giving a lower limit of 119.50 mm, an upper limit of 120.50 mm and a total tolerance width of 1.00 mm. Allow the trays, fixture and instrument to equilibrate in the measurement area, and record ambient temperature. Assume the trays remain dimensionally stable during the study and the measurement does not damage them.

Choose ten trays, P01 through P10, and three trained operators, A, B and C. Each operator measures every tray twice: 10 × 3 × 2 = 60 observations. The design is crossed because every operator measures every part, and balanced because every part-operator combination has two readings. In a real investigation, choose parts covering the actual or expected process variation and operators representative of routine inspection, rather than only the best inspectors. Minitab’s crossed-study data guidance supports this design and cautions that ten parts provide a limited estimate of process variation.

Make the measurement instruction specific enough to expose departures from the intended method:

  • Datum and positioning: Seat the tray’s flat underside, datum A, on three support pads. Locate its designated end face, datum B, against the fixture stop. Measure across the opposed lands at the marked section, 10 mm above datum A. Keep the caliper beam parallel to A and perpendicular to the contacted lands.
  • Instrument and contact: Use the same identified 0–150 mm digital caliper, with 0.01 mm resolution, throughout. Clean the jaws and contact lands; check zero and a suitable reference near the working size under the existing instrument-control procedure. Close the jaws with the thumbwheel to first full contact, without further squeezing. This manual endpoint does not control force in newtons; record it as a method limitation, not as controlled-force measurement.
  • Recording and reset: Record each displayed value in millimeters to 0.01 mm, without averaging. Open the jaws, remove the tray completely and reseat it for every reading. Include instrument identification, method revision and session conditions in the study record.

The coordinator assigns each operator a fresh randomized part order for each trial and randomizes operator order within each trial. Complete all 30 first-trial readings before starting the second trial. Hide previous readings and other operators’ results; do not tell inspectors which trays are expected to pass. Avoid coaching or changing the method midway through collection. Otherwise, the two trials no longer evaluate the same process.

Download the 60 synthetic readings (CSV) to inspect or reproduce this example. The columns are Part, Operator, Trial and Measurement; all measurements are in millimeters. The rows are grouped for checking, not presented as the collection sequence. Keep the actual run order separately. A 0.01 mm display increment is a recording resolution, not proof of 0.01 mm accuracy.

Do not apply this repeated-part design to a changing specimen

A box compression test, or BCT, taken to failure cannot provide a second measurement on the same unchanged box. The first test has altered or destroyed the specimen. A second compression therefore mixes measurement effects with damage from the first loading. Giving each operator a fresh box creates a different problem: differences between boxes are now mixed with operator differences.

The same reasoning applies when repeated contact permanently dents a surface or the specimen changes between readings. Keeping the same part ID does not keep the measured quantity constant. Stop and choose a study design appropriate to the destructive or changing test, with its assumptions justified before collection. Merely selecting a different analysis option does not establish those assumptions. The tray study here is not a substitute protocol for BCT.

Read the results without confusing three percentages

The 60 readings were independently analyzed using a two-way random-effects ANOVA, an analysis of variance that separates part, operator, part-by-operator interaction and residual variation. An interaction means the difference between operators depends on which part they measure. The calculation assumes representative random factors and independent, approximately normally distributed residual errors with a common variance.

For this example, the prespecified rule pools interaction into residual error when its p-value is at least 0.05. The interaction test gives p = 0.6827, so pooling is used. This does not prove that interaction is absent. Any negative final variance-component estimate would be set to zero; none required truncation after pooling. These are point estimates, not confidence bounds. For consequential borderline decisions, examine confidence intervals or extend the study rather than relying on a point estimate alone.

The study-variation multiplier is 6 throughout. The definitions follow Minitab’s documented Gage R&R variance and evaluation formulas. The table describes individual readings, not averages of repeated readings. Values are rounded for display; calculations retain full precision.

SourceStandard deviation (mm)Variance contribution (%)Study variation (%)Tolerance (%)
Repeatability0.018550.345.8511.13
Reproducibility, operator0.061863.8019.5037.12
Total Gage R&R0.064584.1420.3638.75
Part-to-part0.3106295.8697.91186.37
Total0.31727100.00100.00190.36

Scroll to compare all columns.

Variance contribution divides a component’s variance by total estimated variance. Here, total Gage R&R variance is 0.00417104 mm² and total variance is 0.10065840 mm². Their ratio, multiplied by 100, is 4.14%. Repeatability and reproducibility sum to total Gage R&R; Gage R&R and part-to-part sum to total variance. Do not add the subtotal to its components again.

Percent study variation instead compares standard-deviation-based spreads. Define s_GRR as the estimated Gage R&R standard deviation and s_total as the total standard deviation. The calculation is 100 × (6 × s_GRR) / (6 × s_total). Here it is 100 × 0.387502 / 1.903603 = 20.36%. The multiplier cancels because it is identical in numerator and denominator. This is why 4.14% variance contribution does not mean measurement variation is below a 10% study-variation threshold. Standard deviations, unlike variances, do not add directly.

Percent tolerance compares the measurement spread with the drawing’s allowed width: 100 × (6 × s_GRR) / (USL − LSL). The result is 100 × 0.387502 / 1.00 = 38.75%. Using the ±0.50 mm half-width instead would incorrectly double it. This percentage describes a spread relative to tolerance, not the probability of accepting a bad tray and not the instrument’s bias.

The study-variation denominator comes from these ten parts, not a historical production distribution. Their broad spread makes measurement variation look smaller in relative terms. Adding more extreme parts could improve that percentage without improving the measurement process. Values above 100% in the part-to-part and total rows describe modeled 6s spreads relative to tolerance; they are not defect-rate estimates or a capability assessment.

Let the dominant component determine the next action

Minitab’s public summary of AIAG guidance describes below 10%, 10–30% and above 30% bands for standard-deviation-based Gage R&R percentages, with the middle band depending on the application and other factors. Its corresponding variance-contribution boundaries are 1% and 9%, not 10% and 30%. These are contextual guidelines, not a universal packaging acceptance law. Agree the metric and criterion for the actual decision, including failure consequences and customer requirements, before interpreting the result.

Here, 20.36% study variation might support a conditional discussion, but 38.75% tolerance does not justify confident near-limit acceptance. The calculated overall means are 119.9530 mm for A, 120.0145 mm for B and 120.0770 mm for C. Because everyone measured the same parts equally, C’s 0.1240 mm higher mean than A cannot be explained by a different part mix. It identifies disagreement, but without reference values it does not identify which operator is correct.

Operator variation accounts for about 91.75% of the Gage R&R variance. Observe how each inspector seats the tray, aligns the jaws and decides when contact is complete. Revise the fixture or instruction where that observation reveals ambiguity, then train to the revised method. Buying a finer-resolution instrument is not the first conclusion: even eliminating the estimated repeatability component would leave a 6s operator spread of 0.371 mm, or 37.12% of tolerance.

In another study, dominant repeatability would direct attention toward jaw condition, zero behavior, seating, contact force or instrument suitability. A substantial interaction would justify examining part-specific access and alignment rather than applying one operator correction to every reading. These components locate the investigation; they do not prove a physical cause.

Establish the measurement basis before reporting capability

Do not calculate capability from these 60 readings as though they were 60 independent production parts. There are ten physical trays, deliberately remeasured to investigate inspection. Collect appropriate process data only after establishing the measurement basis; the packaging statistics guide covers the subsequent capability and acceptance-sampling questions.

The PackCalc Statistics Suite includes ANOVA-based Gage R&R using part IDs, operator IDs and measurements, with an optional tolerance width. The CSV column names match its import fields; use 1.00 mm as the total tolerance width for this example. The article’s independently computed table uses a pooled-interaction model and a 6σ study spread. PackCalc currently retains the interaction model, sets negative variance components to zero and uses 5.15σ. Those choices change the numerical results, so do not expect the table to match the tool exactly. Confirm the analysis method and spread multiplier whenever comparing software outputs.

Study design and interpretation remain the user’s responsibility. This study evaluates the specified caliper, fixture, dimension and operator scope; it does not establish adequacy for other instruments or packaging materials.

For this tray example, revise the operator-sensitive method and repeat the study before relying on near-limit readings. Resolve disputed trays through an agreed measurement and disposition procedure, not by remeasuring until a preferred result appears. Proceed when the revised evidence meets the justified decision criterion, with bias checks and ongoing stability monitoring covering what the short R&R study cannot establish.

Try it with your dimensions

Statistics Suite

SPC, Capability, GR&R, and Sampling analysis

Open calculator