← ALL ARTICLES
BATTLE HARD AFTER 50September 21, 2026· 6 min read

RETEST, DON'T GUESS: TELLING A REAL GAIN FROM MEASUREMENT NOISE

From Battle Hard After 50 - Chapter 13: Retest: Earn Your Next Standard

By Will Power · Marine veteran, four-time Ironman, 300+ races · Last updated September 21, 2026

Not a physician, dietitian, researcher, healthcare provider, lawyer, or financial advisor. How studies are selected and cited

ANSWER

Twelve weeks on, you run the same five-domain battery again. The hard part is reading the result honestly: published test-retest figures show strength measures repeat tightly while power measures repeat loosest, and grip has to move about 3.5 kg before the change can be told apart from the test's own noise.

Twelve weeks after Chapter 2 handed you five domain scores, you run the same battery again. The hard part isn't the retest. It's reading the result honestly — because some of what looks like progress is measurement noise, and some of what looks like failure is a domain that simply moves on a different schedule than the one beside it.

You Get Five Grades, Not One

Chapter 2 set your Recruit / Field Ready / Battle Hard tier in each of five domains, each from one primary test: Strength (the push-up standard, or the chair-stand entry test), Power (the broad jump), Engine (the Rockport walk), Carry and Grip (the grip dynamometer), and Mobility (the chair sit-and-reach). The retest runs the same five, the same way — and they almost certainly won't all move the same amount. One might jump a full tier, one might hold, one might not move at all. That isn't the program failing. It's the same picture Chapter 3 laid out from the start: your five domains age, and respond to training, on five separate schedules.

Worth saying plainly: the twelve-week retest interval is a Will Power Protocols program standard, matched to this book's phase structure. It is not a claim that twelve weeks is an established scientific checkpoint for every capacity in the book.

A Retest Is Only as Good as the Conditions It Was Run Under

Schaun and colleagues had 43 middle-aged (40–55) and older (over 60) adults — some with mobility limitations, some without — run a battery of strength, power and functional tests twice, four weeks apart, and reported how repeatable each measure was.

Maximal dynamic and isometric strength were the most repeatable measures in that study (coefficient of variation 2.2–7%). Functional tests — sit-to-stand, gait speed, timed up-and-go, stair climb, six-minute walk — came in at CV 4.2–6.8%. Peak power was the loosest of the three at CV 6.6–12.8%, and peak power measured at 80–90% of one-rep max looser still, at 14.4–18.3%.

The authors' own conclusion is the operational one: most of these measures were sufficiently reliable even when the two tests sat a month apart — provided the same equipment and procedures were used. That is the whole argument for the boring instruction. Same dynamometer. Same measured course. Same chair. Same time of day, if you can manage it. Change the equipment and you have changed the measurement, not just the score.

How Much Change Counts as Change

Grip is the domain where this gets concrete, because grip has a published measurement error you can hold against your own number.

Rolsted and colleagues tested 70 adults aged 23 to 88 — five men and five women in each decade from 20 to 80-plus — on two electronic dynamometers, three attempts each. Agreement between the devices was high (ICC 0.98). The number to carry into a retest is the measurement error they reported for a single strongest attempt: a standard error of measurement of 1.64 kg (4.2%) and a minimal detectable change of 3.55 kg (9.0%).

Minimal detectable change is the threshold a result has to clear before it can be told apart from the test's own noise. On that study's figures, a grip score that moved 2 kg is not clearly distinguishable from an identical score measured twice. One that moved 5 kg is. Two limits worth keeping attached to that number: it came from a mixed-age sample spanning six decades rather than adults over 50 specifically, and it was measured on new, factory-calibrated devices — a well-used gym dynamometer is not guaranteed to be that tight.

Flat Is Information, Not a Verdict

The reflex when a number doesn't move is to conclude you didn't train hard enough. Before that, audit the boring explanations: was the retest run under the same conditions as the first? Has recovery genuinely been a problem lately? Was this specific domain actually trained consistently, or did life get in the way of exactly this one? Could an off day, a minor illness, or ordinary test-to-test variability account for some of it?

There is also evidence that the size of a training response varies between people for reasons that aren't effort. Soendenbroe and colleagues ran a secondary analysis of 58 healthy men, average age 72, randomized to thrice-weekly heavy resistance training or continued sedentary living. In the training group at 16 weeks, maximal voluntary contraction strength rose 19 ± 14%, rate of force development 58 ± 80%, quadriceps cross-sectional area 3 ± 4%, and type II fibre cross-sectional area 14 ± 25%.

Read the standard deviations, not just the averages. Same men, same program, four outcomes moving by very different amounts and with very different spread between individuals. When the authors classified individual responses against the typical error of each measure, 82% of the training group came out Robust or Excellent responders and 5% came out Poor — and notably, the variability did not track with training compliance or 1RM progression. Lower baseline values were associated with larger improvements but did not fully account for the differences.

Scope that honestly: healthy men averaging 72, older than this book's audience, over 16 weeks rather than 12. What it supports is narrow and useful — the authors concluded that genuine non-responders were rare, which is an argument for running the next block rather than abandoning the approach after one flat number.

A decline is a different signal. Don't answer it by simply adding volume. Reassess first: an unresolved injury, an accumulating recovery deficit, or a real change in circumstances — illness, major stress, disrupted sleep — before deciding the domain needs a harder push rather than a different kind of attention.

What the New Profile Tells You to Do Next

Chapter 12's phase structure runs again from wherever you actually are today — but you allocate attention differently. The domain still sitting at Recruit gets more of the next twelve weeks than the one already comfortable at Battle Hard. That isn't a punishment for the domain that lagged. It's the principle the whole book runs on: train what needs it, not what already feels good to score. Battle Hard isn't a status you earn once — it's a standard you hold by continuing to meet it.

The Protocol

  • Run the full battery before you read your old scores. Write every new number next to the original, not on its own.
  • Hold the conditions constant — same dynamometer, same course, same chair. Schaun's reliability figures assume it.
  • Give each measure its own tolerance. Strength tests repeat tightly; power tests repeat loosest. A 3% shift means something different in each.
  • On grip, treat roughly 3.5 kg as the bar for a change you can distinguish from noise, on that study's numbers.
  • Track the loaded carry separately from your Carry and Grip tier. Grip strength sets the tier; the carry is additional information, not a second vote on the same one.
  • When a number is flat, audit execution, recovery and test conditions together before concluding anything about the training. When one declines, reassess before you add volume.
  • Retest in another twelve weeks. Numbers holding steady once you're Battle Hard across the board is a legitimate outcome, not a stall.

Sources

  1. Schaun GZ, Raidl P, Andrade LS, et al. Examining the test-retest reliability of commonly used neuromuscular, morphological, and functional measures in aging adults. GeroScience. 2025;47(3):4381–4393. PMID 40067538
  2. Rolsted SK, Andersen KD, Dandanell G, et al. Comparison of two electronic dynamometers for measuring handgrip strength. Hand Surgery and Rehabilitation. 2024;43(3):101692. PMID 38705572
  3. Soendenbroe C, Andersen JL, Heisterberg MF, Kjaer M, Mackey AL. Heavy resistance exercise training in older men: A responder and inter-individual variability analysis. PLoS One. 2026;21(1):e0338775. PMID 41563970

KNOW SOMEONE WHO NEEDS THIS?

FREE TOOL

GET YOUR PERSONALIZED PROTOCOL

Answer a few questions and get a training, nutrition, and recovery protocol built for your body, goals, and schedule.

GET YOUR FREE WILL POWER PROTOCOL →

THIS ARTICLE IS FROM

BATTLE HARD AFTER 50 - CHAPTER 13: RETEST: EARN YOUR NEXT STANDARD

Get the full protocol on Amazon — Kindle and paperback.

GET THE BOOK →

Medical disclaimer. This article is for educational purposes only and is not medical advice. These statements have not been evaluated by the Food and Drug Administration, and nothing on this site is intended to diagnose, treat, cure, or prevent any disease. I share published research as a health enthusiast and endurance athlete, not as a clinician — I do not interpret your results and I do not diagnose. Consult your physician before making changes to your supplement, training, or nutrition regimen, especially if you take prescription medication or have an existing health condition.

THE PROTOCOL NEWSLETTER

BATTLE HARD. IN YOUR INBOX.

An email when a new research breakdown or book is published. No set schedule, so you only hear from us when there is something worth reading. No fluff, no spam.

JOIN THE LIST →

Free. Unsubscribe anytime.

MORE ARTICLES