FTP without a test

FTP stopped improving: are you a non-responder?

You have trained consistently for months and the number has not moved. Somewhere online you found the word non-responder, and it read like a verdict. It is not one. When one dataset was run through several accepted statistical methods for classifying responders, only 11 of 20 subjects kept the same label across all of them [Hecksteden et al. 2018]. Real between-person variation in training response almost certainly exists. A defensible way to tell one specific rider they are a non-responder does not.

By Jim Camut · Former pro & ex-Bruyneel Academy racer

Updated Jul 21, 20264 chapters6 citations

01 / 04

What the non-responder literature actually shows

The idea comes from large training studies reporting wide spreads in VO2max gain. The critique is methodological, not physiological: most of those studies ran no control group, so within-subject noise was never separated from true response [Williamson et al. 2017]. That is no evidence of effect, which is a different claim from evidence of no effect.

The canonical source is the HERITAGE Family Study, whose published gain distributions — some participants improving enormously, some barely at all — became the origin of the responder and non-responder vocabulary. Williamson and colleagues searched more than 180 HERITAGE publications and could not find a comparator arm in any of them [Williamson et al. 2017]. Without a control condition, a spread of observed changes cannot be decomposed into real individual response and ordinary test-retest variation. Their conclusion is blunt: true inter-individual differences in response cannot be quantified, let alone appraised for clinical relevance.

The people who built HERITAGE agree with the design point. The 2019 precision exercise medicine consensus statement in the British Journal of Sports Medicine, co-authored by Claude Bouchard and James Skinner, states that designs without a control group cannot isolate changes due to treatment from changes that would have occurred without it [Ross et al. 2019]. That is the field conceding the methodology, in print, under the names of the investigators who generated the original data.

The same paper carries the other half, and honesty requires carrying it too. Monozygotic twin resemblance and familial aggregation in the response data show that response variance is not randomly distributed — a familial component accounts for roughly 30% to 60% of it depending on the trait [Ross et al. 2019]. A genuine individual-response component is therefore almost certainly non-zero. But that finding is an association measured across a population. It is not a prediction about any one person, and no published method converts it into one.

02 / 04

Why the label flips depending on who does the maths

Hecksteden and colleagues took a single training dataset and applied several accepted analytical approaches for classifying responders. Only 11 of 20 subjects were consistently classified [Hecksteden et al. 2018]. Nine of twenty changed label on nothing but the statistic chosen. A classification that unstable cannot carry a verdict about a person.

The detail that matters is that the data never changed. One year-long training study, one set of VO2max measurements, several defensible ways to draw the responder line — and the labels moved. The authors describe the disagreement between approaches as remarkable [Hecksteden et al. 2018]. If a rider is a responder or a non-responder depending on which reviewer's preferred method gets applied to identical numbers, the label is describing the analyst, not the athlete.

Measurement error is the reason. In a pooled analysis of three randomised trials — 251 intervention participants against 87 controls — only 45% of participants exceeded twice the technical error of measurement for absolute VO2peak [Brennan et al. 2022]. In supervised, funded, laboratory-measured training, most participants did not produce a change large enough to separate from the equipment and the day — even though the group mean improved significantly. Brennan calls those participants uncertain, not unresponsive. Undetectable and absent are different things, and most published training data cannot tell them apart.

Hecksteden's proposed remedy is worth knowing for what it implies. The fix is not more subjects; it is repeated testing during the training phase, so each individual carries their own error estimate [Hecksteden et al. 2018]. It was published as a proof of concept, and where others have applied it the picture did not improve: measuring fitness five times across 24 weeks, Bonafiglia and colleagues could confidently classify only 55 of 109 participants as responders [Bonafiglia et al. 2019]. No protocol specifies how many repeat tests one person needs for a stated misclassification rate, which means nobody currently holds the evidence that would settle it.

03 / 04

The four things far more likely than being a non-responder

Before reaching for genetics, check four things: the two FTP numbers came from different protocols, the stimulus never actually changed, fatigue was never resolved, or the window was too short. Each is more common than true non-response, and each is fixable. A January ramp test against a June 20-minute test is not a comparison.

Measurement first, because it is the cheapest to rule out and the most often guilty. Warm-up structure alone shifts the FTP a 20-minute test reports [Tramontin et al. 2022], and the ramp test's assumption that FTP equals 75% of peak one-minute power is a protocol convention rather than a constant — the true ratio varies rider to rider. Two tests, two protocols, two states of readiness, and most of the difference between them is method. This is exactly the problem our ftp-without-a-test pillar addresses: an FTP modelled from ride history you already have is anchored to dozens of efforts across months rather than two isolated days, so a flat trend line means closer to what you think it means.

Second, the stimulus. Eight months of the same sweet-spot template at the same three weekly durations is not eight months of progressive overload; it is one month repeated eight times. Adaptation follows a change in demand. If volume, session structure, and intensity distribution have all been stable since January, the honest reading is that the plan stopped asking for anything new — not that the body stopped answering.

Third, unresolved fatigue. A threshold estimate taken while carrying accumulated load reads low, and a rider stacking hard weeks on five hours of sleep can hold a suppressed number for a full season without ever seeing the fitness underneath it. The diagnostic is cheap: take a genuine easy week, then re-measure. If the number jumps, the problem was recovery, not response.

Fourth, the window. A change smaller than roughly twice the measurement error of your protocol cannot be told apart from noise [Brennan et al. 2022], and for most amateur field tests that band is wide enough to swallow a real gain. Six weeks is a training block, not an evaluation period. Twelve or more weeks, measured the same way each time, is the minimum before flat means anything — and switching apps or protocols mid-window resets the clock.

04 / 04

What an honest answer looks like

No protocol establishes how many repeat tests are needed to classify one person as a responder at a stated error rate. Until that exists, no coach, laboratory, or app — AdaptCycling included — can tell you that you are a non-responder. What can honestly be said is narrower: this stimulus, measured this way, did not move this number.

The narrower statement is also the more useful one, because every term in it is something you control. The stimulus can change. The measurement method can be held constant. The window can be extended. The genetic verdict, by contrast, is unfalsifiable in practice — there is no test you can run on yourself that returns it, which is precisely why it is a comfortable thing to believe and a poor thing to act on.

Response is also trait-specific, which makes single-number verdicts worse than they look. In the same pooled trial analysis, the share of participants exceeding twice the technical error ranged from 37% for glucose disposal to between 51% and 77% for body-composition measures [Brennan et al. 2022]. Same people, different traits, very different response rates. A rider whose FTP is flat may have added durability late in long rides, repeatability across intervals, or five-minute power, and would still be labelled a non-responder by an assessment that only ever looked at threshold.

AdaptCycling does not classify riders as responders or non-responders, and will not until the literature supports it. What we do is remove two of the four confounders above by default: the FTP estimate is modelled continuously from your Strava history rather than sampled from occasional tests, which holds the measurement method constant, and the plan rotates its stimulus on a periodized cadence rather than repeating one block indefinitely. It does not detect that your numbers have stalled — that read is still yours to make. That is not a genetic answer. It is the part of the question that is answerable.

Common questions

Quick answers

Am I a non-responder to training?

There is currently no defensible way for anyone to answer that about you specifically. Classification of individuals is unstable — in one dataset run through several accepted analytical approaches, only 11 of 20 subjects were consistently classified [Hecksteden et al. 2018] — and no protocol establishes how many repeat tests would settle it for one person at a stated error rate. A flat FTP is far more likely to be a measurement, stimulus, or recovery problem.

What percentage of people are non-responders?

Every figure you will see quoted is method-dependent, which is the whole problem. Nine of twenty subjects in one analysis changed classification based purely on which accepted statistic was applied to the same data [Hecksteden et al. 2018]. Separately, only 45% of participants in a pooled trial analysis exceeded twice the technical error for absolute VO2peak [Brennan et al. 2022] — but that measures whether a change was detectable, not whether a person is incapable of adapting.

Does genetics affect how much I improve?

Almost certainly yes, at the population level. Twin and family data show response variance is not randomly distributed, with a familial component accounting for roughly 30% to 60% of it depending on the trait [Ross et al. 2019]. That is an association across groups, not a prediction for an individual. Knowing a familial component contributes to the spread tells you nothing about where you personally sit within it.

What should I change first if my FTP will not move?

Change one variable and hold the rest still. Start with a real recovery week and re-measure, because a suppressed number from unresolved fatigue is the fastest thing to rule out. If it does not move, change the stimulus rather than the volume of the same stimulus — different session structure, different intensity distribution — and give it twelve weeks measured the same way before judging.
References

Sources cited in this guide

  1. 01
    Hecksteden et al. 2018. Repeated testing for the assessment of individual response to exercise training. Journal of Applied Physiology.
  2. 02
  3. 03
    Ross et al. 2019. Precision exercise medicine: understanding exercise response variability. British Journal of Sports Medicine.
  4. 04
    Brennan et al. 2022. Toward Personalized Exercise Medicine: A Cautionary Tale. Medicine & Science in Sports & Exercise.
  5. 05
  6. 06
    Tramontin et al. 2022. Functional Threshold Power Estimated from a 20-minute Time-trial Test is Warm-up-dependent. International Journal of Sports Medicine.
In this series

More inside FTP without a test

Start here · Foundational guide

FTP without a test: estimating threshold from real rides

How to find FTP without a 20-minute or ramp test — using your power curve, critical-power modeling, and the rides you've already done.

Read the full guide

Other articles in this series

  1. 01

    Indoor vs outdoor FTP: why the numbers differ

    Why your indoor FTP reads lower than outdoor — heat, cooling, motivation, and power-source differences — and whether to keep two numbers.

  2. 02

    How to estimate FTP without a power meter

    Estimating FTP from heart rate, RPE, and Strava when you don't own a power meter — how close you can get and where the method breaks down.

  3. 03

    20-minute vs 8-minute FTP test: which to use

    How the 20-minute and 8-minute FTP tests differ, the multipliers each uses, and which one fits your riding — plus why both are only protocols.

  4. 04

    Why cycling apps show you different FTP numbers

    Strava, Xert, Intervals.icu, and TrainingPeaks can each report a different FTP. Why the estimates diverge and which one to trust.

  5. 05

    Is your ramp test FTP too high? Why it happens

    Ramp tests overestimate FTP for anaerobically-gifted riders and underestimate it for diesels. Why the 75% rule misfires and how to correct it.

  6. 06

    How often should you test your FTP?

    How often to re-test FTP as a self-coached cyclist — twice a year, not every six weeks — and why modeled estimates change the cadence.

  7. 07

    Critical Power vs FTP: which number should you trust?

    CP and FTP track each other across a group but can differ 20-36W for you specifically. Why there is no correction factor, and which to anchor to.

  8. 08

    Can you actually hold your FTP for an hour?

    Published time-to-exhaustion at FTP runs 34 to 51 minutes across studies, never 60, and the individual spread is wide. Why the hour is a bad test.

  9. 09

    Durability: why your power fades late in long rides

    What durability is, why the kJ threshold your app uses is arbitrary, and the honest evidence on whether it is independent of FTP or trainable.

Free training analysis · No card · ~3 minutes

Try the adaptive coach yourself.

Connect Strava and the coach reads your last eight weeks — intensity balance, ramp rate, what's working, what's holding you back — then drafts your training plan around it.

Free 14-day trial, no card.