FTP stopped improving: are you a non-responder?
You have trained consistently for months and the number has not moved. Somewhere online you found the word non-responder, and it read like a verdict. It is not one. When one dataset was run through several accepted statistical methods for classifying responders, only 11 of 20 subjects kept the same label across all of them [Hecksteden et al. 2018]. Real between-person variation in training response almost certainly exists. A defensible way to tell one specific rider they are a non-responder does not.
By Jim Camut · Former pro & ex-Bruyneel Academy racer
Updated Jul 21, 20264 chapters6 citations
What the non-responder literature actually shows
The idea comes from large training studies reporting wide spreads in VO2max gain. The critique is methodological, not physiological: most of those studies ran no control group, so within-subject noise was never separated from true response [Williamson et al. 2017]. That is no evidence of effect, which is a different claim from evidence of no effect.
The canonical source is the HERITAGE Family Study, whose published gain distributions — some participants improving enormously, some barely at all — became the origin of the responder and non-responder vocabulary. Williamson and colleagues searched more than 180 HERITAGE publications and could not find a comparator arm in any of them [Williamson et al. 2017]. Without a control condition, a spread of observed changes cannot be decomposed into real individual response and ordinary test-retest variation. Their conclusion is blunt: true inter-individual differences in response cannot be quantified, let alone appraised for clinical relevance.
The people who built HERITAGE agree with the design point. The 2019 precision exercise medicine consensus statement in the British Journal of Sports Medicine, co-authored by Claude Bouchard and James Skinner, states that designs without a control group cannot isolate changes due to treatment from changes that would have occurred without it [Ross et al. 2019]. That is the field conceding the methodology, in print, under the names of the investigators who generated the original data.
The same paper carries the other half, and honesty requires carrying it too. Monozygotic twin resemblance and familial aggregation in the response data show that response variance is not randomly distributed — a familial component accounts for roughly 30% to 60% of it depending on the trait [Ross et al. 2019]. A genuine individual-response component is therefore almost certainly non-zero. But that finding is an association measured across a population. It is not a prediction about any one person, and no published method converts it into one.
Why the label flips depending on who does the maths
Hecksteden and colleagues took a single training dataset and applied several accepted analytical approaches for classifying responders. Only 11 of 20 subjects were consistently classified [Hecksteden et al. 2018]. Nine of twenty changed label on nothing but the statistic chosen. A classification that unstable cannot carry a verdict about a person.
The detail that matters is that the data never changed. One year-long training study, one set of VO2max measurements, several defensible ways to draw the responder line — and the labels moved. The authors describe the disagreement between approaches as remarkable [Hecksteden et al. 2018]. If a rider is a responder or a non-responder depending on which reviewer's preferred method gets applied to identical numbers, the label is describing the analyst, not the athlete.
Measurement error is the reason. In a pooled analysis of three randomised trials — 251 intervention participants against 87 controls — only 45% of participants exceeded twice the technical error of measurement for absolute VO2peak [Brennan et al. 2022]. In supervised, funded, laboratory-measured training, most participants did not produce a change large enough to separate from the equipment and the day — even though the group mean improved significantly. Brennan calls those participants uncertain, not unresponsive. Undetectable and absent are different things, and most published training data cannot tell them apart.
Hecksteden's proposed remedy is worth knowing for what it implies. The fix is not more subjects; it is repeated testing during the training phase, so each individual carries their own error estimate [Hecksteden et al. 2018]. It was published as a proof of concept, and where others have applied it the picture did not improve: measuring fitness five times across 24 weeks, Bonafiglia and colleagues could confidently classify only 55 of 109 participants as responders [Bonafiglia et al. 2019]. No protocol specifies how many repeat tests one person needs for a stated misclassification rate, which means nobody currently holds the evidence that would settle it.
The four things far more likely than being a non-responder
Before reaching for genetics, check four things: the two FTP numbers came from different protocols, the stimulus never actually changed, fatigue was never resolved, or the window was too short. Each is more common than true non-response, and each is fixable. A January ramp test against a June 20-minute test is not a comparison.
Measurement first, because it is the cheapest to rule out and the most often guilty. Warm-up structure alone shifts the FTP a 20-minute test reports [Tramontin et al. 2022], and the ramp test's assumption that FTP equals 75% of peak one-minute power is a protocol convention rather than a constant — the true ratio varies rider to rider. Two tests, two protocols, two states of readiness, and most of the difference between them is method. This is exactly the problem our ftp-without-a-test pillar addresses: an FTP modelled from ride history you already have is anchored to dozens of efforts across months rather than two isolated days, so a flat trend line means closer to what you think it means.
Second, the stimulus. Eight months of the same sweet-spot template at the same three weekly durations is not eight months of progressive overload; it is one month repeated eight times. Adaptation follows a change in demand. If volume, session structure, and intensity distribution have all been stable since January, the honest reading is that the plan stopped asking for anything new — not that the body stopped answering.
Third, unresolved fatigue. A threshold estimate taken while carrying accumulated load reads low, and a rider stacking hard weeks on five hours of sleep can hold a suppressed number for a full season without ever seeing the fitness underneath it. The diagnostic is cheap: take a genuine easy week, then re-measure. If the number jumps, the problem was recovery, not response.
Fourth, the window. A change smaller than roughly twice the measurement error of your protocol cannot be told apart from noise [Brennan et al. 2022], and for most amateur field tests that band is wide enough to swallow a real gain. Six weeks is a training block, not an evaluation period. Twelve or more weeks, measured the same way each time, is the minimum before flat means anything — and switching apps or protocols mid-window resets the clock.
What an honest answer looks like
No protocol establishes how many repeat tests are needed to classify one person as a responder at a stated error rate. Until that exists, no coach, laboratory, or app — AdaptCycling included — can tell you that you are a non-responder. What can honestly be said is narrower: this stimulus, measured this way, did not move this number.
The narrower statement is also the more useful one, because every term in it is something you control. The stimulus can change. The measurement method can be held constant. The window can be extended. The genetic verdict, by contrast, is unfalsifiable in practice — there is no test you can run on yourself that returns it, which is precisely why it is a comfortable thing to believe and a poor thing to act on.
Response is also trait-specific, which makes single-number verdicts worse than they look. In the same pooled trial analysis, the share of participants exceeding twice the technical error ranged from 37% for glucose disposal to between 51% and 77% for body-composition measures [Brennan et al. 2022]. Same people, different traits, very different response rates. A rider whose FTP is flat may have added durability late in long rides, repeatability across intervals, or five-minute power, and would still be labelled a non-responder by an assessment that only ever looked at threshold.
AdaptCycling does not classify riders as responders or non-responders, and will not until the literature supports it. What we do is remove two of the four confounders above by default: the FTP estimate is modelled continuously from your Strava history rather than sampled from occasional tests, which holds the measurement method constant, and the plan rotates its stimulus on a periodized cadence rather than repeating one block indefinitely. It does not detect that your numbers have stalled — that read is still yours to make. That is not a genetic answer. It is the part of the question that is answerable.
Quick answers
Am I a non-responder to training?
What percentage of people are non-responders?
Does genetics affect how much I improve?
What should I change first if my FTP will not move?
Sources cited in this guide
- 01Hecksteden et al. 2018. Repeated testing for the assessment of individual response to exercise training. Journal of Applied Physiology.
- 02Williamson et al. 2017. Inter-Individual Responses of Maximal Oxygen Uptake to Exercise Training: A Critical Review. Sports Medicine.
- 03Ross et al. 2019. Precision exercise medicine: understanding exercise response variability. British Journal of Sports Medicine.
- 04Brennan et al. 2022. Toward Personalized Exercise Medicine: A Cautionary Tale. Medicine & Science in Sports & Exercise.
- 05Bonafiglia et al. 2019. The application of repeated testing and monoexponential regressions to classify individual cardiorespiratory fitness responses to exercise training. European Journal of Applied Physiology.
- 06Tramontin et al. 2022. Functional Threshold Power Estimated from a 20-minute Time-trial Test is Warm-up-dependent. International Journal of Sports Medicine.
More inside FTP without a test
Start here · Foundational guide
FTP without a test: estimating threshold from real rides
How to find FTP without a 20-minute or ramp test — using your power curve, critical-power modeling, and the rides you've already done.
Read the full guide
Other articles in this series
- 01
Indoor vs outdoor FTP: why the numbers differ
Why your indoor FTP reads lower than outdoor — heat, cooling, motivation, and power-source differences — and whether to keep two numbers.
- 02
How to estimate FTP without a power meter
Estimating FTP from heart rate, RPE, and Strava when you don't own a power meter — how close you can get and where the method breaks down.
- 03
20-minute vs 8-minute FTP test: which to use
How the 20-minute and 8-minute FTP tests differ, the multipliers each uses, and which one fits your riding — plus why both are only protocols.
- 04
Why cycling apps show you different FTP numbers
Strava, Xert, Intervals.icu, and TrainingPeaks can each report a different FTP. Why the estimates diverge and which one to trust.
- 05
Is your ramp test FTP too high? Why it happens
Ramp tests overestimate FTP for anaerobically-gifted riders and underestimate it for diesels. Why the 75% rule misfires and how to correct it.
- 06
How often should you test your FTP?
How often to re-test FTP as a self-coached cyclist — twice a year, not every six weeks — and why modeled estimates change the cadence.
- 07
Critical Power vs FTP: which number should you trust?
CP and FTP track each other across a group but can differ 20-36W for you specifically. Why there is no correction factor, and which to anchor to.
- 08
Can you actually hold your FTP for an hour?
Published time-to-exhaustion at FTP runs 34 to 51 minutes across studies, never 60, and the individual spread is wide. Why the hour is a bad test.
- 09
Durability: why your power fades late in long rides
What durability is, why the kJ threshold your app uses is arbitrary, and the honest evidence on whether it is independent of FTP or trainable.
Free training analysis · No card · ~3 minutes
Try the adaptive coach yourself.
Connect Strava and the coach reads your last eight weeks — intensity balance, ramp rate, what's working, what's holding you back — then drafts your training plan around it.
Free 14-day trial, no card.