FTP without a test

Durability: why your power fades late in a long ride even when your FTP is fine

Your FTP is 280 and you can hold it in a 20-minute effort on a Tuesday. Four hours into a Saturday group ride the same wattage feels like 330. That gap has a name — durability — and since 2021 it has become the most-hyped number in cycling analytics [Maunder et al. 2021]. It is real, it is measurable from ride files you already own, and almost everything sold around it runs well ahead of the evidence. Here is what the research establishes and what it does not.

By Jim Camut · Former pro & ex-Bruyneel Academy racer

Updated Jul 21, 20264 chapters7 citations

01 / 04

What durability actually means and where the term came from

Ed Maunder and colleagues coined the term in a 2021 Sports Medicine paper, defining durability as the time of onset and magnitude of deterioration in physiological-profiling characteristics during prolonged exercise [Maunder et al. 2021]. The definition is stated in time terms. It specifies no kilojoule criterion at all — that part was added later, by other people.

The observation behind the term is mundane once stated. Everything in a standard physiological profile — VO2max, lactate threshold, gross efficiency, critical power — is measured on a fresh rider in a rested state. Maunder and colleagues argued those characteristics are not static, and that a profile ignoring their deterioration describes an athlete who no longer exists after hour two [Maunder et al. 2021]. Valenzuela and colleagues put a number on it: twelve male professional cyclists lost a mean 2.9% of time-trial power after roughly four hours of submaximal riding [Valenzuela et al. 2023].

The mean hides the interesting part. Individual responses in that same group ranged from an 8.5% loss to a 1.1% gain [Valenzuela et al. 2023]. One rider was measurably worse late; another was fractionally better. Both had elite fresh numbers. That spread is the entire case for measuring durability at all — if every trained rider decayed by the same 3%, the metric would be a constant and you could safely ignore it.

It also explains a mismatch most self-coached riders feel before they have a word for it. A single FTP describes a rider at minute zero. The power-duration curve our ftp-without-a-test pillar shows you how to fit from ride history you already have — the estimate that needs no formal test — is the curve you carry into a ride, not the one you carry out of it. Hour four has its own curve, and it sits lower.

02 / 04

Why the kJ number your app uses is arbitrary

Every app reporting durability picks a work threshold — 1000 kJ, 2000 kJ, 40 kJ/kg — and calls everything past it fatigued. A systematic review catalogued the thresholds used across the literature and concluded that kilojoules alone are insufficient, because the metric ignores the intensity at which that work was done [Sanchez-Jimenez et al. 2025]. The number is an investigator's choice, not a physiological boundary.

List them and the arbitrariness is hard to miss. Published work has used 2.5, 5 and 7.5 kJ/kg; 15, 25, 35 and 45 kJ/kg; a 0 to 50 kJ/kg sweep; roughly 40 kJ/kg; and a flat absolute 2000 kJ [Sanchez-Jimenez et al. 2025]. The most-quoted of them, 40 kJ/kg, traces back to a study with exactly two conditions: 0 and 40 [Valenzuela et al. 2023]. That study was not designed to locate a threshold and could not have, because you cannot find a breakpoint from two points.

When the range is swept, the decline starts earlier. Mateo-March and colleagues tracked mean-maximal power across accumulated work from 0 to 40 kJ/kg in professional cyclists and found progressive decline detectable after 20 kJ/kg [Mateo-March et al. 2025]. The threshold in wide circulation is roughly double the point at which the effect becomes measurable, which means a durability score gated at 40 kJ/kg quietly discards the first half of the decline it claims to describe.

The deeper problem is that total work is the wrong axis. The same review found efforts above critical power produced 10 to 20% power declines at only 2.5 to 15 kJ/kg, while similar or greater volumes at lower intensities produced smaller decrements [Sanchez-Jimenez et al. 2025]. A 2000 kJ steady endurance ride and a 2000 kJ road race leave you in completely different states. Our own gate is 20 kJ/kg of body mass — closer to Mateo-March than the 40 kJ/kg default, still a field heuristic, and still blind to how those kilojoules were earned.

03 / 04

Is durability independent of FTP? The honest answer

The marketing claim is that durability is a fourth axis beside FTP, VO2max and anaerobic capacity. The evidence is split. One study found change in critical power correlated 0.891 with relative peak power, 0.835 with relative VO2max and 0.869 with gross efficiency at 300 W [Spragg physiology 2023]. Another, in a similar cohort, found no association with any laboratory endurance measure [Valenzuela et al. 2023].

That evidence comes from ten under-23 professionals. The headline was an R-squared of 0.96 to 0.98 for predicting durability from physiological measures, but it came from backwards stepwise selection fitting three predictors with six residual degrees of freedom and no out-of-sample validation [Spragg physiology 2023]. A model at that ratio fits nearly anything you hand it. And it does not replicate. Valenzuela and colleagues asked the same question of twelve professionals and found no significant association between power decay and ventilatory threshold, peak power output or VO2max [Valenzuela et al. 2023]. Two small studies, opposite answers — that is what an open question looks like, not a settled one.

The study usually cited as proof durability wins races is more careful than its reputation. Van Erp and colleagues examined 26 male professionals across 85 rider-seasons and found a smaller decline in power after accumulated work was associated with season-level competitive success, with the discriminating durations differing by specialization — 20-minute and 5-minute power for climbers, 10-second and 1-minute for sprinters [Van Erp et al. 2021]. The design was retrospective, exposure and outcome came from the same season, success was a ranking-points proxy, and the study never measured FTP or critical power at all. An association observed within a season, in a sample where nobody's threshold was measured, cannot establish prediction, and it cannot establish independence from a variable that was never collected.

None of this makes durability fake. It makes its independence an open question rather than the settled fourth axis it is sold as. AdaptCycling computes a durability index — the median ratio of your best late-ride 20-minute power to a matched fresh 20-minute effort earlier in the same ride, across long rides in the past year — and reports it as descriptive, never as a verdict. One limit worth stating plainly: we score everyone on the 20-minute and 5-minute windows Van Erp found discriminating for climbers, so a sprinter or crit racer is being graded on the wrong pair.

04 / 04

What we know about training it — and what to do in the meantime

No controlled training trial in trained cyclists has measured a change in durability. The nearest work is uncontrolled dose-response, and the one randomised trial to move a durability measure ran in runners. That is no evidence of a trainability effect size, not evidence that durability cannot be trained. Every durability protocol sold to cyclists rests on observational data.

The strongest observational signal comes from Spragg and colleagues tracking 30 professional cyclists across a competitive season: time spent below the first ventilatory threshold correlated with improvement in fatigued 2-minute power at r = 0.43 (p = 0.018) [Spragg training 2023]. That is the empirical basis for the standard advice to ride more easy volume. The caveats are load-bearing. The analysis was observational and uncorrected for multiple comparisons, and both the training exposure and the fatigued-power outcome were extracted from the same pooled training and race files — riders accumulating more easy hours were also riders racing and training more overall.

We act ahead of that evidence, and it is worth saying so. When a learned durability index lands below 88% of the matched fresh effort, that deficit can pull a fatigue-resistance long ride into a build week — quality efforts placed after an aerobic pre-load rather than at the fresh start of the session. The reasoning is specificity: if you fade at hour three, practice the hard part at hour three. That is a defensible coaching argument. It is not a proven dose-response, and no controlled trial has shown that session moves the number.

In the meantime, the intensity finding is the more actionable one. If work above critical power costs 10 to 20% at 2.5 to 15 kJ/kg while steady riding at the same or greater volume costs measurably less [Sanchez-Jimenez et al. 2025], then how you rode the first 90 minutes may matter more for your hour-four watts than how far you went. The rider who contests every roller in the opening hour of a five-hour ride has spent their durability before the ride is a third done. Audit pacing and fueling discipline before buying a training block aimed at a score nobody has yet shown is trainable.

Common questions

Quick answers

Why does my power fade at hour three when my FTP is fine?

Because FTP describes you fresh. Professional cyclists lost a mean 2.9% of time-trial power after roughly four hours of submaximal riding, with individual responses ranging from an 8.5% loss to a 1.1% gain [Valenzuela et al. 2023]. Some fade is normal physiology. An unusually large fade is worth measuring against your own ride history rather than a population average.

How many kilojoules before durability starts to matter?

Nobody knows. The thresholds apps use vary widely across the literature, and a systematic review concluded kilojoules alone are insufficient because they ignore intensity [Sanchez-Jimenez et al. 2025]. The widely quoted 40 kJ/kg comes from a study with only two conditions, 0 and 40 [Valenzuela et al. 2023]. The one study that actually swept the range found progressive decline detectable after 20 kJ/kg [Mateo-March et al. 2025].

Is durability a fourth training axis alongside FTP and VO2max?

Unresolved. One study of ten under-23 professionals found change in critical power correlated 0.891 with relative peak power and 0.835 with relative VO2max [Spragg physiology 2023]. Another found no association between power decay and any laboratory endurance measure [Valenzuela et al. 2023]. Two small samples, opposite results. Treat a durability score as informative, not as an established independent axis.

Can I train durability with long slow rides?

Possibly. One season-long observational study found time below the first ventilatory threshold correlated with improvement in fatigued 2-minute power at r = 0.43 [Spragg training 2023]. No controlled trial in trained cyclists has measured a change in durability, so this is no evidence of an effect size rather than evidence of no effect.
References

Sources cited in this guide

  1. 01
  2. 02
  3. 03
    Spragg physiology 2023. The Relationship between Physiological Characteristics and Durability in Male Professional Cyclists. Medicine & Science in Sports & Exercise.
  4. 04
  5. 05
  6. 06
    Mateo-March et al. 2025. Reliability of the durability concept in professional cyclists: a field-based study. International Journal of Sports Medicine.
  7. 07
    Valenzuela et al. 2023. Durability in Professional Cyclists: A Field Study. International Journal of Sports Physiology and Performance.
In this series

More inside FTP without a test

Start here · Foundational guide

FTP without a test: estimating threshold from real rides

How to find FTP without a 20-minute or ramp test — using your power curve, critical-power modeling, and the rides you've already done.

Read the full guide

Other articles in this series

  1. 01

    Indoor vs outdoor FTP: why the numbers differ

    Why your indoor FTP reads lower than outdoor — heat, cooling, motivation, and power-source differences — and whether to keep two numbers.

  2. 02

    How to estimate FTP without a power meter

    Estimating FTP from heart rate, RPE, and Strava when you don't own a power meter — how close you can get and where the method breaks down.

  3. 03

    20-minute vs 8-minute FTP test: which to use

    How the 20-minute and 8-minute FTP tests differ, the multipliers each uses, and which one fits your riding — plus why both are only protocols.

  4. 04

    Why cycling apps show you different FTP numbers

    Strava, Xert, Intervals.icu, and TrainingPeaks can each report a different FTP. Why the estimates diverge and which one to trust.

  5. 05

    Is your ramp test FTP too high? Why it happens

    Ramp tests overestimate FTP for anaerobically-gifted riders and underestimate it for diesels. Why the 75% rule misfires and how to correct it.

  6. 06

    How often should you test your FTP?

    How often to re-test FTP as a self-coached cyclist — twice a year, not every six weeks — and why modeled estimates change the cadence.

  7. 07

    Critical Power vs FTP: which number should you trust?

    CP and FTP track each other across a group but can differ 20-36W for you specifically. Why there is no correction factor, and which to anchor to.

  8. 08

    Can you actually hold your FTP for an hour?

    Published time-to-exhaustion at FTP runs 34 to 51 minutes across studies, never 60, and the individual spread is wide. Why the hour is a bad test.

  9. 09

    FTP stopped improving: are you a non-responder?

    Non-responder labels flip for 45% of people depending only on the statistics used. What a flat FTP usually means instead, and what to check first.

Free training analysis · No card · ~3 minutes

Try the adaptive coach yourself.

Connect Strava and the coach reads your last eight weeks — intensity balance, ramp rate, what's working, what's holding you back — then drafts your training plan around it.

Free 14-day trial, no card.