FTP without a test

Critical Power vs FTP: why your two numbers disagree, and which one should anchor your training

Your software shows you two numbers: a Critical Power fitted from your power curve, and an FTP from a 20-minute test. They are not the same construct, and for you specifically the gap can run up to about 36 watts in either direction. Across three studies of trained cyclists the average gap was +7 W, +16 W, and -3 W. Small, and inconsistent even in sign. That is exactly why no correction factor exists. Here is what the agreement data shows, and which number should anchor your prescriptions.

By Jim Camut · Former pro & ex-Bruyneel Academy racer

Updated Jul 21, 20264 chapters7 citations

01 / 04

CP and FTP are not the same construct

Critical Power is the asymptote of your power-duration curve — a modelled boundary above which the physiological response stops being steady-state [Jones & Vanhatalo 2017]. FTP is an operational convention: roughly the power you could hold for an hour, usually estimated as 95% of a 20-minute effort. One is a fitted parameter. The other is prescription shorthand.

The distinction is not academic hair-splitting, because the two have very different evidence behind them. Jamnick and colleagues reviewed the methods used to mark the boundary between the heavy and severe intensity domains and concluded there is little evidence supporting the validity of most commonly used approaches — with critical power and critical speed named as the exceptions [Jamnick et al. 2020]. FTP is not among the validated markers. It is a field convention that has proven extremely useful for organising training, which is a different claim from being a measured physiological threshold.

FTP is also not one-hour power, despite the definition. Wong and colleagues had 13 cyclists ride to exhaustion at their measured FTP and got 33.7 plus or minus 7.6 minutes — barely half the implied hour. Add 15 watts and time to exhaustion collapsed to 22.0 plus or minus 5.7 minutes [Wong et al. 2022]. The authors concluded FTP is not a valid marker of the maximal metabolic steady state. So when your software puts CP and FTP side by side, it is comparing a modelled domain boundary against a number whose own definition does not survive a stopwatch. A gap between them is the expected output of that comparison, not a fault in either estimator.

02 / 04

What the agreement studies actually found

Three agreement studies in trained cyclists, all reaching the same verdict: not interchangeable. Karsten found CP 7 W above FTP [Karsten et al. 2021]. McGrath found it 16 W above [McGrath et al. 2021]. Morgan found it 3 W below [Morgan et al. 2019]. Every one of them reported individual limits of agreement spanning 40 watts or more.

Karsten and colleagues tested 17 trained cyclists and triathletes: CP 256 plus or minus 50 W against FTP 249 plus or minus 44 W. Mean bias +7 W, about 2.8%. Correlation r = 0.969. And 95% limits of agreement of -19 to +33 W [Karsten et al. 2021]. Read those last two numbers together. The correlation is near-perfect, and the interval inside which any individual rider's gap will fall is 52 watts wide. Both facts come from the same 17 riders.

McGrath and colleagues, working with highly-trained athletes, found CP 282 plus or minus 53 W against FTP 266 plus or minus 55 W — a 16 W bias, roughly 6%, at p < 0.001, exceeding the agreement thresholds they had set in advance [McGrath et al. 2021]. Morgan and colleagues went the other way: 12 competitive male cyclists, CP 275 plus or minus 40 W against FTP 278 plus or minus 42 W, a non-significant -3 W bias, limits of agreement -36 to +30 W [Morgan et al. 2019]. That non-significant result is worth stating carefully. It is no evidence of a difference in the group mean. It is not evidence that the two numbers are the same, and the 66-watt agreement band in that same study is the proof.

This is the association-versus-prediction distinction, and it is the whole article. A correlation of 0.969 answers a ranking question: among these riders, does a high CP travel with a high FTP? Yes, almost perfectly. The limits of agreement answer a prediction question: given one rider's FTP, how tightly can I pin their CP? Within roughly 40 to 66 watts, depending on which cohort you ask. High correlation paired with wide limits of agreement is the signature of a metric that describes a cohort well and prescribes to an individual badly. On a 250-watt rider, a 30-watt error moves a sweet-spot target from a productive 218 W to a session-ending 244 W.

03 / 04

Why there is no correction factor

The group-mean gap is small and its sign flips between studies: +7 W, +16 W, -3 W. A constant that adds 6% in one cohort subtracts 1% in another, while individual limits of agreement run 40 to 66 watts wide in all three. There is nothing stable enough to correct with.

McGrath and colleagues did propose a correction factor, derived from the bias they measured in their own athletes [McGrath et al. 2021]. That is a defensible thing to publish and a dangerous thing to generalise. Apply their +16 W adjustment to Morgan's cyclists, where CP actually sat 3 W below FTP, and you have made the estimate 19 watts worse on average [Morgan et al. 2019]. A correction factor that flips sign between two cohorts of trained cyclists is not describing a physiological offset. It is describing the two protocols each lab happened to run.

Which is the second reason to distrust a tidy conversion: the CP your software reports is partly a protocol artifact. The reference protocol is 3 to 5 fresh maximal trials on separate days, each exhausting the rider in roughly 2 to 15 minutes [Jones & Vanhatalo 2017]. Fitting CP from a whole-season best-effort power curve — what nearly every consumer tool does, including us — violates every clause of that. Season data mixes fresh efforts with fatigued ones, race-day motivation with Tuesday-night motivation, and durations chosen by the terrain rather than by the protocol. Borszcz and colleagues, meta-regressing 36 studies, found agreement between critical power and the maximal lactate steady state shifted with the longest predictive trial duration, the type of predictive trials, and the model-fitting parameters [Borszcz et al. 2024]. Change the protocol and you change the number.

04 / 04

Which number should anchor your training

Anchor to one number and know which one it is. If your intervals are prescribed as a percentage of FTP, CP is a cross-check, not a substitute — and vice versa. Do not spend a season trying to reconcile the two into a single true watt figure. The disagreement is structural, and no amount of extra data closes it.

The failure mode is specific and easy to miss. Take a 4x8 prescribed at 105% of a 249 W FTP: 261 watts. Recompute the same session off a 256 W CP and you get 269 watts. Eight watts of silent drift, entirely because both numbers were sitting in the app under the word threshold. The mistake is never having two numbers on screen. The mistake is letting a target computed against one of them get applied to the other, which is what happens every time a rider copies an interval target out of one platform and into another.

This sits inside the larger argument our ftp-without-a-test pillar makes — that estimating threshold from the data you already have is good enough for nearly all self-coached training, and a forced test is rarely the best use of a Saturday. The caveat this spoke adds is about what that estimate is. Whatever the model reports, CP or FTP, is a parameter estimate carrying a confidence interval, not a measurement. Treat your anchor as stable rather than exact, and let it move slowly, on trend rather than on any single ride.

Here is our own position, stated plainly. We fit CP from a season-best mean-maximal-power curve using efforts between 2 and 20 minutes, drop the longest effort when leaving it out shifts CP by more than 5% — the signature of a 20-minute personal best that was not actually maximal — floor the result at 85% of a demonstrated 5-minute maximum, and then convert CP to an FTP-equivalent using a fixed 0.97 multiplier. That multiplier is precisely the correction factor this article has just argued does not exist for an individual. We use it because it is the best available population approximation, and because a field-fitted number that actually moves your training beats a lab number you will never collect. It is a pragmatic field approximation, not the Jones and Vanhatalo protocol [Jones & Vanhatalo 2017], and we are not going to pretend otherwise. Intervals.icu, Xert, TrainingPeaks and Garmin all make some version of the same compromise. The part we will commit to is naming it.

Common questions

Quick answers

Is Critical Power the same thing as FTP?

No. Critical Power is the asymptote of your power-duration curve, fitted from maximal efforts. FTP is a convention approximating one-hour power, usually taken as 95% of a 20-minute test. They correlate strongly in trained cyclists — r = 0.969 in one cohort — but the 95% limits of agreement in that same study ran -19 to +33 W [Karsten et al. 2021]. Strong relationship, unreliable individual conversion.

My CP reads 15 watts higher than my FTP. Which one is wrong?

Probably neither. A 15-watt gap sits comfortably inside the individual limits of agreement reported in every published comparison [Karsten et al. 2021, McGrath et al. 2021, Morgan et al. 2019]. It tells you the two estimators disagree about you by a normal amount. Chasing a third test to break the tie usually just adds a third number.

Can I convert CP to FTP with a multiplier?

Not reliably for one rider. The published mean biases are +7 W, +16 W and -3 W across three cohorts of trained cyclists — a multiplier fitted to one of them is wrong in the others, sometimes in the opposite direction [Karsten et al. 2021, McGrath et al. 2021, Morgan et al. 2019]. Population conversions are fine for describing a group. They are not accurate enough to price an individual interval.

Should I train to CP or to FTP?

Train to whichever number your plan is anchored to, and keep it consistent. If your workouts arrive as percentages of FTP, switching the anchor to CP without recomputing every target shifts your whole intensity distribution upward. Consistency matters more here than picking the physiologically purer construct, because the prescription error from mixing anchors is larger than the difference between them.
References

Sources cited in this guide

  1. 01
  2. 02
    McGrath et al. 2021. Do Critical and Functional Threshold Powers Equate in Highly-Trained Athletes?. International Journal of Exercise Science.
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
In this series

More inside FTP without a test

Start here · Foundational guide

FTP without a test: estimating threshold from real rides

How to find FTP without a 20-minute or ramp test — using your power curve, critical-power modeling, and the rides you've already done.

Read the full guide

Other articles in this series

  1. 01

    Indoor vs outdoor FTP: why the numbers differ

    Why your indoor FTP reads lower than outdoor — heat, cooling, motivation, and power-source differences — and whether to keep two numbers.

  2. 02

    How to estimate FTP without a power meter

    Estimating FTP from heart rate, RPE, and Strava when you don't own a power meter — how close you can get and where the method breaks down.

  3. 03

    20-minute vs 8-minute FTP test: which to use

    How the 20-minute and 8-minute FTP tests differ, the multipliers each uses, and which one fits your riding — plus why both are only protocols.

  4. 04

    Why cycling apps show you different FTP numbers

    Strava, Xert, Intervals.icu, and TrainingPeaks can each report a different FTP. Why the estimates diverge and which one to trust.

  5. 05

    Is your ramp test FTP too high? Why it happens

    Ramp tests overestimate FTP for anaerobically-gifted riders and underestimate it for diesels. Why the 75% rule misfires and how to correct it.

  6. 06

    How often should you test your FTP?

    How often to re-test FTP as a self-coached cyclist — twice a year, not every six weeks — and why modeled estimates change the cadence.

  7. 07

    Can you actually hold your FTP for an hour?

    Published time-to-exhaustion at FTP runs 34 to 51 minutes across studies, never 60, and the individual spread is wide. Why the hour is a bad test.

  8. 08

    FTP stopped improving: are you a non-responder?

    Non-responder labels flip for 45% of people depending only on the statistics used. What a flat FTP usually means instead, and what to check first.

  9. 09

    Durability: why your power fades late in long rides

    What durability is, why the kJ threshold your app uses is arbitrary, and the honest evidence on whether it is independent of FTP or trainable.

Free training analysis · No card · ~3 minutes

Try the adaptive coach yourself.

Connect Strava and the coach reads your last eight weeks — intensity balance, ramp rate, what's working, what's holding you back — then drafts your training plan around it.

Free 14-day trial, no card.