Critical Power vs FTP: why your two numbers disagree, and which one should anchor your training
Your software shows you two numbers: a Critical Power fitted from your power curve, and an FTP from a 20-minute test. They are not the same construct, and for you specifically the gap can run up to about 36 watts in either direction. Across three studies of trained cyclists the average gap was +7 W, +16 W, and -3 W. Small, and inconsistent even in sign. That is exactly why no correction factor exists. Here is what the agreement data shows, and which number should anchor your prescriptions.
By Jim Camut · Former pro & ex-Bruyneel Academy racer
Updated Jul 21, 20264 chapters7 citations
CP and FTP are not the same construct
Critical Power is the asymptote of your power-duration curve — a modelled boundary above which the physiological response stops being steady-state [Jones & Vanhatalo 2017]. FTP is an operational convention: roughly the power you could hold for an hour, usually estimated as 95% of a 20-minute effort. One is a fitted parameter. The other is prescription shorthand.
The distinction is not academic hair-splitting, because the two have very different evidence behind them. Jamnick and colleagues reviewed the methods used to mark the boundary between the heavy and severe intensity domains and concluded there is little evidence supporting the validity of most commonly used approaches — with critical power and critical speed named as the exceptions [Jamnick et al. 2020]. FTP is not among the validated markers. It is a field convention that has proven extremely useful for organising training, which is a different claim from being a measured physiological threshold.
FTP is also not one-hour power, despite the definition. Wong and colleagues had 13 cyclists ride to exhaustion at their measured FTP and got 33.7 plus or minus 7.6 minutes — barely half the implied hour. Add 15 watts and time to exhaustion collapsed to 22.0 plus or minus 5.7 minutes [Wong et al. 2022]. The authors concluded FTP is not a valid marker of the maximal metabolic steady state. So when your software puts CP and FTP side by side, it is comparing a modelled domain boundary against a number whose own definition does not survive a stopwatch. A gap between them is the expected output of that comparison, not a fault in either estimator.
What the agreement studies actually found
Three agreement studies in trained cyclists, all reaching the same verdict: not interchangeable. Karsten found CP 7 W above FTP [Karsten et al. 2021]. McGrath found it 16 W above [McGrath et al. 2021]. Morgan found it 3 W below [Morgan et al. 2019]. Every one of them reported individual limits of agreement spanning 40 watts or more.
Karsten and colleagues tested 17 trained cyclists and triathletes: CP 256 plus or minus 50 W against FTP 249 plus or minus 44 W. Mean bias +7 W, about 2.8%. Correlation r = 0.969. And 95% limits of agreement of -19 to +33 W [Karsten et al. 2021]. Read those last two numbers together. The correlation is near-perfect, and the interval inside which any individual rider's gap will fall is 52 watts wide. Both facts come from the same 17 riders.
McGrath and colleagues, working with highly-trained athletes, found CP 282 plus or minus 53 W against FTP 266 plus or minus 55 W — a 16 W bias, roughly 6%, at p < 0.001, exceeding the agreement thresholds they had set in advance [McGrath et al. 2021]. Morgan and colleagues went the other way: 12 competitive male cyclists, CP 275 plus or minus 40 W against FTP 278 plus or minus 42 W, a non-significant -3 W bias, limits of agreement -36 to +30 W [Morgan et al. 2019]. That non-significant result is worth stating carefully. It is no evidence of a difference in the group mean. It is not evidence that the two numbers are the same, and the 66-watt agreement band in that same study is the proof.
This is the association-versus-prediction distinction, and it is the whole article. A correlation of 0.969 answers a ranking question: among these riders, does a high CP travel with a high FTP? Yes, almost perfectly. The limits of agreement answer a prediction question: given one rider's FTP, how tightly can I pin their CP? Within roughly 40 to 66 watts, depending on which cohort you ask. High correlation paired with wide limits of agreement is the signature of a metric that describes a cohort well and prescribes to an individual badly. On a 250-watt rider, a 30-watt error moves a sweet-spot target from a productive 218 W to a session-ending 244 W.
Why there is no correction factor
The group-mean gap is small and its sign flips between studies: +7 W, +16 W, -3 W. A constant that adds 6% in one cohort subtracts 1% in another, while individual limits of agreement run 40 to 66 watts wide in all three. There is nothing stable enough to correct with.
McGrath and colleagues did propose a correction factor, derived from the bias they measured in their own athletes [McGrath et al. 2021]. That is a defensible thing to publish and a dangerous thing to generalise. Apply their +16 W adjustment to Morgan's cyclists, where CP actually sat 3 W below FTP, and you have made the estimate 19 watts worse on average [Morgan et al. 2019]. A correction factor that flips sign between two cohorts of trained cyclists is not describing a physiological offset. It is describing the two protocols each lab happened to run.
Which is the second reason to distrust a tidy conversion: the CP your software reports is partly a protocol artifact. The reference protocol is 3 to 5 fresh maximal trials on separate days, each exhausting the rider in roughly 2 to 15 minutes [Jones & Vanhatalo 2017]. Fitting CP from a whole-season best-effort power curve — what nearly every consumer tool does, including us — violates every clause of that. Season data mixes fresh efforts with fatigued ones, race-day motivation with Tuesday-night motivation, and durations chosen by the terrain rather than by the protocol. Borszcz and colleagues, meta-regressing 36 studies, found agreement between critical power and the maximal lactate steady state shifted with the longest predictive trial duration, the type of predictive trials, and the model-fitting parameters [Borszcz et al. 2024]. Change the protocol and you change the number.
Which number should anchor your training
Anchor to one number and know which one it is. If your intervals are prescribed as a percentage of FTP, CP is a cross-check, not a substitute — and vice versa. Do not spend a season trying to reconcile the two into a single true watt figure. The disagreement is structural, and no amount of extra data closes it.
The failure mode is specific and easy to miss. Take a 4x8 prescribed at 105% of a 249 W FTP: 261 watts. Recompute the same session off a 256 W CP and you get 269 watts. Eight watts of silent drift, entirely because both numbers were sitting in the app under the word threshold. The mistake is never having two numbers on screen. The mistake is letting a target computed against one of them get applied to the other, which is what happens every time a rider copies an interval target out of one platform and into another.
This sits inside the larger argument our ftp-without-a-test pillar makes — that estimating threshold from the data you already have is good enough for nearly all self-coached training, and a forced test is rarely the best use of a Saturday. The caveat this spoke adds is about what that estimate is. Whatever the model reports, CP or FTP, is a parameter estimate carrying a confidence interval, not a measurement. Treat your anchor as stable rather than exact, and let it move slowly, on trend rather than on any single ride.
Here is our own position, stated plainly. We fit CP from a season-best mean-maximal-power curve using efforts between 2 and 20 minutes, drop the longest effort when leaving it out shifts CP by more than 5% — the signature of a 20-minute personal best that was not actually maximal — floor the result at 85% of a demonstrated 5-minute maximum, and then convert CP to an FTP-equivalent using a fixed 0.97 multiplier. That multiplier is precisely the correction factor this article has just argued does not exist for an individual. We use it because it is the best available population approximation, and because a field-fitted number that actually moves your training beats a lab number you will never collect. It is a pragmatic field approximation, not the Jones and Vanhatalo protocol [Jones & Vanhatalo 2017], and we are not going to pretend otherwise. Intervals.icu, Xert, TrainingPeaks and Garmin all make some version of the same compromise. The part we will commit to is naming it.
Quick answers
Is Critical Power the same thing as FTP?
My CP reads 15 watts higher than my FTP. Which one is wrong?
Can I convert CP to FTP with a multiplier?
Should I train to CP or to FTP?
Sources cited in this guide
- 01Karsten et al. 2021. Relationship Between the Critical Power Test and a 20-min Functional Threshold Power Test in Cycling. Frontiers in Physiology.
- 02McGrath et al. 2021. Do Critical and Functional Threshold Powers Equate in Highly-Trained Athletes?. International Journal of Exercise Science.
- 03Morgan et al. 2019. Road Cycle TT Performance: Relationship to the Power-Duration Model and Association with FTP. Journal of Sports Sciences.
- 04Jones & Vanhatalo 2017. The Critical Power Concept: Applications to Sports Performance with a Focus on Intermittent High-Intensity Exercise. Sports Medicine.
- 05Jamnick et al. 2020. An Examination and Critique of Current Methods to Determine Exercise Intensity. Sports Medicine.
- 06Wong et al. 2022. Functional Threshold Power is Not a Valid Marker of the Maximal Metabolic Steady State. Journal of Sports Sciences.
- 07Borszcz et al. 2024. Agreement Between Maximal Lactate Steady State and Critical Power in Different Sports: A Systematic Review and Bayesian Meta-Regression. Journal of Strength and Conditioning Research.
More inside FTP without a test
Start here · Foundational guide
FTP without a test: estimating threshold from real rides
How to find FTP without a 20-minute or ramp test — using your power curve, critical-power modeling, and the rides you've already done.
Read the full guide
Other articles in this series
- 01
Indoor vs outdoor FTP: why the numbers differ
Why your indoor FTP reads lower than outdoor — heat, cooling, motivation, and power-source differences — and whether to keep two numbers.
- 02
How to estimate FTP without a power meter
Estimating FTP from heart rate, RPE, and Strava when you don't own a power meter — how close you can get and where the method breaks down.
- 03
20-minute vs 8-minute FTP test: which to use
How the 20-minute and 8-minute FTP tests differ, the multipliers each uses, and which one fits your riding — plus why both are only protocols.
- 04
Why cycling apps show you different FTP numbers
Strava, Xert, Intervals.icu, and TrainingPeaks can each report a different FTP. Why the estimates diverge and which one to trust.
- 05
Is your ramp test FTP too high? Why it happens
Ramp tests overestimate FTP for anaerobically-gifted riders and underestimate it for diesels. Why the 75% rule misfires and how to correct it.
- 06
How often should you test your FTP?
How often to re-test FTP as a self-coached cyclist — twice a year, not every six weeks — and why modeled estimates change the cadence.
- 07
Can you actually hold your FTP for an hour?
Published time-to-exhaustion at FTP runs 34 to 51 minutes across studies, never 60, and the individual spread is wide. Why the hour is a bad test.
- 08
FTP stopped improving: are you a non-responder?
Non-responder labels flip for 45% of people depending only on the statistics used. What a flat FTP usually means instead, and what to check first.
- 09
Durability: why your power fades late in long rides
What durability is, why the kJ threshold your app uses is arbitrary, and the honest evidence on whether it is independent of FTP or trainable.
Free training analysis · No card · ~3 minutes
Try the adaptive coach yourself.
Connect Strava and the coach reads your last eight weeks — intensity balance, ramp rate, what's working, what's holding you back — then drafts your training plan around it.
Free 14-day trial, no card.