You finished a ride with an average power of 195W. Last week you also averaged 195W. But that ride felt easy, and this one was brutal. You slept well, you ate well, you were not sick. So what happened? Here is the answer: last week was a flat road with steady effort. This week had three nasty climbs and a group ride with constant surges. Same average power. Completely different stress on your body. And the number that reveals the difference is not average power -- it is Normalized Power.
The Simple Version
Average power lies to you because your body does not respond to effort in a straight line. Pushing 500 watts for 5 seconds and then coasting for 5 seconds does not feel the same as holding 250 watts for 10 seconds, even though the average is identical. Your nervous system, your cardiovascular system, your hormonal responses -- they all have inertia. They ramp up, overshoot, and take time to settle. Normalized Power (NP) accounts for this biological lag by weighting surges more heavily than steady effort. And everything downstream -- Intensity Factor, Training Stress Score, your entire load tracking system -- is built on NP, not average power. Get NP wrong and every training decision based on it is wrong too.
How It Works
Why Your Body Does Not Do Averages
The 30-Second Lag
When you slam the pedals for a sprint, your heart rate does not instantly jump to 180 bpm. It takes 20-30 seconds to respond. Lactate does not flood your bloodstream immediately -- there is a production and clearance delay. Stress hormones like cortisol and adrenaline ramp gradually. Your oxygen consumption drifts upward over seconds, not milliseconds.
This means a 10-second surge at 400W imposes a metabolic cost that lingers for 20-30 seconds after you stop. Your body is still paying for that effort long after the power meter reads zero. Average power ignores this entirely. It treats a second at 400W and a second at 0W as equivalent to two seconds at 200W. Your muscles know better.
The NP Algorithm: Four Steps
Dr. Andrew Coggan developed Normalized Power in the early 2000s to solve this problem. The algorithm has four steps:
-
Calculate the 30-second rolling average of power. This smooths out instantaneous spikes and coasting, mimicking the body's physiological lag time. A 1-second sprint at 800W gets spread across 30 seconds, just like its metabolic impact.
-
Raise each 30-second average to the 4th power. This is the key step. By using the 4th power (not the 2nd or 3rd), the algorithm heavily penalizes surges. A value of 300W contributes 8.1 billion to the sum, while 150W contributes only 506 million -- a 16x difference for a 2x difference in power. This captures the nonlinear physiological cost of hard efforts.
-
Take the mean of all those 4th-power values.
-
Take the 4th root of the mean. This brings the number back to watts, giving you a single figure in familiar units.
The result is a number that always equals or exceeds average power. For a perfectly steady ride on a velodrome, NP equals average power exactly. For a chaotic criterium with constant attacks, NP can be 20-30% higher than average power.
xPower: The Alternative
Dr. Philip Skiba developed xPower as an alternative to NP, using a 25-second exponentially weighted moving average instead of a 30-second simple rolling average. The difference is subtle -- xPower responds slightly faster to changes in effort and typically produces values 1-3% different from NP. Both are valid. Most platforms use NP because of its longer history and wider adoption.
Variability Index: How Smooth Is Your Ride?
Variability Index (VI) is the ratio of NP to average power. It tells you how steady or surgy a ride was.
- VI = 1.00 to 1.02: Nearly perfect pacing. Velodrome, flat time trial, indoor trainer.
- VI = 1.05 to 1.15: Normal variation. Rolling terrain, moderate group rides, most structured workouts.
- VI = 1.15 to 1.25: Significant surges. Hilly rides, aggressive group rides, criteriums.
- VI > 1.25: Extremely variable. Attacks, breakaways, mountain stages. Very costly metabolically.
For Ironman and long-distance racing, a VI below 1.05 on the bike is the gold standard. Every point above 1.05 represents wasted energy -- effort that raised your physiological cost without increasing your average speed.
Intensity Factor and TSS
Intensity Factor (IF) puts NP in context by comparing it to your FTP:
IF = NP / FTP
An IF of 1.0 means you rode at your FTP intensity -- the kind of effort you could sustain for about 60 minutes. An IF of 0.75 represents a solid aerobic endurance ride in Zone 2-3. An IF above 0.90 for more than 2 hours means you had an exceptionally hard day.
Training Stress Score brings duration into the equation:
TSS = (duration in hours) x IF^2 x 100
Because TSS uses the square of IF, and IF is based on NP, the entire load management system -- CTL, ATL, TSB, your fitness and fatigue tracking -- ultimately depends on NP capturing the true cost of your rides.
Example
Worked Example: Three Rides, Three Realities
All three rides are 60 minutes by the same athlete with an FTP of 250W.
Ride A -- Flat Road Solo: Average power: 200W. NP: 204W. VI: 1.02. IF: 0.82. TSS: 67. A steady tempo ride. Almost no variation. The athlete held 195-210W for the entire hour. Low stress, good aerobic development, easy recovery.
Ride B -- Hilly Route: Average power: 188W. NP: 223W. VI: 1.19. IF: 0.89. TSS: 79. Average power is actually lower than Ride A because the descents dragged the average down. But the climbs pushed power to 280-320W for minutes at a time. NP captures this: the body paid for those surges. TSS is 18% higher than the "harder" flat ride.
Ride C -- Group Ride with Surges: Average power: 195W. NP: 241W. VI: 1.24. IF: 0.96. TSS: 92. Constant attacks, chasing wheels, sprinting out of corners. Average power looks moderate. NP reveals the truth: this was nearly an FTP-level effort. TSS is 37% higher than Ride A despite similar average power.
Without Normalized Power, you would look at these three rides and think they were roughly equivalent. Your legs know otherwise, and now you have the math to prove it.
Practical Rules
Practical Rules
-
Always look at NP, not average power. Average power is useful only for perfectly steady efforts like time trials or indoor trainer sessions. For anything with variation -- which is every outdoor ride -- NP is the true measure of how hard you worked.
-
Target IF 0.75-0.80 for Zone 2-3 endurance rides. If your IF regularly exceeds 0.80 on what should be easy rides, you are going too hard. This is the most common training mistake among age-group triathletes -- turning every ride into a tempo session.
-
IF above 0.90 for 2+ hours demands 2 days of recovery. This is race-level intensity sustained over a long duration. Your neuromuscular system, hormonal system, and glycogen stores all need time to rebuild. Do not stack hard days after a ride like this.
-
Aim for VI below 1.05 on Ironman bike legs and time trials. Every surge costs you disproportionately. A perfectly paced Ironman bike at VI 1.03 will leave you with fresher legs for the marathon than a surgy ride at VI 1.15, even if the average power is identical.
-
VI between 1.10 and 1.20 is normal for intervals and hilly rides. Do not stress about a high VI during a structured interval session -- that is by design. The concern is when VI is high during rides that should be steady.
-
VI above 1.25 is a red flag for long-distance racing. If your race files consistently show VI above 1.25, you are racing inefficiently. Practice pacing on hilly courses, resist the urge to match surges, and let gaps open on climbs that you can close on descents.
-
Updating your FTP invalidates old TSS comparisons. When your FTP changes, IF changes, and TSS changes retroactively in concept. A ride that was IF 0.85 at FTP 240 becomes IF 0.82 at FTP 250. Most platforms recalculate automatically, but be aware of this when comparing training blocks across fitness changes.
Evidence Base
Evidence Base
Normalized Power was developed by Dr. Andrew Coggan in the early 2000s and formalized in "Training and Racing with a Power Meter" (Allen and Coggan, 2006), the foundational text of power-based cycling training. The choice of a 30-second rolling average was based on the approximate time constant of physiological responses to changes in exercise intensity. The 4th power exponent was determined empirically -- it best matched the observed relationship between variable-intensity exercise and physiological markers like blood lactate, heart rate, and perceived exertion across a wide range of trained cyclists.
Dr. Philip Skiba introduced xPower in 2006 as a refinement, using a 25-second exponentially weighted moving average. The exponential weighting gives more influence to recent power values, arguably better reflecting how the body "remembers" recent effort. In practice, NP and xPower differ by 1-3% for most rides, and either provides a substantial improvement over average power for quantifying training load.
A key limitation of NP is that it was designed specifically for cycling. Running does not have an equivalent power-duration dataset with the same depth of validation. Garmin and Stryd offer running power metrics, but there is no scientific consensus on a "Normalized Running Power" standard. For running, Grade Adjusted Pace (GAP) serves a similar conceptual purpose -- accounting for the extra cost of hills -- but it is based on pace, not power, and uses different underlying assumptions. Swimmers have no widely accepted equivalent. For multisport athletes, NP remains the gold standard for the bike leg, while run and swim training load are best tracked through their own sport-specific metrics.