The appeal of a lactate meter is that it replaces argument with a number. No more wondering whether that was really threshold pace -- prick a finger, read 3.2, adjust. Except that when researchers measured the same blood on four handheld analysers, they found that at a true concentration of 3 mM the 95% prediction interval was 0.80 mM on a Lactate Pro2. Two pricks from the same sample can legitimately read 2.6 and 3.4. One of those tells you to hold the interval, the other tells you to back off. The number is real. The precision people believe it has is not.
The Simple Version
A portable lactate meter measures something genuine, and it correlates well with laboratory analysers. What it does not have is the resolution that lactate-guided training talk implies. The device noise alone is a few tenths of a millimole; on top of that sit sampling-site differences, day-to-day biology, and -- largest of all -- the fact that fourteen different accepted methods will pull fourteen different thresholds out of the same curve. And the question that ought to settle it has not been answered: no randomised trial has shown that training guided by lactate produces better results than training guided by power, pace or heart rate. Meanwhile, free alternatives agree with the aerobic threshold about as closely as most people need.
How It Works
What the Device Actually Does
Correlation Is Not Precision
Portable analysers correlate strongly with laboratory reference devices. That fact gets quoted constantly and it is true. It is also not the property you need.
Correlation tells you the device tracks changes. Precision tells you how much a single reading can wander. Mentzoni and colleagues compared four handhelds -- Lactate Plus, Lactate Pro2, Lactate Scout 4 and TaiDoc -- against two stationary analysers, across blood lactate from 0.88 to 4.89 mM. At a reference value of 3 mM, the 95% prediction intervals came out at 0.72 mM for Plus, 0.80 mM for Pro2 and 0.87 mM for Scout 4. The stationary YSI managed 0.23 mM.
Their conclusion, verbatim: the precision of the four handheld lactate analysers evaluated in this study was poor.
There is a second issue at the top of the range. Bonaventura and colleagues found every portable analyser they tested under-read relative to the criterion device, and the negative bias grew worst at high concentrations, around 15-23 mM. If your meter tells you an all-out effort produced 12 mM, the real figure was probably higher.
The Honest Counterpoint
It would be easy to stop there and conclude the devices are junk. They are not, and the same study that documented the bias also found that within-brand analytical error was substantially smaller than biological variation.
That is the important framing. Your meter is not the main source of noise in your lactate number. Your body is. Which means buying a more expensive meter fixes very little, and the interesting question is what else is moving.
Fourteen Methods, Fifty-Six Answers
Here is the part that should give anyone pause before building a training plan on a lactate curve.
Jamnick and colleagues took 17 trained cyclists, ran them through graded exercise tests with stage durations of 1, 3, 4, 7 and 10 minutes, and then calculated 56 different lactate thresholds using 14 accepted methods across four protocols. They also measured each rider's actual maximal lactate steady state directly, which came out at 264 +/- 39 W.
Their finding: many of the traditional threshold methods were not valid. The best performer was the Modified Dmax method derived from the four-minute-stage test, which landed within 1.1 W of the measured steady state with an ICC of 0.96. Methods based on fixed concentrations -- the familiar 4 mmol/L -- overestimated it substantially.
Two things follow. First, "my threshold is 260 W" is an incomplete statement; without the protocol and the method it is not reproducible even by you. Second, the choice of method can move your prescribed training zone by more than a season of training would.
The Question Nobody Has Answered
Set aside precision for a moment. The premise of buying a meter is that intensity guided by lactate produces better outcomes than intensity guided by watts.
There does not appear to be a randomised controlled trial testing that. The research digest prepared for this article found none, and an independent search found none either -- what exists compares intensity distributions (polarised against threshold, and so on), or validates lactate as a measurement, or documents what elite programmes do. None of it isolates the guidance method while holding total volume and intensity constant.
This is absence of evidence rather than evidence of absence, and it does not mean the practice is wrong. Elite coaches using lactate in-session are mostly using it as a brake -- a way to stop athletes drifting from the heavy domain into the severe one during high-volume threshold blocks. That is a coherent purpose. It is just not the same as a demonstrated performance advantage over a power meter.
What Free Alternatives Achieve
Two non-invasive markers have been validated against gas exchange.
The heart rate variability index DFA a1 hits 0.75 close to the first ventilatory threshold: in 15 participants, r = 0.99 against VO2 and 0.97 against heart rate, with group means of 39.8 versus 40.1 mL/kg/min and 152 versus 154 bpm. It needs a chest strap that records beat-to-beat intervals and an app.
The talk test is cruder and free. Scored on a visual analogue scale, it identified the first ventilatory threshold with a mean difference of -1.3 W and the second with 11.1 W. The caveat matters: that study was run in 19 overweight and obese patients, not trained athletes, so the numbers should not be assumed to transfer to a cyclist with a 300 W threshold.
Example
Example: Signal and Noise on a Single Reading
An athlete is running a threshold session with a target of "hold 3.0 mmol/L". They prick, read 3.2, and decide to hold the pace. What do we actually know?
Start with the device. On a Lactate Pro2, the 95% prediction interval at 3 mM is 0.80 mM wide. So a displayed 3.2 is consistent with a true value roughly between 2.8 and 3.6.
Now add the things around it:
| Source of movement | Direction and rough size |
|---|---|
| Analyser prediction interval at 3 mM | +/- 0.4 mM (Pro2) |
| Sampling site, if you switched finger and earlobe | systematic, and site-dependent |
| Glycogen state after yesterday's session | shifts the whole curve |
| Which threshold method you used to set the 3.0 target | worth more watts than a training block |
The target itself deserves scrutiny too. There is no universal 3.0 mmol/L boundary: an individual's actual maximal lactate steady state can sit well below or well above that figure. Prescribing everyone the same absolute concentration ignores that.
What to take from this. The number is not useless -- it is directionally right, and a reading of 6 mM tells you something a reading of 2 mM does not. What it cannot support is the decision our athlete just made. Distinguishing 3.2 from 2.9 is below the resolution of the instrument, and both are inside the range that a different threshold method would have called something else entirely.
Track the shape of your curve across a season and its movement rightward. Do not manage a single interval to one decimal place.
Practical Rules
Practical Rules
-
Buy a meter to build a curve, not to chase a number. Periodic step tests that show your curve shifting right over months are within what the device can resolve. Adjusting an interval because you read 3.2 instead of 2.9 is not.
-
Fix your protocol and never change it. Stage duration alone moves the answer. Record stage length, increment size, starting power, sampling timing, and repeat them exactly. A test compared against a differently-run test tells you nothing.
-
Use one sampling site, always. Fingertip and earlobe do not give the same value, and the difference is systematic. Pick one and stay with it for as long as you want your results to be comparable.
-
Wipe the first drop away. Sweat contains lactate, and so does anything sugary you have touched. Clean the site, let it dry fully, discard the first drop and test the second. This is standard practice rather than a finding from a trial, but the failure mode -- a wildly inflated reading -- is well known.
-
Never test in a fatigued or glycogen-depleted state. Depletion flattens the curve and pushes it right, which looks exactly like improved fitness. An athlete who tests after a hard week can conclude they have made a breakthrough when the opposite is true.
-
Try the free options first. If DFA a1 tracks the aerobic threshold at r = 0.99 against gas exchange and a chest strap costs a fraction of a meter plus a year of strips, start there. Add the meter when you can name the specific decision it will change.
-
Be clear about what the elites are using it for. In-session lactate in the Norwegian model functions as a ceiling -- it stops athletes from turning a controlled threshold session into a hard one, which is what makes doing several per week survivable. That is a workload-management tool. Nobody has demonstrated it beats a power meter for the same purpose.
Evidence Base
Evidence Base
On device performance the evidence is recent and clean. Mentzoni et al. (2024) tested four handhelds against two stationary analysers over a physiologically relevant range and reported prediction intervals directly, which is the number a practitioner actually needs and which most validation papers do not give. Bonaventura et al. (2015) is older and small -- 22 samples from 5 subjects -- but its finding of consistent negative bias at high concentrations is specific and has been reproduced.
The most useful thing in that literature is also the most easily overlooked: analytical error within a brand is smaller than biological variation. The device is not your limiting factor.
On threshold determination, Jamnick et al. (2018) is the reference point. Its design is what makes it valuable -- rather than comparing methods to each other, it compares them to a directly measured maximal lactate steady state in the same riders. The sample is 17 trained male cyclists, so the specific watt figures should not be read as universal, but the structural finding -- that method choice and protocol design move the answer more than most people assume -- is not sample-dependent.
On lactate-guided training outcomes there is a gap, and it should be stated as a gap. No randomised trial isolating lactate as the guidance variable was located, either in the research prepared for this article or in an independent search. Two searches failing to find something is weak evidence, and it is possible such work exists and was missed. What can be said is that the confident claims made for lactate-guided training rest on physiological reasoning and elite practice rather than on a controlled comparison.
On the alternatives, note the populations. Rogers et al. (2021) validated DFA a1 in 15 participants on a treadmill, and subsequent work has found the index behaves less predictably in some athletes and with some heart rate sensors, so treat the r = 0.99 as a best case rather than a guarantee. The talk test validation was conducted in overweight and obese patients, and its agreement figures cannot simply be transferred to trained endurance athletes.
Two further limitations run through everything here. Sample sizes are small throughout -- 15, 17, 19, 22 samples from 5 people. And the day-to-day reproducibility of an individual's lactate threshold is not well characterised, which means the size of change you need to see before calling it real improvement is genuinely uncertain.