跳至主要內容

Validity of Wearable Devices in Training Monitoring: A Comparative Study of Commercial Products vs. Laboratory Equipment

訓練科學

Foreword: A Scientific Bridge from the Lab to Taiwan’s Roads

Smartwatches, bands, and rings—wearable devices allow everyone to quantify training and recovery. But how trustworthy are these numbers? Treating commercial products as laboratory gold standards for decision-making can be misleading. Understanding the validity differences across metrics is key to “using the right data without being held hostage by it.” This article compares the accuracy of commercial wearables against laboratory equipment based on validation studies, and provides principles for rational use.

Heart Rate and Step Count: Relatively Reliable

Optical heart rate (PPG) shows acceptable agreement with chest straps/ECG during steady-state, low-to-moderate intensity exercise, but errors increase significantly during intense interval changes, high-intensity intervals, or when wrist movement is large, due to motion artifacts and blood flow changes. Step count is accurate during walking/running, but tends to undercount in situations like cycling or pushing a cart. Overall, heart rate trends and step count are suitable for daily monitoring, but for key training sessions (such as interval heart rate zones), a chest strap is still recommended to ensure accuracy.

Metric Validity Usage Recommendation
Heart rate (steady-state) Acceptable Fine for daily use
Heart rate (intense interval changes) Increased error Use chest strap for key sessions
Step count Accurate for walking/running Prone to undercounting while cycling
Energy expenditure 20–30%+ error Look at trends, not absolutes
Sleep staging Total duration acceptable, staging limited Reference trends

Energy Expenditure and Sleep Staging: Larger Errors

Energy expenditure (calories) is one of the least accurate metrics on wearables, with validation studies commonly showing 20–30% or even larger errors, because it is indirectly estimated from heart rate, movement, and other factors, with large individual variability. Sleep staging (deep/light/REM), while acceptable for “total sleep time,” has limited staging accuracy compared to polysomnography (PSG), with deep sleep/REM often over- or underestimated. These metrics are suitable for viewing “relative trends” rather than absolute values.

Usage Principle Explanation
Same-device comparison Absolute values are not comparable across brands
Fixed conditions Measure HRV at rest in the morning
Look at trends Relative changes are more reliable than absolute values
Cross-validation Use subjective feelings and performance as the benchmark

HRV and Recovery Scores: Look at Trends, Not Absolutes

Many devices provide HRV and “recovery scores.” HRV measurements are more stable under fixed conditions (such as morning rest), but algorithms and measurement timing vary by brand, making absolute values incomparable across devices. Recovery scores are black-box composite metrics that can serve as personal trend references, but should not be treated as gospel. The principle is: same device, fixed conditions, look at trends, and always cross-validate with subjective feelings and training performance to avoid letting a single number dictate decisions.

The Principles and Error Sources of Optical Heart Rate

Wrist-worn optical heart rate uses photoplethysmography (PPG)—shining green light onto the skin and detecting changes in reflected light caused by blood flow to estimate heart rate. Its error sources include: motion artifacts from wrist movement (especially during intense or variable-speed exercise), strap tightness and fit, skin pigmentation, ambient light interference, and changes in blood perfusion (such as reduced peripheral blood flow in cold conditions). Therefore, optical heart rate is acceptable during steady-state, low-to-moderate intensity, but errors increase significantly during intervals, sprints, or high-frequency wrist movements, sometimes even producing the illusion of “locking onto cadence.” Understanding these limitations clarifies why chest straps (which directly measure the heart’s electrical signals) are significantly more accurate than wrist optical heart rate for key high-intensity or interval sessions—worth using for precise training.

Data Literacy: How Not to Be Held Hostage by Wearable Numbers

The proliferation of wearables has brought “data anxiety”—excessive fixation on sleep scores, recovery metrics, or step counts that affects physical and mental well-being. Healthy data literacy includes: understanding the validity and limitations of each metric (such as large errors in energy expenditure, limited sleep staging), interpreting trends rather than single-day absolute values, comparing with the same device under fixed conditions (not comparable across brands), and always using subjective feelings and actual performance as the final benchmark. Devices should be a “dashboard” that aids self-understanding, not an “examiner” that creates pressure. If data causes anxiety or sleep pressure instead (such as feeling tense about sleep scores), then step back—the body’s direct sensations are always closer to reality than algorithmic estimates. Using data wisely without being held hostage by it is an essential skill in the wearable era.

Methods of Wearable Validation Studies and Consumer Insights

Wearable device validity validation studies use laboratory gold standards (ECG, metabolic carts, polysomnography) as references to quantify errors across metrics. The consumer insights from these studies are highly practical: heart rate is acceptable at steady state but has large errors during intense interval changes; step count is accurate for walking/running but prone to undercounting while cycling; energy expenditure errors often reach 20–30%; total sleep time is acceptable but staging is limited; HRV is more stable under fixed conditions but not comparable across brands. Understanding these, consumers can “use the right metrics, in the right way”—treating devices as personal trend dashboards, using chest straps for key sessions, treating energy expenditure as reference only, and not comparing absolute values across brands. Validation studies also remind us that marketing often exaggerates accuracy. The pragmatic approach is: make good use of the convenience and trend insights wearables provide, but understand their limitations, use subjective feelings and actual performance as the final benchmark, and avoid being misled by inaccurate numbers or having them create anxiety.

Cross-Disciplinary Integrated Perspective: Measurement Science and Data Literacy

Research on wearable device validity integrates measurement science, engineering, and exercise physiology, highlighting the importance of “data literacy” in the wearable era. It uses laboratory gold standards to test the accuracy of commercial products, revealing validity differences across metrics—heart rate is acceptable at steady state, energy expenditure has large errors, and sleep staging is limited. The value of this cross-disciplinary integration lies in teaching us how to use wearable data rationally, rather than blindly believing or following it. From a measurement perspective, different metrics have varying accuracy and limitations; from an engineering perspective, there are error sources like motion artifacts in optical heart rate; from an application perspective, interpretation should rely on trends rather than absolute values, and same-device rather than cross-brand comparisons. This perspective cultivates essential “data literacy”—understanding the validity and limitations of data, interpreting trends, and using subjective feelings and actual performance as the final benchmark. It also reminds us that marketing often exaggerates accuracy, and consumers need to remain critical. In an era where everyone wears devices, this literacy helps us harness the convenience and trend insights of technology while avoiding being misled by inaccurate numbers or having them create anxiety. Understanding measurement science makes us masters of data, not slaves to it.

From Research to Daily Life: A Rational Framework for Using Data

Rational use of wearable data can follow the framework of “Trend Dashboard—Metric Grading—Fixed Conditions—Body as the Benchmark.” Trend Dashboard: treat the device as a “personal trend dashboard,” tracking your own long-term changes rather than fixating on single-day absolute values or comparing with others. Metric Grading: understand the validity of each metric—heart rate (acceptable at steady state, large errors during intense interval changes, use chest strap for key sessions), step count (accurate for walking/running, prone to undercounting while cycling), energy expenditure (20–30% error, trend reference only, don’t calculate diet precisely from it), sleep (total duration acceptable, staging limited), HRV (more stable under fixed conditions, not comparable across brands). Fixed Conditions: measure metrics like HRV with the same device under fixed conditions (morning rest) and look at trends; absolute values are not comparable across brands. Body as the Benchmark: always use subjective feelings and actual performance as the final benchmark; if data causes anxiety or sleep pressure instead, step back—the body’s direct sensations are closer to reality than algorithmic estimates. The core of this framework is: make good use of the convenience and trend insights of wearable technology, understand its validity and limitations, ground decisions in bodily sensations and performance, and be a master of data rather than a slave held hostage by inaccurate numbers.

Taiwan-Specific Applications: Climate, Events, and Cultural Context

Taiwan’s sports enthusiasts have high wearable adoption rates, making rational use especially important. Recommendations: treat the device as a “trend dashboard” rather than a precision laboratory—track your own long-term changes rather than fixating on single-day absolute numbers or comparing with others. Use a chest strap for key sessions (intervals, lactate threshold) to ensure heart rate accuracy. Hot and humid environments affect optical heart rate and sleep data (sweating, body temperature), so factor in context when interpreting. The core principle remains unchanged: data assists decision-making, but the body’s subjective feelings and actual performance are the final benchmark.

Taiwan’s wearable adoption rate is high, making rational use especially important. It is recommended to treat the device as a “trend dashboard,” tracking personal long-term changes rather than fixating on single-day numbers or comparing with others. Use a chest strap for key sessions to ensure heart rate accuracy. High heat and humidity affect optical heart rate and sleep data, so factor in context when interpreting. The core principle remains unchanged: data assists decision-making, but bodily sensations and actual performance are the final benchmark.

Frequently Asked Questions and Myth Clarification

Myth 1: Are the calories displayed on your watch accurate? Energy expenditure is one of the least accurate metrics, with errors often reaching 20–30%. Don’t use it to precisely calculate your diet; treat it only as a reference for trends.

Myth 2: Can sleep scores reflect true sleep quality? Total sleep duration is somewhat reliable, but the accuracy of deep sleep/REM staging is limited. Judge based on how refreshed you feel upon waking and long-term trends—don’t obsess over the score.

Myth 3: Can data from different brands be compared? Algorithms vary, so absolute values across brands are not comparable. Track personal trends using the same device under consistent conditions.

How to Read Sports Science Research: Developing Evidence Literacy

This article cites 4 studies from top international journals (such as Journal of Applied Physiology, Medicine & Science in Sports & Exercise, Sports Medicine, Nature, Cell series, etc.), but as a reader, cultivating “evidence literacy” can help you absorb this knowledge more rationally rather than accepting it at face value. First, distinguish study types: randomized controlled trials (RCTs) have the strongest causal inference power, while observational studies (cohort, cross-sectional) can only show associations rather than causation. Animal and cellular studies reveal mechanisms, but translation to humans requires caution. Second, pay attention to samples and contexts: results from small samples or specific populations (such as elite athletes or particular age groups) may not apply to you; studies predominantly based on European and American populations also warrant consideration regarding applicability to Taiwanese populations. Third, emphasize effect size rather than just looking at “statistical significance”: statistical significance does not equal a practically meaningful benefit—ask “is this difference important in real training or health terms?” Fourth, be wary of over-extrapolation and commercialization: preliminary findings from single studies are often exaggerated into “miracle” products or methods; wait for replication and systematic reviews. Fifth, judge comprehensively based on the “consistency” of mechanistic, associative, and interventional evidence, rather than rejecting everything due to flaws in a single study or accepting everything because of one impressive result. Sixth, understand that “individual variability” is the norm in sports science: the same intervention elicits different responses in different people due to genetics, training background, lifestyle, and environment. Studies present group averages—when applying to yourself, be sure to observe your own actual responses and adjust accordingly. Seventh, prioritize the “fundamentals”: sleep, nutrition, consistent training, and recovery—these have overwhelming evidence support and clear benefits—are always worth investing in before any novel supplements, gadgets, or methods. Many seemingly sophisticated interventions yield far less marginal benefit than getting the basics right. Sports science is a constantly evolving field. Maintaining an open yet critical attitude, updating your understanding as evidence evolves, while respecting individual differences and valuing fundamentals, is the only way to truly translate cutting-edge research from international journals into training and health decisions that are useful, safe, and sustainable long-term—without falling into blind trend-chasing or deference to a single authority.

Key Takeaways from This Article

Synthesizing the above interdisciplinary research and mechanistic analyses, the core points can be distilled as follows: Treat devices as trend dashboards: track personal long-term changes, don’t obsess over single-day absolute values. Use a chest strap for key workouts: for heart rate accuracy during interval/threshold training, chest straps outperform wrist watches. Energy expenditure is for reference only: errors are large, don’t use it to precisely calculate your diet. Measure HRV with the same device under consistent conditions: absolute values across brands are not comparable. Use bodily sensation as the ultimate criterion: data assists but does not replace subjective feeling and performance. Behind these points lies the convergence of multiple fields—sleep science, immunology, genomics, neuroscience, microbiology, endocrinology, and data science—which together convey a core message: the benefits and adaptations of exercise are the integrated result of multiple body systems working in coordination, not something captured by any single factor. Understanding this interdisciplinary, integrative perspective helps us move beyond fragmented “treat-the-symptom” thinking and view training, recovery, and health more holistically. Only by incorporating these principles into daily training and life, and dynamically adjusting based on individual circumstances, actual responses, and professional advice, can we translate cutting-edge findings from top international journals into practices that are truly feasible, safe, and sustainable within Taiwan’s climate, racing calendar, and lifestyle context. The value of sports science ultimately lies in helping every athlete—elite or amateur, young or old—exercise smarter, healthier, and with more enjoyment, achieving physical and mental growth along the way.

Practical Recommendations for Taiwanese Athletes

  1. Treat devices as trend dashboards: Track personal long-term changes, don’t obsess over single-day absolute values.
  2. Use a chest strap for key workouts: For heart rate accuracy during interval/threshold training, chest straps outperform wrist watches.
  3. Energy expenditure is for reference only: Errors are large, don’t use it to precisely calculate your diet.
  4. Measure HRV with the same device under consistent conditions: Absolute values across brands are not comparable.
  5. Use bodily sensation as the ultimate criterion: Data assists but does not replace subjective feeling and performance.

Research Citations and Further Reading

  • Bunn, J. A., et al. (2018). Current state of commercial wearable technology in physical activity monitoring. International Journal of Exercise Science, 11(7), 503–515.
  • Nelson, B. W., & Allen, N. B. (2019). Accuracy of consumer wearable heart rate measurement during an ecologically valid 24-hour period. JMIR mHealth and uHealth, 7(3), e10828.
  • Düking, P., et al. (2018). Recommendations for assessment of the reliability, sensitivity, and validity of data provided by wearable sensors. JMIR mHealth and uHealth, 6(4), e102.
  • Chevance, G., et al. (2022). Accuracy of wearable devices. npj Digital Medicine / related validation studies.

This article is a translation of sports science knowledge. Individual physiological responses vary. Please consult professional coaches and sports medicine physicians before making any training or intervention adjustments, and proceed gradually according to your personal health status.

相關影片
訂閱CT的頻道

訂閱 CT Yeh,看武嶺實測與路線攻略

北進武嶺、西進武嶺、經典百K,每條路線都親自騎過,配速、爬升、補給點全部實拍實測。

467 部影片 · 累計 838 萬次觀看