跳至主要內容

Effect Size and Practical Significance: The Truth Behind the Numbers—When Research Says "Effective," Is It Actually Useful for You?

訓練科學

Effect Size and Practical Significance: The Truth Behind the Numbers—When Research Says "Effective," Is It Actually Useful for You?

Opening: A Debate Sparked by a Cup of Coffee

Last month at a coffee shop in Taipei, an amateur mountain biker I coach, A-Kai, pushed his phone across the table toward me. The screen showed a shared article with a bold headline: “Study Confirms: This Supplement Significantly Boosts Athletic Performance.” His eyes lit up as he asked, “Coach, should I take this? It says it has a ‘significant’ effect!”

I didn’t answer right away. Instead, I asked him a question first: “Do you know that ‘significant’ might just mean ‘0.3 seconds faster’? And that 0.3 seconds was measured in a meticulously controlled lab with 300 subjects?” He was stunned.

This is the thing I’ve spent the most energy unpacking over my fifteen years coaching athletes of all levels and everyday exercisers: “Whether something works” and “how big the effect is, and whether it matters to you” are two completely different questions. The former is the p-value talking; the latter is what you actually feel when climbing Wuling, or doing intervals along the riverside.

In this article, I want to walk you into the world of “effect size” and the “minimal important difference.” They sound like ivory-tower statistical terms, but I promise—once you understand them, your eyes will be different the next time you read any sports science study, any piece of gear or supplement marketing. You’ll go from being led around by numbers to being someone who can question them back.

Conceptual Foundation: The p-value Myth and the Arrival of Effect Size

The p-value Only Answers a Very Narrow Question

Let’s start with a harsh truth: most people (including many journalists) misunderstand the p-value.

The p-value (we usually get excited and say “there’s a significant difference” when we see p < 0.05) actually only answers a very narrow question: “Assuming the two groups actually have no difference, how unlikely is it that I observed this result, or something even more extreme?” If the probability is low enough (below 5%), we say “this difference is unlikely to be pure chance.”

Note that it never tells you how big the difference is.

Here’s a trap many people don’t know about: as long as the sample size is large enough, almost any tiny difference can become “statistically significant.” If I tested a training method on 5,000 people, even if it only raised FTP (Functional Threshold Power) by an average of 1 watt, I could still get a pretty number like p < 0.001. But you and I both know that 1 watt is imperceptible to a 250-watt rider—it’s even smaller than measurement error.

Effect Size: Quantifying “How Much Difference”

Effect size was born to fill this gap. It doesn’t care about sample size; it directly answers: “On a standardized scale, how big is this difference, exactly?”

The most common effect size metric is Cohen’s d, and its logic is intuitive: divide the difference between the two group means by the variability of the data (the standard deviation). In other words, it asks, “Relative to the differences that already exist between individuals, is this difference big or small?”

According to the empirical conventions proposed by Jacob Cohen back in the day (note: conventions, not iron laws), the general interpretation is:

Cohen’s d Value Conventional Interpretation Plain-Language Translation
Around 0.2 Small effect It exists, but you need to look closely to see it
Around 0.5 Medium effect Visible to the naked eye, starts to feel meaningful in practice
Around 0.8 Large effect Obvious, usually worth taking seriously

But I want to add a very important reminder here: these thresholds are rough cross-disciplinary references, and Cohen himself emphasized that context matters. In the world of elite sport, the rules are completely different—the gap between gold and fourth place at the Olympics might translate to an effect size of only 0.1 to 0.2, tiny by any standard, but for the athlete, it’s the difference between standing on the podium and not, between getting a national bonus and not. So a “small effect” doesn’t equal “unimportant,” and a “large effect” doesn’t guarantee it matters to you. That’s exactly what we’re going to talk about next.

(For Cohen’s d thresholds and interpretation, see the Simply Psychology and LibreTexts statistics textbook links at the end of this article.)


Core Concept: Minimal Important Difference (MID / MCID)

From “Statistically Significant” to “Worth Your Attention”

If effect size standardizes the difference, then the Minimal Important Difference (MID; often called the Minimal Clinically Important Difference, MCID in clinical settings) brings the question back to the real world:

How big does this change need to be before you, as a person, actually feel it and find it meaningful?

MCID originated in clinical medicine—for example, how much does a pain-relief treatment need to lower the average pain score before patients feel “I really am in less pain,” rather than just seeing the questionnaire numbers move? Sports science borrowed this way of thinking: a training method or intervention needs to improve your performance beyond a certain threshold before it’s worth your time, money, and recovery cost.

There’s a harsh but practical point I love to cite: MCID is actually a “low bar.” It represents “barely noticeable,” not “works great.” Even if an intervention crosses the MCID, it only means it’s “worth considering”—it’s a long way from “strongly recommended.” (The ScienceDirect article “MCID Is a Low Bar” linked at the end says this very bluntly.)

Three Numbers You Must Always Look At Together

I often tell my athletes that when reading research, you need to hold three questions in your mind simultaneously—none can be skipped:

The Question You Should Ask Corresponding Statistical Concept What Happens If You Ignore It
“Is this difference real, or just luck?” p-value / confidence interval Mistaking noise for signal
“How big is this difference?” Effect size (Cohen’s d, etc.) Celebrating 1 watt as if it were 30
“Is this difference big enough for me to feel?” Minimal important difference (MCID) Spending big money on something you can’t even perceive

If you only look at the first one, you’ll end up like A-Kai, ready to pull out your wallet at the sight of “significant.” Look at all three together, and you’ll have the ability to make cost-effective training decisions.

(For the definition of MCID and the discussion of it being a low bar, see the PubMed review and ScienceDirect links at the end.)

A Comparison Table to Understand “Four Scenarios”

Cross the two dimensions of “statistically significant or not” and “effect size has practical meaning or not,” and you get four completely different scenarios. I often draw this table on the whiteboard for my athletes because it clears up the most misunderstandings in one go:

Large effect, practically meaningful Small effect, not practically meaningful
Statistically significant Ideal: genuinely effective and noticeable, worth adopting Trap: large sample turns “almost no difference” into significant—the most misleading cell
Not significant A pity: might actually work but sample too small to detect it, worth further research Clear: neither confident nor meaningful, skip it entirely

The most dangerous cell is the top-right—statistically significant, but the effect is too small to matter. This is exactly what large-sample studies produce most often, and what marketing most frequently cherry-picks. From now on, when you see “large-scale study confirms significant effect,” you should be even more careful about asking how big the effect actually is. The bottom-left cell reminds us: “not detecting significance” doesn’t equal “proving it doesn’t work” —it might just mean the sample wasn’t big enough, or we haven’t looked closely enough yet.


Practical Methods: Translating Research into Training Decisions

Theory is done. I know what you really want is “so how do I actually use this.” Here’s the translation process I teach my athletes in practice.

Step One: First Ask, “What Does This Effect Translate to for Me?”

Studies often report results in percentages or standardized numbers. The first thing you need to do is convert that into your own absolute numbers.

Here’s a concrete example (numbers are illustrative scenarios, not citations of specific studies): suppose a study says a certain interval training method improved subjects’ VO2max-related metrics by “an average of about 3%.” Sounds like not much? Let’s convert. Say you currently ride a familiar climb (like Eighteen Peaks Mountain in Hsinchu or Fengguizui) in about 20 minutes:

Your Baseline Estimated After 3% Improvement Practical Meaning
20:00 climb About 30 to 40 seconds faster Noticeable; you won’t get dropped on weekend group rides
250 W FTP About 7 to 8 watts higher Usually exceeds measurement error, worth pursuing
40-minute 10K run About 60 to 70 seconds faster A clear improvement for an amateur runner

See that? The same “3%” suddenly becomes concrete and judgeable once you convert it to scenarios you know. This step is the key to grounding abstract research in reality.

Step Two: Compare Against “Measurement Error” and “Daily Fluctuation”

This is the step most people miss. Every measurement has error, and your body’s daily state fluctuates too. If a claimed effect is smaller than your own normal daily fluctuation, it’s meaningless in the real world.

I usually have my athletes do “self-variability” homework first: for several consecutive weeks, under the same conditions (same route, similar sleep, similar temperature), record the same performance metric and see how much it naturally swings.

Metric Typical Natural Daily Fluctuation Range (highly individual, conceptual reference only) Interpretation Principle
Morning resting heart rate A few bpm up or down Don’t panic over a single high day; look at trends
Same climb time On the order of tens of seconds Anything smaller than this isn’t progress or regression
Power meter readings Usually a few percent error Can’t directly compare before and after changing bikes or head units

The principle is simple: if an intervention’s effect drowns in your own daily noise, then no matter how impressive the study sounds, it doesn’t count for you.

Step Three: Estimate the “Return on Investment”

Even if the effect truly exceeds the noise and the MCID, you still have to ask one final question: is the cost worth it?

I once coached a rider named Xiao-Lin who was preparing for the amateur category of Taiwan’s KOM (King of the Mountain) Challenge. At one point he wanted to try several supposedly “effective” methods all at once. I had him make a very low-tech but incredibly useful table:

Intervention Estimated Effect Magnitude Cost Required My Recommendation
Regular structured training (periodization) Large Time, discipline First priority; the most solid effect
Sleep and recovery management Medium to large Lifestyle adjustments High return; many people underestimate it
Losing excess weight (within a healthy range) Medium to large (especially for climbing) Dietary discipline Extremely effective for climbing events
Certain marginal supplements Small Money, hassle Consider only after the first three are done

This table woke him up: he’d been fixated on things with “small effects but real money and effort costs,” while ignoring the “large effects that just require commitment”—sleep and periodized training. When prioritizing decisions, always eat the fruit with the biggest effect and lowest cost first.

Step Four: Run a “Small-Scale Validation” on Yourself

Research gives you group averages, but the real answer has to be found on your own body. I teach my athletes a very down-to-earth but effective self-validation method I call the “N-of-1 trial”:

  1. Establish a baseline first. Pick a test you can reliably repeat—e.g., the same climb, the same warm-up, as similar temperature and sleep as possible. Measure it several times in a row to capture your natural range “without changing anything.”
  2. Change only one thing at a time. If you want to test an intervention, change only that one variable and keep everything else identical. Change three things at once and you’ll never know which one is doing the work (or canceling the others out).
  3. Give it enough time. Many physiological adaptations take weeks to appear. Don’t give up after two days without feeling anything, and don’t take two days of feeling good as proof (that might just be a good day).
  4. Compare against your baseline fluctuation, not against yesterday. A change only counts if it exceeds your natural fluctuation range.

The spirit of this method is to take the concepts of “effect size” and “daily noise” and run them directly on yourself. It’s not perfect (no control group, placebo effects exist), but for the general exercising population, it’s far more reliable than blindly trusting advertisements.

A Concrete Conversion Example: Caffeine and a Time Trial

Let me walk through the process with another common question (numbers are conceptual illustrations, not citations of specific studies). Suppose someone tells you “caffeine can improve endurance performance by about 2%.”

Interpretation Step Applied to You
Convert to absolute value In a 40-minute time trial, 2% is roughly 40-plus seconds faster
Compare to daily fluctuation If your natural fluctuation under the same conditions is about 20 to 30 seconds, then 40 seconds “might” rise above the noise
Does it exceed the MCID? If you care about your placing, 40 seconds is noticeable; if you’re purely riding for leisure, maybe it doesn’t matter
Cost and risk Cheap and convenient, but watch out for heart palpitations, GI issues, sleep disruption, and individual tolerance differences
Individual response Some people respond strongly, some barely at all, and some even feel unwell

The conclusion won’t be “you must use it” or “absolutely don’t use it,” but rather: “This is an option with a medium effect size, low cost, but high individual variability—worth trying once on a small scale yourself before deciding.” That’s what mature interpretation looks like—not dogmatic, evidence-based, and actionable.


Common Mistakes and Corrections

In all my years coaching, I’ve seen too many people stumble on “interpreting numbers.” Here are the most common mistakes, along with the corrections.

Mistake One: Treating “Significant” as “Large Effect”

This is the number-one problem. As mentioned earlier, “statistically significant” only means “probably not luck”—it has nothing to do with “how big the effect is.”

Correction: Every time you see “significant improvement,” immediately ask yourself—“Improved by how much? Translated to me, how many seconds, watts, or kilograms?” If they can’t give you a concrete number, or the number is laughably small, you know to take it with a grain of salt.

Mistake Two: Only Looking at the Average, Ignoring Individual Differences

Studies report the “average effect,” but you are not the average person. For the same intervention, some people respond strongly and others barely at all—this is called “responder variability.” An average improvement of 3% might be the result of half the people improving 6% and half not improving at all.

Correction: Treat research as “a hypothesis worth trying,” not “a guaranteed promise.” The real validation is a small, documented trial on yourself (more on this below).

Mistake Three: Applying Perfect Lab Conditions to Real Life

Studies are often conducted in highly controlled environments—subjects are well-rested, caffeine-free, on standardized diets. What about your real life in Taiwan? Overtime work, eating out for a pork chop bento, riding the riverside after work, summer heat so humid you’re drenched in sweat. The effects measured in the lab often shrink in your chaotic daily life.

Correction: Mentally discount study effects by default, especially for interventions that depend on “perfect execution.” An effect that survives real-world chaos is a good effect.

Mistake Four: Getting Emotionally Amplified by “Relative Numbers”

“Risk reduced by 50%!” sounds terrifying, but if the original absolute risk was 0.002%, cutting it in half to 0.001% makes virtually no practical difference. Sports marketing also loves relative numbers (“improved 30%!”) to amplify the feeling, without telling you what the baseline was.

Correction: Always look for “absolute numbers.” A relative percentage without a baseline value is just being dishonest.

Mistake Five: Ignoring the Width of the Confidence Interval

Next to the “point estimate” of an effect size, there’s usually a confidence interval. If that interval is wide (e.g., spanning from “almost no effect” to “large effect”), the result is highly uncertain—especially common in small-sample studies.

Correction: When you see a small sample, no reported confidence interval, or an absurdly wide one, dial down the certainty of the conclusion.

Mistake Six: Ignoring “Who the Subjects Were” for Transferability

Many studies use subjects who aren’t the same kind of person as you. Effects found in untrained beginners often shrink dramatically when applied to athletes who’ve trained for years—because experienced athletes have much less “room for improvement” to begin with (this is what sports science calls the effect of “training status”). Conversely, studies done on professional athletes may not apply to us amateur working folks either.

Correction: When reading a study, ask “Who were the subjects? Are they like me?” If the subjects differ greatly from you in training level, age, sex, or lifestyle, discount the transferability. A common situation in Taiwan: tons of marketing cites data from young foreign athletes, but you’re in your forties, eating out every day, and only riding on weekends—applying that data directly will only lead to disappointment.


Taiwan-Specific Context: Grounding the Concepts

After all this theory, I want to highlight a few practical points that Taiwanese readers are most likely to trip over, because our living conditions are far from the lab.

Humid Heat Will Eat Your Effect Size

Taiwan’s summers are humid and hot, and this has a huge impact on endurance performance. Many foreign studies produce beautiful effects measured in cool environments, but when you ride the riverside on a July or August afternoon, heat stress alone is enough to swamp that intervention’s effect. You have to factor environmental variables into your interpretation—the difference in performance between winter and midsummer for the same training may be far larger than any marginal supplement’s effect. This is also why summer demands more attention to hydration, electrolytes, and choosing the right time of day (early morning or evening)—these are the “large effect” fundamentals.

The Nutritional Reality of Eating Out

Study subjects often eat standardized diets, but you might have a pork chop bento for lunch and braised pork rice with bubble tea for dinner. When an intervention’s effect is premised on “good dietary control,” it will shrink when applied to someone who eats out all the time. Instead of obsessing over some marginal supplement, first shore up the bigger-effect basics: “Am I getting enough protein at every meal? Am I eating enough carbs to support my training?”

Easy Healthcare Access Is Your Advantage

Taiwan’s National Health Insurance is highly accessible and convenient—this is actually a huge advantage when interpreting health information. When you have doubts about a supplement’s safety, a physical warning sign, or whether an intervention suits you, the cost of making an appointment to ask a doctor, physical therapist, or dietitian is very low. Don’t let online statistics replace a professional individual assessment. Especially if you have a chronic condition (such as diabetes, hypertension, or heart-related issues), any training or nutritional adjustment should first be discussed with your primary care physician. I’m deliberately being conservative here—these situations must be individualized and require medical attention; they’re not something you can decide on your own by reading articles.


Actionable Advice for Readers at Different Levels

Understanding the concepts is one thing; using them is another. I’ve divided the advice into three levels—find where you fit.

If You’re a Beginner Just Starting to Exercise

You don’t need to read papers or calculate Cohen’s d right now. What you need is to build the right “priority intuition”:

  • Put the big rocks in the jar first. Regular exercise, sleep, basic nutrition, gradual progression—these “largest effect size” fundamentals matter far more than any trick. Don’t get distracted by supplement and gear ads before you’ve even laid the foundation.
  • When you see “significant” or “proven effective” marketing, take a deep breath. Ask: “What exactly is the effect, in concrete terms? What does it cost?” If they can’t answer, set it aside.
  • Record your own baseline. Even just noting your ride time and how you felt in your phone’s notes app—three months later, your own data is more relevant to you than any study.

If You’re an Experienced, Advanced Athlete

You already have a training foundation. The next step is improving your “interpretive ability”:

  • Learn to convert study results into your own absolute numbers (see Step One of the practical methods). This is the translation skill advanced athletes should practice most.
  • Build your “personal MCID” concept. Think clearly: for your goals at this stage, how much does performance need to improve before it’s worth chasing? For example, if you care about breaking a PB, first calculate the seconds needed to break it, then evaluate which interventions have effect sizes that can reach it.
  • Use confidence intervals and sample size as filters. Small-sample, wide-interval studies should be treated as “interesting leads” at most—don’t overhaul your entire training plan over them.

If You’re a Coach or Team Leader

Every recommendation you make affects a group of people, so the responsibility is heavier:

  • When communicating with athletes, use absolute numbers they understand, not a pile of p-values and d-values. “This approach might make you about 20 seconds faster on that climb” is far more motivating than “effect size 0.4.”
  • Manage expectations. Remind athletes that studies report averages and individual responses vary; treat new methods as “hypotheses worth testing” and use small-scale trials to see who’s a responder.
  • When prioritizing decisions, always ask about effect size and cost first. Resources are limited; invest them where the effect is largest and most solid (training structure, recovery, basic nutrition), rather than chasing a bunch of marginal gimmicks.

A Complete Case Study: Walking Through the Entire Mindset

Let me close with A-Kai’s story to tie all the concepts together.

After that day at the coffee shop, I didn’t just tell him not to buy the supplement. Instead, I walked him through the full process.

Step one, translate the effect. We found the rough numbers behind that marketing claim, converted them to his situation, and estimated the effect was “maybe a few seconds faster” on the riverside time-trial segment he rides regularly.

Step two, compare against noise. I asked him to pull up his records for the same segment over the past two months. His own natural fluctuation was tens of seconds—far larger than the few seconds the supplement claimed to provide. The conclusion was clear: that effect drowned directly in his daily fluctuation.

Step three, calculate the return. That supplement cost a fair amount per month and required remembering to take it daily. But when we reviewed his life together, we found he’d been sleeping less than six hours a night and often stayed up late binge-watching shows the week before races.

I told him: “Instead of spending money on a few seconds you can’t even feel, why not start by getting your sleep from six hours to seven and a half? That effect size, I guarantee you’ll feel it on the climbs.”

Three months later, without taking any new supplement—only adjusting his sleep and training structure—he improved his riverside PB by nearly a minute. The first thing he said when he came back was: “Coach, from now on, whenever I see the word ‘significant,’ I’m going to ask: how many seconds, exactly?”

At that moment, I knew he truly got it. That’s what I most want to give you with this article—not some specific conclusion, but a way of thinking that will serve you for a lifetime.


FAQ

Q: Are the effect size thresholds (0.2 / 0.5 / 0.8) absolute standards?
A: No. They’re rough cross-disciplinary conventions proposed by Cohen, and he himself emphasized that context matters. In elite sport, a very small effect size can decide victory or defeat; in general fitness, you might need a larger effect to justify changing your habits. Thresholds are references, not scripture.

Q: So do I need to calculate Cohen’s d myself from now on?
A: Absolutely not. What you need to learn is the “mindset,” not the “calculation”—when you see a conclusion, ask “What does this translate to for me, does it exceed my daily fluctuation, and is it worth the cost?” These three questions are more practical than any formula.

Q: The study says it works, but I tried it and felt nothing. Is something wrong with me?
A: Probably not. Studies report averages, and individual responses vary greatly—you might simply be a “low responder.” That’s exactly why small-scale self-validation is so important. Don’t doubt your own body because of one paper.

Q: Then should I just ignore research entirely and rely on how I feel?
A: That’s also wrong—it goes to the other extreme. Research is valuable “group-level evidence” that can help you filter out a lot of unfounded folk remedies. The right approach: use research for direction, use your own records for validation—the two complement each other.

Q: Can I judge health- or injury-related interventions this way on my own?
A: The concepts can help you understand information, but when it comes to disease, injury, medication, or supplement safety, you must seek professional individual assessment. Statistical thinking helps you read numbers, but it can’t replace medical judgment. Taiwan has convenient healthcare with highly accessible NHI—when in doubt, making an appointment to consult a doctor, physical therapist, or dietitian is always the safest course.

Q: Is effect size the same thing as “correlation coefficient”?
A: Not exactly, but the correlation coefficient (like r) is also a type of effect size metric, used to describe the strength of association between two things. The key point is the same: don’t just look at “whether there’s an association”—look at “how strong the association is and whether it’s practically meaningful.” And correlation doesn’t equal causation—two things changing together doesn’t mean one causes the other. That’s another big pitfall that’s often misused.

Q: Why are some studies with large effects not widely adopted?
A: There could be many reasons: small samples, failure to replicate in other studies, flawed experimental design, or the effect only holds under very specific conditions. No matter how impressive a single study is, it’s just “one piece of evidence.” Truly reliable conclusions usually require multiple independent studies showing consistent directions across different populations. When you see “a single study’s shocking discovery,” be conservative.

Q: There’s so much to remember. What if I can’t keep track as a regular person?
A: You don’t need to remember it all. You just need to develop one reflex—every time you see “effective,” ask yourself three questions: “How big is the effect (converted to my seconds or watts)? Does it exceed my daily fluctuation? Is it worth the cost?” These three questions condense ninety percent of this article’s value.


Conclusion: Become Someone Who Can Question Numbers

We live in an era bombarded by numbers. Open your phone and it’s full of “research confirms,” “significant improvement,” “amazing results.” These phrases have a kind of magic that makes people believe without thinking.

But after reading this, I hope you have a layer of immunity. You know “significant” doesn’t equal “large effect,” you know to translate research into your own absolute numbers, you know to compare against daily noise and the MCID, and you know to calculate the return on investment.

The most formidable athletes aren’t the ones who read the most research—they’re the ones who ask the best questions. The next time someone brings you an article “proving effectiveness,” you won’t have your eyes light up like A-Kai used to. Instead, you’ll smile and ask:

“So, what exactly is the effect, in concrete terms? And for me, does it really matter?”

That question is the watershed between being led around by numbers and mastering them.

One last thing I want to emphasize: developing this mindset isn’t about becoming cynical or believing in nothing. Quite the opposite—it’s about investing your limited time, money, and recovery capacity where it’s truly worth it. When you stop being distracted by a bunch of marginal effects, you can focus more deeply and steadily on the few things with the largest effect sizes (consistent training, adequate sleep, solid nutrition, gradual progression) and do them well, for the long haul. That’s the real engine of long-term progress.

Numbers are tools, not masters. Understand them, question them, and then put them to work for you—that’s the posture of a mature athlete.

May you ride every climb, every riverside path, and every race in Taiwan with more understanding and more wisdom than yesterday. See you out on the road.


This article is educational content and does not replace individual diagnosis and treatment advice from a physician, physical therapist, or dietitian. For issues involving disease (such as diabetes, hypertension, heart disease, etc.), injury, or supplement safety, please be sure to seek individualized assessment and assistance from professional medical personnel.

References

相關影片
訂閱CT的頻道

訂閱 CT Yeh,看武嶺實測與路線攻略

北進武嶺、西進武嶺、經典百K,每條路線都親自騎過,配速、爬升、補給點全部實拍實測。

467 部影片 · 累計 838 萬次觀看