OUR EVIDENCE
Every claim has a source. Here they are.
Most performance brands tell you what works. Few show you why they believe it. Fewer still tell you what does not work, and why they would not recommend it even if it would sell.
This page is where SPMD shows its work.
Every framework element, every recommendation, and every clinical claim at SustainablePerformanceMD is grounded in published literature. The evidence is graded honestly. Strong evidence is labeled strong. Promising but limited evidence is labeled accordingly. Overhyped interventions are named and explained. The conflict of interest question is asked on every source.
That is what the fiduciary standard looks like in practice.
WHAT IS SPMD’S EVIDENCE GRADING STANDARD?
SPMD ranks evidence from Strong to Overhyped.
Strong Evidence: Multiple well-designed randomized controlled trials with adequate sample size, appropriate controls, and reproducible outcomes. Consistent findings across independent research groups. Meta-analyses available.
Moderate Evidence: Randomized controlled trials with methodological limitations, smaller samples, or limited replication. Mechanistically plausible with supportive observational data.
Promising but Limited: Early-stage research, small samples, pilot studies, or strong mechanistic rationale without definitive clinical trials. Worth monitoring. Not yet a firm recommendation.
Overhyped: Widely promoted with weak, conflicted, or misrepresented evidence. Often profitable. Rarely delivers on the claim. SPMD names these and explains why.
Personal and Operational Evidence: Case series, clinical observation, and documented operational experience. A legitimate level of the evidence hierarchy, and a low one. It is not a substitute for peer-reviewed literature, but it is not nothing either, and it is always labeled as what it is.
This scale grades sources. The MOVES Database uses a separate scale that grades tools, because a tool can rest on strong evidence and still be a poor fit, and a tool can be worth trying while its evidence question stays open.
THE 5P PERFORMANCE DOMAINS: OUR EVIDENCE BASE
Physical Readiness
The physical substrate governs everything downstream. No cognitive, psychological, or behavioral intervention outperforms chronic sleep deprivation and physical inactivity as performance impairments. This is the organizing clinical principle behind Physical Readiness as the first of the five domains.
Sleep and cumulative cognitive impairment: Van Dongen HPA, et al. The cumulative cost of additional wakefulness: dose-response effects on neurobehavioral functions and sleep physiology. Sleep. 2003. (PMID12683469). Evidence grade: Strong. Key finding: partial sleep restriction to six hours per night for two weeks produced cognitive deficits equivalent to two full nights of total deprivation, while subjects did not accurately perceive their own impairment level.
Cardiorespiratory fitness and all-cause mortality: Kodama S, et al. Cardiorespiratory fitness as a quantitative predictor of all-cause mortality and cardiovascular events in healthy men and women. JAMA. 2009. Evidence grade: Strong. The relationship between VO2 max and mortality is among the most consistent dose-response findings in preventive medicine.
Exercise and cognitive performance: Hillman CH, et al. Be smart, exercise your heart: exercise effects on brain and cognition. Nature Reviews Neuroscience. 2008. Evidence grade: Strong. Aerobic exercise directly improves prefrontal cortex function, executive attention, and cognitive control across the lifespan.
Resistance training and longevity: Stamatakis E, et al. Associations of strength training with all-cause, cardiovascular disease, and cancer mortality in US older adults. British Journal of Sports Medicine. 2022. Evidence grade: Strong. Two or more sessions per week of muscle-strengthening activity is associated with significant reduction in all-cause and cancer mortality independent of aerobic activity.
Exercise and sleep quality: Li L, et al. Front Psychol. 2024. PMCID: PMC11484100. Yoga and combined training modalities produced the largest improvements in sleep quality across populations. Moderate aerobic exercise consistently improves sleep onset latency and slow-wave sleep duration. Evidence grade: Moderate-Strong.
Prefrontal Operations
The prefrontal cortex is disproportionately sensitive to sleep debt compared to other brain regions. Decision-making, working memory, and executive function degrade before subjective awareness of impairment does. This is the clinical mechanism behind Prefrontal Operations as a distinct performance domain, not a synonym for intelligence or personality.
Sleep deprivation and prefrontal function: Harrison Y, Horne JA. The impact of sleep deprivation on decision making. Journal of Sleep Research. 2000. Evidence grade: Strong. Innovative thinking and flexible decision-making are preferentially impaired before rote task performance degrades, meaning high performers lose their highest-value cognitive functions first.
Decision fatigue and the willpower model: Carter EC, Kofler LM, Forster DE, McCullough ME. A series of meta-analytic tests of the depletion effect: self-control does not seem to rely on a limited resource. Journal of Experimental Psychology: General. 2015;144:796-815. Hagger MS, Chatzisarantis NLD, Alberts H, et al. A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science. 2016;11(4):546-573. Vohs KD, Schmeichel BJ, Lohmann S, et al. A multisite preregistered paradigmatic test of the ego-depletion effect. Psychological Science. 2021;32(10):1566-1581. Danziger S, Levav J, Avnaim-Pesso L. Extraneous factors in judicial decisions. PNAS. 2011;108(17):6889-6892. Weinshall-Margel K, Shapard J. Overlooked factors in the analysis of parole decisions. PNAS. 2011;108(42):E833. Glockner A. The irrational hungry judge effect revisited: simulations reveal that the magnitude of the effect is overestimated. Judgment and Decision Making. 2016. Evidence grade: Overhyped for the model. Promising but Limited for the phenomenon.
The ego depletion model holds that self-control runs on a limited shared resource that empties with use. It has not survived testing. Carter 2015 found substantial small-study bias across the published literature, with a corrected effect near zero. Hagger 2016 ran a preregistered replication across 23 laboratories and 2,141 participants: effects were trivial, and for most labs the confidence interval included zero. Vohs 2021 tried again across 36 laboratories and 3,531 participants, using a design built to give the effect its best available shot. The confirmatory result was d = 0.06, nonsignificant, with a Bayesian analysis finding the data four times more likely under the null than under the hypothesis.
The parole-board study most often cited as proof of decision fatigue carries a live confound rather than a limitation. Case ordering was not random: prisoners without attorneys were typically heard last and were less likely to be granted parole regardless of the hour. Glockner's simulations concluded the effect magnitude has been substantially overestimated.
What survives is narrower and real. In Vohs 2021, participants reporting more fatigue after the first task performed worse on the second. Cognitive load and working memory limits are separately and robustly established, below. Fatigue is doing something. It is not behaving like a fuel gauge.
The honest reading is that the resource model is unsupported rather than disproven, and its proponents have not specified a procedure that reliably produces the effect in a large sample. Baumeister and Vohs have argued the replication paradigms were poor tests of the theory, and that objection is on the record.
SPMD does not use the willpower-as-resource model. SPMD does not cite the parole study as evidence of anything. Where SPMD states that a DRAINS is a mechanism rather than a discipline problem, this is the literature behind that claim, including the part that constrains it.
Working memory limits: Cowan N. The magical number 4 in short-term memory: a reconsideration of mental storage capacity. Behavioral and Brain Sciences. 2001;24:87-185. Cowan N. The Magical Mystery Four: how is working memory capacity limited, and why? Current Directions in Psychological Science. 2010;19(1):51-57. (PMC4673075). Evidence grade: Strong. Working memory capacity is limited to approximately four chunks of information in active processing simultaneously. This is the clinical justification for Limit applied to Avalanche.
Mental rehearsal and prefrontal preparation: Gabbott B, et al. BJS Open. 2020. PMCID: PMC7709374 Evidence grade: Moderate-Strong. Mental simulation activates the same neural circuits as physical execution, supporting its use in HOPE architecture as a pre-loading mechanism.
Psychological Flexibility
Psychological flexibility is the capacity to contact the present moment fully, as a conscious human being, and to change or persist in behavior when doing so serves valued ends. It is not positive thinking. It is not emotional suppression. It is the ability to carry difficult internal experience while continuing effective action. SPMD trains it as a clinical target, not a personality trait.
ACT meta-analysis: Gloster AT, et al. The empirical status of acceptance and commitment therapy. Psychological Research and Behavior Management. 2020. Evidence grade: Strong. ACT demonstrates efficacy across anxiety, depression, chronic pain, substance use, and work performance outcomes.
Psychological flexibility as mechanism: Levin ME, et al. The impact of treatment components suggested by psychological flexibility theory. Behaviour Therapy. 2012. Evidence grade: Moderate. Acceptance and defusion components independently mediate outcomes beyond exposure alone.
Mindfulness and physiological stress markers: Heckenberg RA, et al. Do workplace-based mindfulness meditation programs improve physiological indices of stress? Journal of Psychosomatic Research. 2018. Evidence grade: Moderate. Structured mindfulness programs produce measurable reductions in cortisol and blood pressure in occupational settings. Effect sizes are real and clinically meaningful, not transformative.
Personal Systems
Environment and system design are more reliable performance levers than motivation. Motivation is a state. Systems are architecture. The evidence for implementation intentions and friction-based behavior design is among the most consistent in behavioral science.
Implementation intentions: Gollwitzer PM, Sheeran P. Implementation intentions and goal achievement: a meta-analysis of effects and processes. Advances in Experimental Social Psychology. 2006. Effect size d = 0.65. Evidence grade: Strong. If-then planning (when X occurs, I will do Y) significantly increases goal follow-through across health behaviors, academic performance, and occupational tasks.
Habit formation timelines: Lally P, et al. How are habits formed: modelling habit formation in the real world. European Journal of Social Psychology. 2010. Evidence grade: Moderate. Habit automaticity develops across a range of 18 to 254 days depending on behavior complexity and individual factors. The 21-day claim has no empirical basis. SPMD does not cite it.
Choice architecture and behavior design: Thaler RH, Sunstein CR. Nudge: Improving Decisions About Health, Wealth, and Happiness. Yale University Press. 2008. Mertens S, Herberz M, Hahnel UJJ, Brosch T. The effectiveness of nudging: a meta-analysis of choice architecture interventions across behavioral domains. PNAS. 2022;119(1):e2107346118. Correction for Mertens et al. PNAS. 2022;119(19):e2204059119. Maier M, Bartos F, Stanley TD, Shanks DR, Harris AJL, Wagenmakers EJ. No evidence for nudging after adjusting for publication bias. PNAS. 2022;119(31):e2200300119. Szaszi B, et al. No reason to expect large and consistent effects of nudge interventions. PNAS. 2022;119(31):e2200732119. Mertens S, et al. Reply to Maier et al., Szaszi et al., and Bakdash and Marusich: the present and future of choice architecture research. PNAS. 2022;119(31):e2202928119. Evidence grade: Reference for the book. Promising but Limited for the mechanism. Not Strong.
Nudge is an accurate popularization of the choice architecture literature. It is not research, and a book cannot carry a grade this scale was built to assign.
The literature underneath it is contested in a way worth naming precisely. Mertens 2022 meta-analyzed 447 choice architecture experiments and reported an overall effect of d = 0.40, concluding that nudging is an effective and widely applicable behavior change tool. The same paper reported moderate publication bias in its own data. PNAS subsequently issued a correction: the analyzed sample had included four observations from a paper that has since been retracted, along with erroneous values from a second paper and several coding errors. On the corrected data, adjusting for moderate publication bias attenuates the effect by 22.5 percent, from d = 0.40 to d = 0.31. Adjusting for severe publication bias attenuates it to d = 0.08.
Three independent critique letters followed. Maier 2022, applying robust Bayesian meta-analysis, found no evidence for nudging after adjusting for publication bias. Szaszi 2022 concluded that nudge interventions may work under conditions the literature has barely identified. In their published reply, Mertens and colleagues acknowledged that these methods report smaller effect sizes or even null effects.
Note the shape of that sequence, because it is not unique. A popular construct, a headline meta-analysis, a publication bias correction that collapses the effect, and proponents arguing the correction methods are wrong. That is the same sequence as ego depletion, above. SPMD grades both the same way, for the same reason, and would grade its own work the same way if it came apart like that.
What survives: specific, tested applications. Defaults and friction reduction have their own evidence in their own contexts. What does not survive is choice architecture as a general theory of behavior change, and SPMD does not use it as one.
Procedural Competency
Skill is acquired through deliberate practice, not time on task. The distinction matters clinically. Performing a task and improving at a task are not the same activity. SPMD treats Procedural Competency as a trainable performance domain with a defined evidence base, not a fixed attribute.
Deliberate practice: Ericsson KA, et al. The role of deliberate practice in the acquisition of expert performance. Psychological Review. 1993. Evidence grade: Strong for near-transfer effects. Key limitation: meta-analysis shows deliberate practice explains approximately 26% of performance variance in games and 21% in music. [VERIFY: this is Macnamara BN, Hambrick DZ, Oswald FL. Deliberate practice and performance in music, games, sports, education, and professions: a meta-analysis. Psychological Science. 2014. Confirm and write the citation in. On the page carrying the 10,000-hour correction, the meta-analysis that carries it should be named.] The 10,000-hour rule as commonly stated is a popularization that removed all the conditions that made the original finding meaningful. SPMD does not cite that figure.
Growth mindset in context: Yeager DS, et al. A national experiment reveals where a growth mindset improves achievement. Nature. 2019. Evidence grade: Moderate. Effect sizes are real and statistically significant (d = 0.14 for academic achievement). Standalone mindset interventions without behavioral architecture produce minimal results. Mindset work in SPMD is always paired with a procedural change, not offered as a standalone.
Multimodal delivery and performance outcomes: Pearce et al. Annals of Behavioral Medicine. 2023. 865,000-plus participants. [VERIFY: no title. Complete the citation before publishing.] Evidence grade: Strong. Multimodal delivery formats produce superior adherence and outcome compared to single-channel delivery.
THE DRAINS AND CLEARS: OUR EVIDENCE BASE
The DRAINS taxonomy is derived from the clinical and behavioral science literature on performance-impairing states. Each CLEARS is matched to its DRAINS based on the mechanism literature for that specific state. The matching is not intuitive or conventional-wisdom-based. It follows the clinical mechanism.
Drift and Clarify
Drift is the progressive misalignment between stated values and actual daily behavior, occurring without deliberate choice and often without awareness. It is not laziness. The clinical mechanism is values-behavior incongruence accumulating below the threshold of conscious monitoring.
Values-behavior gap: Hayes SC, Strosahl KD, Wilson KG. Acceptance and Commitment Therapy: The Process and Practice of Mindful Change. Guilford Press. 2012. Evidence grade: Textbook, authoritative for the model. Written by ACT's originator; cited here for the model itself, not as trial evidence. Trial evidence for ACT is Gloster 2020, above. Clarification of values-behavior discrepancy is a primary ACT intervention mechanism.
Goal-setting as clarification: Locke EA, Latham GP. Building a practically useful theory of goal setting and task motivation. American Psychologist. 2002. Evidence grade: Strong. Specific, challenging goals with feedback outperform do-your-best goals across task types and populations.
Resistance and Subtract
Resistance is friction, avoidance, and activation energy mismatch between intention and execution. The environment generates it. The clinical intervention reduces the friction, not the person's willpower.
Behavior design and friction: Fogg BJ. Tiny Habits. Houghton Mifflin Harcourt. 2019. Evidence grade: Reference. An accurate popularization of behavior design work, not research. Motivation is unreliable; ability and prompt are the clinical levers.
Friction reduction architecture: Thaler and Sunstein, above, at the grade recorded there. Friction reduction is one of the specific applications that carries its own evidence, distinct from choice architecture as a general theory.
Avalanche and Limit
Avalanche is cognitive overload: the simultaneous activation of more demands than working memory and executive function can process without degraded output. The CLEARS is Limit, which reduces the active demand set to match cognitive capacity.
Cognitive load theory: Sweller J. Cognitive load during problem solving: effects on learning. Cognitive Science. 1988. Evidence grade: Strong. Working memory limitations impose a hard ceiling on simultaneous information processing. Schema formation requires load management, not willpower.
Working memory limits: Cowan, above.
Identity Lock and Anchor
Identity Lock is rigidity in self-concept that prevents behavioral adaptation when adaptation is required. It presents as the conviction that one cannot change a particular behavior because it is too central to who they are. The CLEARS is Anchor, which does not challenge the identity but expands it.
Identity-based behavior change: Oyserman D, Elmore K, Smith G. Self, self-concept, and identity. Handbook of Self and Identity. 2012. Evidence grade: Review chapter, Moderate. A synthesis of the identity literature, not a primary study.
ACT self-as-context: Hayes et al., above. Psychological flexibility includes the capacity to hold self-concept lightly enough to permit behavioral change without identity threat.
Nerve Failure and Execute
Nerve Failure is action inhibition in the presence of adequate knowledge and intention. The person knows what to do and intends to do it. They do not start. The clinical mechanism is prefrontal initiation deficit, often potentiated by anticipated negative outcome or perfectionism. The CLEARS is Execute, which bypasses the initiation threshold.
Behavioral activation: Martell CR, et al. Behavioral Activation for Depression. Guilford Press. 2010. Evidence grade: Strong. Activation precedes and generates motivation. The clinical sequence is action first, not motivation first.
Prefrontal initiation and executive function: Miller EK, Cohen JD. An integrative theory of prefrontal cortex function. Annual Review of Neuroscience. 2001. Evidence grade: Strong.
Spent and Reset
Spent is physiological and psychological depletion beyond the normal fatigue that exercise and cognitive work produce. It is a state of reserve deficit, not a motivation problem. The CLEARS is Reset, which targets the parasympathetic recovery pathway.
Overtraining syndrome and recovery: Meeusen R, et al. Prevention, diagnosis and treatment of the overtraining syndrome. European Journal of Sport Science. 2013. Evidence grade: Moderate-Strong.
Parasympathetic activation through controlled breathing: Zaccaro A, et al. How breath-control can change your life: a systematic review on psycho-physiological correlates of slow breathing. Frontiers in Human Neuroscience. 2018. Evidence grade: Moderate.
Heart rate variability biofeedback: Lehrer PM, Gevirtz R. Heart rate variability biofeedback: how and why does it work? Frontiers in Psychology. 2014. Evidence grade: Moderate.
THE MODIFIED SPIRAL PRINCIPLE: OUR EVIDENCE BASE
In 1960, cognitive psychologist Jerome Bruner described the spiral curriculum: a model in which learners return to foundational material multiple times, each pass adding depth anchored to what is already understood. The mechanism is not repetition. It is the recognition that working memory has hard limits, and that sequencing from simple to complex while gradually reducing instructional scaffolding is what makes learning hold rather than overwhelm and evaporate.
The SPMD tier structure is a spiral curriculum. That is not a marketing choice. It is the design principle that was operating correctly before it had a name.
Original source: Bruner JS. The Process of Education. Harvard University Press. 1960.
Spiral curriculum in medical education: Harden RM, Stamper N. What is a spiral curriculum? BMC Med Educ. 2007 Dec 21;7:52. doi:10.1186/1472-6920-7-52.
Working memory limits and learning: Sweller, above. Cowan, above.
Mastery-based learning: Bloom BS. Learning for mastery. Evaluation Comment. 1968. Evidence grade: Strong for sequenced, competency-gated instruction. Mastery at each level before advancing is a validated instructional model.
SPMD's modification to Bruner's original model adds two gates the classroom-pacing version never had: demonstrated use, not calendar exposure, and tiered commitment. You do not advance because time passed. You advance because you have lived inside the current layer long enough for the next one to be actionable.
WHAT SPMD WILL NOT RECOMMEND, AND WHY?
Exogenous ketones for cognitive performance in healthy individuals. The mechanistic rationale is plausible. The clinical evidence for meaningful cognitive benefit in non-ketogenic individuals without neurological impairment is weak. The cost is high. The recommendation does not pass the Honest Objective Science standard.
Proprietary nootropic supplement stacks as a category. Individual components may have modest evidence in isolation. Stack combinations are rarely tested as formulated. Conflict of interest in manufacturer-funded research is significant. Regulatory oversight under DSHEA is minimal. SPMD does not recommend products whose evidence base was funded by the product manufacturer without independent replication.
Unsupervised cold immersion combined with breath-holding. The American Heart Association and British Heart Foundation have issued warnings. Arrhythmia risk in this specific combination is documented. Cold exposure benefits are real and achievable through safer protocols.
The 10,000-hour rule as popularly stated. A popularization of Ericsson's work that discarded all the conditions that made the original finding meaningful. Deliberate practice matters. Hours of practice without structured feedback, defined objectives, and edge-of-ability tasking do not meet the definition used in the original research and do not produce the same outcomes.
THE FIVE SENSES CLINICAL DECISION STANDARD
When strong randomized controlled trial evidence is absent but a clinical decision is still required, SPMD applies the Five Senses Clinical Standard, derived from the five Operational Principles. This is not an abandonment of evidence. It is a structured approach to using it correctly when the evidence hierarchy does not offer a definitive answer.
Does it fit right: Medical Readiness. Can the person's biological foundation support and tolerate this intervention right now.
Does it look right: Clinical Lens. Does it align with established mechanisms. Is the proposed pathway plausible given known physiology and psychology.
Does it smell right: Honest Objective Science. No counter-evidence, no known harm profile, no red flags in the conflict-of-interest audit.
Does it sound right: Psychological Flexibility. Does it fit this individual, their values, their actual life, and what they will realistically execute.
Does it taste right: Process Over Motivation. Does it follow logically from what has already been established and does it produce a next step.
A recommendation clears all five before it reaches a client.
THIS PAGE IS UPDATED
The evidence base is not static. When new research is published that changes a clinical recommendation, this page is updated. When a previously recommended intervention fails replication, this page is updated. The fiduciary standard requires it.
If you have a specific citation question or want to verify a claim made elsewhere on this site, send it to will@sustainableperformancemd.com.
SPMD assesses performance, not health. SPMD is not medical care.
See how the evidence becomes a clinical architecture. Read the Method.