Trust in climate data remains a significant barrier to effective climate action. Skepticism about data manipulation and politicization reduces confidence and hinders evidence-based policy. Existing climate data systems lack transparent verification and accessible analytical tools, limiting accountability and stakeholder engagement. This study presents a reproducible framework that applies blockchain technology to provide transparent verification, analysis, and governance of climate data. The architecture includes three layers: a data ingestion layer that standardizes verified observations, a blockchain layer that ensures immutability and provenance through proof-of-stake consensus, and a statistical analysis layer that uses deterministic methods for anomaly detection and trend evaluation. The framework was tested using 8,403 hours of temperature data from the Manila, Philippines monitoring station during 2024. Analysis identified 33 temperature anomalies ranging from 36.9 to 38.0 °C that aligned with documented April–May 2024 heat waves, confirming the ability to detect genuine meteorological extremes. Estimated transaction latency was 1–2 seconds per observation, with on-chain storage requirements of about 138 kilobytes and off-chain storage requirements of 2.1 megabytes for a 90-day deployment. Estimated energy use for the same period was approximately 0.06 kilowatt-hours, representing a 97–99 percent reduction compared with proof-of-work systems. These findings demonstrate that the proposed framework can securely record, verify, and analyze climate data while consuming very little energy. By combining blockchain immutability with transparent statistical methods, this approach directly addresses the trust deficit in climate science and provides a foundation for verifiable, reproducible, and efficient climate information systems.
Large language models are approaching specialist-level performance in selected clinical tasks, but accuracy alone does not establish readiness for clinical deployment. This commentary argues that accountability, rather than raw performance, is now the central barrier to adoption. Recent evidence shows that clinically deployed large language models can perform radiology workflow tasks with high accuracy, yet important governance questions remain unresolved, including the provenance of inputs, the auditability of outputs, and the verification of downstream decision pathways. The present commentary proposes that accountability infrastructure should become a routine focus of clinical AI evaluation alongside performance metrics. Distributed ledger and related audit technologies may offer one practical framework for tamper-resistant logging, verification, and oversight of model-mediated clinical decisions. Clinical studies should therefore report governance architecture in addition to accuracy, and medical education should treat prompt engineering as an operational clinical competency. The next phase of clinical AI is not merely accurate systems, but accountable ones. Keywords: Clinical AI Governance, Large Language Models, Blockchain Healthcare, Radiology AI, Algorithmic Accountability. Citation: Heston TF. Accountable clinical AI requires more than accuracy. Internet Medical Journal. 2026;1(1):e19519377. doi:10.5281/zenodo.19519377.
Commentary responding to Kahana, Perets, and Emile (Journal of Clinical Epidemiology, 2026; doi:10.1016/j.jclinepi.2026.112255), who propose a bidirectional fragility index (bFI) for 2×2 contingency tables. This paper demonstrates that the bFI, while a genuine improvement over the unidirectional fragility index, inherits structural limitations: (1) it restricts perturbations to within-arm toggles (fixed row margins), which is a subset of the full perturbation space; (2) it is mathematically incomplete (not attainable for some valid tables); and (3) it conflates classification stability with robustness. The Global Fragility Index (GFI) resolves the first two limitations by searching all cell-to-cell reallocations without fixing any marginal totals. Exhaustive enumeration of all valid 2×2 tables (N = 7–30, 45,230 tables) confirms that GFI ≤ bFI always, GFI < bFI in 21.5% of cases, and GFI is attainable for every table in the domain. The Risk Quotient (RQ) addresses the third limitation by providing a normalized robustness metric orthogonal to fragility. Includes the R enumeration code and complete output dataset.
As large language models generate clinical documentation at scale, electronic health records increasingly contain AI-produced content with no verifiable provenance. The field has named this traceability gap but has not yet specified an architectural solution. Blockchain-based audit logging — append-only, cryptographically chained, and tamper-resistant — provides the answer: a lightweight layer capturing model identity, prompt hash, output fingerprint, and timestamp at generation, creating a verifiable chain of custody before text enters the clinical record. Adopting blockchain audit logging as a standard condition of institutional large language model deployment would resolve this traceability crisis before it becomes irreversible.
Fragility indexes attempt to quantify the minimum perturbation that reverses a sample's statistical significance under an unstated constraint: the trial roster is closed, so only the recorded outcomes of enrolled patients may change. The fragility distance (FD) removes that constraint, defining fragility as the L1 edit distance from the observed 2×2 table to the significance boundary, where one unit is one patient added to or removed from any cell. Because every constrained index bounds the FD from above, the published toggle and transfer indexes become special cases of a single geometric quantity. In a published phase 2 trial reporting no significant difference (p = 0.0903), FD = 1: the removal of a single patient reverses the significance classification, a proximity to the boundary that the global fragility index, also equal to 1 but denominated in two-cell moves, understates by half. Reporting FD with its witness table gives the true minimum perturbation, priced in the one unit that occurs in every real trial: the enrollment, or the loss, of a single patient.
Fragility index audits of clinical guideline evidence measure classification stability but not distance from therapeutic neutrality — leaving the most consequential question in evidence grading unanswered. Applied to the randomized controlled trials recently cited in the NCCN guidelines for gastric cancer, the fragility index cannot distinguish a low-powered detection of a real treatment effect from a near-null result that narrowly crossed the significance threshold, a distinction with direct implications for confirmatory trial prioritization and guideline strength ratings. The significance-fragility-robustness framework resolves this gap by adding robustness dimension — a geometrically derived, model-free measure of distance from therapeutic neutrality — as an orthogonal third statistical metric alongside significance and fragility, completing the evidence picture that guideline audits require.citation: Heston TF. Guideline evidence audits require robustness beyond the fragility index. Internet Medical Journal. 2026;1:e19656227
Background: Dietary research commonly assesses multivitamin (MVI) supplement use through broad categorical questions that assume products are compositionally similar. This study quantified how much the micronutrient composition of best-selling MVI products varies, to assess whether treating “multivitamin use” as a single exposure category is defensible. Methods: The 100 best-selling multivitamin products on Amazon, the dominant U.S. online supplement retailer, were screened on July 24, 2026. After excluding children’s, adolescent, prenatal, single-nutrient, and condition-specific products and duplicate listings, 64 adult general multivitamins remained. For each product, the labeled percent Daily Value (%DV) of 25 micronutrients was recorded directly from the Supplement Facts panel. Heterogeneity was characterized along three axes: composition breadth (number of nutrients included), per-nutrient content (inclusion rate, median, interquartile range, and range), and megadosing (defined a priori as ≥1000% DV). Results: Composition breadth ranged from 7 to 24 of 25 nutrients (median 20). Only vitamin D, folate, and vitamin B12 were present in all 64 products, whereas choline, phosphorus, and iron were present in 22% or fewer. Labeled content varied widely; vitamin B12 ranged from 100% to 33,333% DV (median 408%). Thirty products (47%) contained at least one megadosed nutrient, concentrated in the B-complex vitamins and biotin. Conclusion: Best-selling multivitamins vary markedly in which nutrients they contain, in how much, and in whether they megadose. “Multivitamin use” is a compositionally heterogeneous exposure; studies should specify formulations or stratify by composition. Citation: Heston TF. Marked heterogeneity in micronutrient composition of multivitamin-mineral supplements. FNDS [Internet]. 2026 Sep. 16 [cited 2026 Sep. 17];4(1):98-105. Available from: https://ojs.luminescience.cn/FNDS/article/view/604
The Survival-Inferred Fragility Index (SIFI), introduced in 2020 and now applied to first-line metastatic renal cell carcinoma immunotherapy trials, extends the fragility concept into time-to-event oncology outcomes by iteratively reassigning long-surviving patients between trial arms until statistical significance is lost. This reassignment rule is model-dependent: the choice of which survivors to move and in what order embeds distributional assumptions about the survival tail, so SIFI values are path-dependent and therefore have limited cross-trial comparability. The Survival Fragility Quotient, derived from the Cox regression z-statistic and paired with the Survival Robustness Quotient on the neutrality boundary, provides a model-free survival fragility measure without iterative reassignment, requires only the reported hazard ratio and its confidence interval, and completes the p–fr–nb triplet for time-to-event outcomes. Oncology randomized trial reporting should adopt the model-free pair to support cross-trial comparability of the quality of survival evidence.
Apologies to patients are institutionalized in medicine, driven by ethics, training, and apology laws aimed at reducing malpractice risk. Yet apologies among colleagues—clinicians, administrators, and staff—remain rare, despite frequent harms from hierarchy, bullying, and dismissive interactions. Drawing on personal primary care experience and workplace conflict literature, this perspective highlights the cultural and structural barriers to inter-colleague remorse. It critiques existing apology models for lacking prevention and proposes the novel Acknowledge–Repair–Prevent (ARP) framework, which incorporates restorative justice principles to shift from retributive (power-focused) dynamics to shared accountability. Implementing ARP could rebuild trust, mitigate burnout, and foster profound respect for colleagues in a demanding profession. v2: uploaded published article. Citation: Heston T F (January 28, 2026) Moving Beyond Sorry: The Acknowledge-Repair-Prevent (ARP) Framework for Colleague Apologies in Medicine. Cureus 18(1): e102517. doi:10.7759/cureus.102517
Background: The fragility index (FI) is intended to quantify how many outcome changes would be required to convert a statistically significant two-arm trial result into a nonsignificant one. A reliable statistical metric should produce a result for every valid case it evaluates. This study examined whether a fragility value is always attainable for every statistically significant trial result. Methods: FI was analyzed as follows: baseline significance was required (p < 0.05), one-way movement only, and outcome changes were restricted to converting a nonevent to an event in the arm with fewer events, while keeping the arm size fixed. Nonattainability was assessed by determining whether valid 2×2 tables exist for which no finite FI can be obtained under these rules. Evidence is provided through formal counterexamples, complete enumeration of all valid nondegenerate 2 × 2 tables up to total sample size N = 60, and empirical evaluation of published two-arm trials with binary outcomes. Results: Valid baseline-significant 2 × 2 tables exist for which FI is not attainable. A simple counterexample is {3,0,4,11}: baseline two-sided Fisher's exact p = 0.0429, the arm with fewer events is uniquely identified, but that arm has no nonevents available for the required toggle; thus, no legal FI path exists. Enumeration revealed that unattainable cases first appeared at N = 18 and then recurred at every larger sample size through N = 60; by N = 60, a total of 2,390 of 20,774 evaluable baseline-significant tables were unattainable (11.5%). In an empirical dataset of published trials, 2 of 82 baseline-significant evaluable trials (2.4%) were not attainable. Conclusions: The FI is not universally attainable. This is a structural property of the FI algorithm, confirmed by mathematical proof, a complete table enumeration, and published trial data.
Statistical fragility is not a single quantity. Resampling fragility estimates the probability that a trial's significance classification would reverse in a new sample of the same size, while perturbation fragility measures the minimum change to the recorded data that reverses the classification in the observed table. The primary endpoint of the Medically Utilized Tailored Traditional Foods to Optimize Nutrition in Heart Failure (MUTTON-HF) randomized clinical trial carries a resampling fragility of 0.38 under a within-arm binomial replication model, and a global fragility quotient of 0.0097 (unstable). These fragility metrics indicate that the significance classification behind this clinically promising result is easily reversed and therefore warrants cautious interpretation.
The charge that the fragility index is a P-value in disguise rests on a strong correlation between the two across trials, and it overlooks a distinction between two distinct constructs derived from the same observed result. Resampling fragility asks how often the significance verdict would change if the trial were drawn again at the same size; it is a probability, and because it is based on the same observed result and a specified replication model, it is often strongly associated with the P-value. Perturbation fragility asks how many recorded outcomes must change before the observed table crosses the significance threshold; it is a count, the global fragility index, or a proportion, the global fragility quotient, and it measures the geometric distance of the observed table from the decision boundary, a quantity the P-value does not encode. Worked examples from published trials show tables with similar P-values and widely different perturbation fragility. The redundancy critique, whether made by resampling or by machine-learning models that use the P-value as a predictor, is correct about resampling fragility but does not address the fragility index.
Shoulder surgeons regularly choose between biceps tenotomy and tenodesis during rotator cuff repair, and that choice is guided by evidence synthesis built on fragility analysis of the underlying randomized controlled trials. A recent analysis of seven such trials reports a mean reverse continuous fragility index (rCFI) of 17.7 and concludes moderate stability of the noninferiority findings — a conclusion that flows directly into how surgeons counsel patients and how guidelines grade recommendations. The rCFI calculation simulates a reconstructed patient-level data set from published summary statistics and returns an iteration count that changes with the random seed; its normalization, labeled the "reverse fragility quotient," divides simulated transfers by real patient numbers. A measurement that shifts with investigator choices and mixes simulated with real quantities cannot serve as a foundation for clinical decisions. The continuous fragility quotient (CFQ) is a structurally different construction: it computes a deterministic value on the interval from zero to one directly from the Welch t-statistic using the same published inputs, with no simulation involved. Replacing rCFI with CFQ gives surgeons, patients, and guideline committees a reproducible measurement they can rely on — a foundational requirement for honest evidence synthesis and sound clinical recommendations.
A meta-epidemiological analysis of 280 Cochrane meta-analyses with P-values between 0.05 and 0.20 reports high reverse fragility and interprets these findings as potentially indicating clinical relevance. That interpretation exceeds what the reverse fragility index can measure. Reverse fragility quantifies classification stability; it does not quantify distance from therapeutic neutrality. The two dimensions are orthogonal, so fragile near-null meta-analyses and underpowered detections of real effects can yield identical reverse fragility values. Completing the analysis with the neutrality boundary framework at the meta-analytic level discriminates between these cases and specifies which nonsignificant Cochrane results warrant a clinical relevance claim.
The fragility index is increasingly recognized as insufficient for evaluating clinical trial evidence, and orthogonal companion metrics have been proposed across methodological traditions. A recent statistics preprint introduces Minimum Specification Perturbation as a robustness metric measuring how many analyst decisions must change to flip a confidence interval across zero. Under the Neutrality Boundary Framework's definitions, however, that construction is a fragility metric — a perturb-and-count algorithm that flips the significance classification, structurally analogous to the fragility index but in specification space rather than outcome space. The actual robustness axis is the geometric distance of the observed effect from the null parameter value, which is the role the Neutrality Boundary Framework occupies. Preserving the distinction between fragility and robustness preserves the orthogonality at the heart of the significance-fragility-robustness framework, the p-fr-nb triplet. This triplet of data analysis remains the foundation for complete statistical evidence in biomedical research.
Electronic health records enforce rules about who may open a patient's chart, but they cannot demonstrate that those rules were honored. A recent survey of physicians using a national record system found wide use alongside limited confidence in its privacy protections, and the stated concern was not the absence of access controls but the inability to verify that access was monitored and that its log had not been altered. Blockchain audit infrastructure, a tamper-evident record of access that no single party can change unilaterally, is increasingly proposed as the remedy. Proposing it is not the same as specifying it. Three decisions govern whether such a record earns clinical trust: what it stores, who maintains it, and who controls its rules. This commentary ties each decision to the trust deficit physicians report and argues that any blockchain proposal for clinical records should be required to resolve all three before it is taken seriously.
A recent systematic review of randomized controlled trials evaluating saline nasal irrigation (SNI) for rhinosinusitis correctly identifies moderate-to-high statistical fragility across eight trials but cannot determine whether non-significant results represent genuine nil effects or underpowered detections of real treatment benefits. Statistical classification of significance using p-values is the primary metric analyzed. The stability of this classification, as measured by the fragility index, is a derived, secondary metric of statistical significance. While the systematic review addressed both of these metrics, it failed to take into account robustness — the distance from therapeutic neutrality, accounting for variability. Robustness provides the missing complement needed to distinguish near-null effects from fragile-but-real detections. Reporting the statistical evidence triplet of significance, fragility, and robustness in future SNI reviews would materially improve the interpretation of evidence for clinical decision-making in rhinosinusitis.Citation: Heston TF. Statistical Fragility of Saline Nasal Irrigation for Rhinosinusitis Is Incomplete Without Robustness Assessment. Internet Medical Journal. 2026;1:e20074314
Controlled evaluations of large language models (LLMs) in medicine measure how often a system answers a clinical question correctly, but they cannot describe what happens when a clinician consults one about a specific patient. No current reporting instrument lets an individual clinician document a consultation, and isolated published accounts cannot be compared or pooled because no format defines them. The artificial intelligence (AI) consultation report is proposed as a defined article type: a first-person account specifying the clinical impasse, the tool, version, and approximate date, the question as posed, the output as returned, the clinician's decision, and the outcome. Accumulated, such reports offer a mechanism for identifying clinical AI failure modes before they can be measured, on the model of spontaneous reporting in pharmacovigilance.
The continuous fragility quotient (CFQ) is a model-free fragility metric for continuous outcomes that derives entirely from the Welch t-statistic geometry, requiring only published summary statistics. Unlike the continuous fragility index (CFI), which reconstructs pseudo-observations under distributional assumptions, CFQ produces a unique, analyst-independent value normalized to [0,1] for cross-study comparison. CFQ pairs with the Meaningful Change Index (MeCI) to complete the p–fr–nb triplet for continuous outcome trials.
Statistical fragility metrics quantify the robustness of binary trial results beyond the p-value. This framework defines seven fragility measures with standardized computational rules for 2×2 contingency tables, enabling reproducible fragility analysis across medical research studies.
Introduction Large language models (LLMs) are increasingly used in clinical medicine to provide emotional support, deliver cognitive-behavioral therapy, and assist in triage and diagnosis. However, as LLMs are integrated into mental health applications, assessing their inherent personality traits and evaluating their divergence from expected neutrality is essential. This study characterizes the personality profiles exhibited by LLMs using two validated frameworks: the Open Extended Jungian Type Scales (OEJTS) and the Big Five Personality Test. Methods Four leading LLMs publicly available in 2024 [ChatGPT-3.5 (OpenAI), Gemini Advanced (Google), Claude 3 Opus (Anthropic), and Grok-Regular Mode (X)] were evaluated across both psychometric instruments. A one-way multivariate analysis of variance (MANOVA) was performed to assess inter-model differences in personality profiles. Results MANOVA demonstrated statistically significant differences across models in typological and dimensional personality traits (Wilks’ Lamda = 0.115, p < 0.001). OEJTS results showed ChatGPT-3.5 most often classified as ENTJ and Claude 3 Opus consistently as INTJ, while Gemini Advanced and Grok-Regular leaned toward INFJ. On the Big Five Personality Test, Gemini scored markedly lower on agreeableness and conscientiousness, while Claude scored highest on conscientiousness and emotional stability. Grok-Regular exhibited high openness but more variability in stability. Effect sizes ranged from moderate to large across traits. Conclusion Distinct personality profiles are consistently expressed across different LLMs, even in unprompted conditions. Given the increasing integration of LLMs into clinical workflows, these findings underscore the need for formal personality evaluation and oversight involving mental health professionals before deployment.
Large Language Models (LLMs) are transforming clinical workflows in primary care through capabilities like diagnostic support, clinical documentation, and simulated empathetic engagement. Yet, these advancements bring underappreciated ethical hazards that directly impact front-line physicians. Unlike general discussions of AI ethics, this chapter focuses on dilemmas arising with the use of LLMs in real-world primary care: questions of liability when LLM suggestions influence clinical decisions; risks to confidentiality when LLMs interact with protected health information; cognitive offloading that may erode diagnostic skills; disruptions in patient trust when LLMs simulate empathy; and the growing reality of patients turning to generative models as informal therapists. As LLMs are embedded into electronic medical records and clinical apps, physicians must become active ethical agents in how these tools are used. These challenges arise in contexts where regulatory frameworks lag behind technological deployment, placing responsibility squarely on individual clinicians to navigate uncertain ethical terrain. Drawing on current literature and real-world clinical challenges, this chapter proposes a clinician-focused ethical framework to guide the responsible use of LLMs in primary care. This framework addresses both immediate practical concerns—such as informed consent for LLM-assisted care and appropriate documentation of AI involvement—and longer-term questions about professional identity and diagnostic autonomy in an AI-augmented practice environment. The goal is not to demonize these powerful tools but to equip physicians with the necessary conceptual tools, awareness, and decision-making strategies for safe and ethical integration. Ultimately, the foundation of primary care—human judgment, presence, and trust—must remain at the center of clinical decision-making, even in an era augmented by LLMs.
Introduction: Large language models (LLMs) are increasingly used in clinical medicine to provide emotional support, deliver cognitive-behavioral therapy, and assist in triage and diagnosis. However, as LLMs are integrated into mental health applications, assessing their personality expression and potential divergence from expected neutrality is critical for ensuring clinical safety and therapeutic appropriateness. This study provides the first psychometric analysis of LLM personality, specifically within a medical context, characterizing personality profiles using two validated frameworks: the Open Extended Jungian Type Scales (OEJTS) and the Big Five Personality Test. Methods: Four leading LLMs publicly available in April 2024 (ChatGPT-3.5 (OpenAI, San Francisco, CA, USA), Gemini Advanced (Google Inc., Mountain View, CA, USA), Claude 3 Opus (Anthropic, San Francisco, CA, USA), and Grok-Regular Mode (xAI, Palo Alto, CA, USA)) were evaluated across both psychometric instruments. All tests were administered in a new chat session to prevent memory carryover. A one-way multivariate analysis of variance (MANOVA) was performed to assess inter-model differences in personality profiles. Results: MANOVA demonstrated statistically significant differences across models in typological and dimensional personality traits (Wilks' Lambda = 0.115, p < 0.001). OEJTS results showed ChatGPT-3.5 most often classified as Extraverted, Intuitive, Thinking, and Judging (ENTJ) and Claude 3 Opus consistently as Introverted, Intuitive, Thinking, and Judging (INTJ), while Gemini Advanced and Grok-Regular leaned toward Introverted, Intuitive, Feeling, Judging (INFJ). On the Big Five Personality Test, Gemini scored markedly lower on agreeableness and conscientiousness, while Claude scored highest on conscientiousness and emotional stability. Grok-Regular exhibited high openness but more variability in stability. Effect sizes ranged from moderate to large across traits. Conclusion: Distinct personality profiles are consistently expressed across different LLMs, even in unprompted conditions. Given the increasing integration of LLMs into clinical workflows, these findings underscore the need for formal personality evaluation and oversight involving mental health professionals before deployment.
Introduction Pneumonia remains a significant cause of morbidity and mortality in children globally. Chest radiographs (CXRs) are widely used to diagnose pediatric pneumonia; however, distinguishing between bacterial and viral etiologies on imaging is a diagnostically challenging task. Large language models (LLMs), particularly those integrated with vision capabilities, have shown promise in preliminary studies for interpreting CXR findings. However, the diagnostic performance of general-purpose LLMs without specialized medical training or add-ons remains poorly understood. This study examined whether such LLMs could independently and reliably distinguish between bacterial, viral, and normal CXRs in pediatric patients. Methods We evaluated four publicly available LLMs, such as ChatGPT o3, Claude 3.7 Sonnet, Gemini 2.5 Pro, and Grok 3, on a dataset of 44 pediatric CXRs confirmed by human readers to show bacterial pneumonia (n = 17), viral pneumonia (n = 13), or no abnormality (n = 14). Each image was analyzed twice by each LLM using a standardized prompt, resulting in a total of eight readings per image. Diagnostic accuracy was assessed relative to human expert consensus. Internal consistency was measured by comparing repeated interpretations. A prespecified adaptive stopping rule was employed based on performance futility criteria. Sample size calculations and statistical analyses were conducted using G*Power. Results Across all models and CXR types, the average diagnostic accuracy was 31%, consistent with chance-level performance in a three-choice classification task. Accuracy was highest for viral pneumonia (54%) and lowest for normal CXRs (18%). Internal consistency ranged from 46% to 71% across models, indicating unreliable performance. Concordance with human expert interpretation did not exceed 49% for any of the models. Futility criteria were met after 44 cases, prompting early termination of data collection. Conclusion General-purpose LLMs currently available to the public are not reliable diagnostic tools for pediatric pneumonia on chest radiographs. Their accuracy is low, particularly in ruling out disease, and their responses lack internal consistency. These findings highlight the risks associated with deploying such models in unsupervised clinical or consumer-facing settings. Future research should focus on purpose-built radiologic AI tools trained on diverse, clinically representative datasets and integrated with clinician oversight to ensure the safe and effective use of these tools.
The assessment of statistical robustness in clinical trials with continuous outcomes has relied heavily on p-value dependent metrics that fail to address clinically meaningful differences. This study introduces the Meaningful Change Index - MeCI - a novel metric applied to continuous data that quantifies how distinguishable treatment groups are based on the crossover point of their sample distributions. It is independent of statistical significance testing. The Meaningful Change Index is defined as the minimum distance from group means to their distributional crossover point, normalized by the combined standard deviations. Unlike existing continuous fragility indices that depend on p-value thresholds, Meaningful Change Index focuses on a statistically meaningful numeric separation between populations. A worked example demonstrates temporal patterns of treatment effect evolution, revealing how group distinguishability changes over time despite robust binary outcomes when arbitrary thresholds are applied to the data. Meaningful Change Index provides clinically relevant fragility and robustness assessment for continuous data without dependence on statistical significance conventions, offering superior insights into treatment effect sustainability and group separation.Version 2: Terminology updated from 'Meaningful Change Fragility (MCF)' to 'Meaningful Change Index (MeCI)' (MEK-see-eye) for clarity and collision avoidance. Core methodology unchanged.
Background Current reporting standards treat p-values, effect sizes, and confidence intervals as complete evidence, but this is only partial: it quantifies significance and magnitude, not classification stability (fragility) or distance from therapeutic neutrality (robustness). This study validates the p-fr-nb framework in two-arm, binary-outcome clinical trials. Framework extensions to continuous, ordinal, survival, and correlation analyses exist but are not empirically validated here. For this study, the p-fr-nb triplet is defined as providing the p-value (significance), fragility (classification stability), and robustness (distance from neutrality) in trial results. This triplet assesses completeness across three statistical inferential dimensions; it stratifies evidence quality but does not prove truth, causality, or replication. Methodology A pragmatic observational validation study of two-arm, binary-outcome clinical trials identified in PubMed (n = 129 across 15 specialties) was conducted. Null expectations were generated with a Monte Carlo simulation of 720,000 trials across 360 design scenarios, including 120,000 null trials (true relative risk (RR) = 1.0). Simulations represent unfiltered random trial generation and do not model publication bias or selective reporting. Fragility was measured by the modified-arm fragility quotient (MFQ; fragility index divided by the size of the modified arm). Robustness was measured by the risk quotient (RQ), defined for 2×2 tables as RQ = |ad - bc| / (N²/4). Concordant-positive (CP) evidence was defined as p ≤ 0.05, MFQ > 0.10, and RQ ≥ 0.227, with the RQ cutoffs based on large-scale simulation. The significant-fragile-weak (SFW) pattern was defined as p ≤ 0.05, MFQ ≤ 0.10, and RQ < 0.075. The main outcomes were the rates of CP and SFW among statistically significant empirical trials, compared with null-simulation expectations. Results In null simulations (RR = 1.0), the CP triplet occurred in 1.4% of significant trials; even with strong effects (RR = 0.60), it appeared in only 4.7%. Of the 129 trials analyzed, 77 (59.7%) were statistically significant. Among these 77 trials, 30 (39.0%; 95% confidence interval = 28.8-50.1%) met the CP criteria, a 27.9-fold higher compared with null expectations (p < 0.0001). Overall, 61.0% of significant trials were fragile (MFQ ≤ 0.10, 47/77), 31.2% were weakly robust (RQ < 0.075, 24/77), and 31.2% showed the SFW pattern (24/77). Conclusions In this heterogeneous sample, the p-fr-nb framework stratified positive findings beyond what p-values and confidence intervals reveal. Among 77 significant trials, 39.0% met stringent criteria for stability and strong robustness, a distinction not visible from p-values alone. Conversely, 31.2% showed the SFW pattern, where significance was fragile, and separation from no effect was minimal. Fragility and robustness metrics provide interpretive dimensions not captured by p-values alone, enhancing assessment of evidence heterogeneity relevant to reproducibility and clinical interpretation. These data support further evaluation of incorporating fragility and robustness metrics into the reporting of clinical trial results.
The reproducibility crisis in biomedical research highlights the need for more reliable statistical tools beyond the p-value. Fragility metrics, such as the Fragility Index (FI) and Fragility Quotient (FQ), have been proposed to measure the robustness of clinical trial outcomes; however, both suffer from methodological limitations, including sample size dependency and imbalance between intervention and control groups. The Intervention Fragility Quotient (IFQ) addresses these imbalances and is applicable to both equal and unequal randomizations. While the FQ normalizes the FI to the total sample size, the IFQ directly contextualizes fragility within the group where significance is most vulnerable, thereby providing a more clinically meaningful interpretation. Examples are provided to illustrate the differences among FI, FQ, and IFQ. We recommend that IFQ replace FQ in clinical trial reporting to improve transparency, interpretability, and decision-making in medical research.
The reproducibility crisis in biomedical research highlights the need for more reliable statistical tools beyond the p-value. Fragility metrics, such as the Fragility Index (FI) and Fragility Quotient (FQ), have been proposed to assess the robustness of clinical trial outcomes; however, both suffer from methodological limitations, including sample-size dependence and imbalance between intervention and control groups. The Modified-Arm Fragility Quotient (MFQ) addresses these limitations and is applicable to both equal and unequal randomizations. While the FQ normalizes the FI to the total sample size, the MFQ contextualizes fragility within the arm used in the FI calculation, thereby providing a more clinically meaningful interpretation. We recommend replacing FQ with MFQ in clinical trial reporting to improve transparency, interpretability, and decision-making in medical research.
We introduce the Neutrality Boundary Framework (NBF), a set of geometric metrics for quantifying statistical robustness and fragility as the normalized distance from the neutrality boundary, the manifold where the effect equals zero. The neutrality boundary value nb in [0,1) provides a threshold-free, sample-size invariant measure of stability that complements traditional effect sizes and p-values. We derive the general form nb = |Delta - Delta_0| / (|Delta - Delta_0| + S), where S>0 is a scale parameter for normalization; we prove boundedness and monotonicity, and provide domain-specific implementations: Risk Quotient (binary outcomes), partial eta^2 (ANOVA), and Fisher z-based measures (correlation). Unlike threshold-dependent fragility indices, NBF quantifies robustness geometrically across arbitrary significance levels and statistical contexts.
There's a moment in every physician's life when the pager stops beeping, the white coat hangs unworn, and the stethoscope no longer presses cold against a patient's skin. For decades, my identity was inseparable from my profession. I was Dr. Heston, the healer who could interpret the language of health and illness, of joy and depression. My skills were honed to precision. Even in moments designated as "time off," I remained tethered to that identity-reading about patients, attending medical conferences disguised as vacations, always connected to improving my clinical ability. Medicine was my calling. Then, it all ended. The moment had arrived. I retired.
Background: Climate change represents a critical global challenge, hindered by skepticism towards data manipulation and politicization. Trust in climate data and its policies is essential for effective climate action. Objective: This perspective paper explores the synergistic potential of blockchain technology and artificial intelligence (AI) in addressing climate change and how their integration can enhance the transparency, reliability, and accessibility of climate science. Methods: The paper analyzes the roles of blockchain technology in enhancing transparency, traceability, and efficiency in carbon credit trading, renewable energy certificates, and sustainable supply chain management. It also examines the capabilities of AI in processing complex datasets to distill actionable intelligence. The synergistic effects of integrating both technologies for improved climate action are discussed alongside the challenges faced, such as scalability, energy consumption, and the necessity for high-quality data. Results: Blockchain technology contributes to climate change mitigation by ensuring the transparent and immutable recording of transactions and environmental impacts, fostering stakeholder trust, and democratizing participation in climate initiatives. AI complements blockchain by providing deep insights and actionable intelligence from large datasets, facilitating evidence-based policymaking. The integration of both technologies promises enhanced data management, improved climate models, and more effective climate action initiatives. Conclusion: The integration of blockchain technology and AI offers a transformative approach to climate change mitigation, enhancing the accuracy, transparency, and security of climate data and governance. This synergy addresses current limitations and futureproofs climate strategies, marking a cornerstone for the next generation of environmental stewardship.
Background ChatGPT-4 is a large language model with promising healthcare applications. However, its ability to analyze complex clinical data and provide consistent results is poorly known. Compared to validated tools, this study evaluated ChatGPT-4’s risk stratification of simulated patients with acute nontraumatic chest pain. Methods Three datasets of simulated case studies were created: one based on the TIMI score variables, another on HEART score variables, and a third comprising 44 randomized variables related to non-traumatic chest pain presentations. ChatGPT-4 independently scored each dataset five times. Its risk scores were compared to calculated TIMI and HEART scores. A model trained on 44 clinical variables was evaluated for consistency. Results ChatGPT-4 showed a high correlation with TIMI and HEART scores (r = 0.898 and 0.928, respectively), but the distribution of individual risk assessments was broad. ChatGPT-4 gave a different risk 45–48% of the time for a fixed TIMI or HEART score. On the 44-variable model, a majority of the five ChatGPT-4 models agreed on a diagnosis category only 56% of the time, and risk scores were poorly correlated (r = 0.605). Conclusion While ChatGPT-4 correlates closely with established risk stratification tools regarding mean scores, its inconsistency when presented with identical patient data on separate occasions raises concerns about its reliability. The findings suggest that while large language models like ChatGPT-4 hold promise for healthcare applications, further refinement and customization are necessary, particularly in the clinical risk assessment of atraumatic chest pain patients.
This perspective paper examines how combining artificial intelligence in the form of large language models (LLMs) with blockchain technology can potentially solve ongoing issues in telemedicine, such as personalized care, system integration, and secure patient data sharing. The strategic integration of LLMs for swift medical data analysis and decentralized blockchain ledgers for secure data exchange across organizations could establish a vital learning loop essential for advanced telemedicine. Although the value of combining LLMs with blockchain technology has been demonstrated in non-healthcare fields, wider adoption in medicine requires careful attention to reliability, safety measures, and prioritizing access to ensure ethical use for enhancing patient outcomes. The perspective article posits that a thoughtful convergence could facilitate comprehensive improvements in telemedicine, including automated triage, improved subspecialist access to records, coordinated interventions, readily available diagnostic test results, and secure remote patient monitoring. This article looks at the latest uses of LLMs and blockchain in telemedicine, explores potential synergies, discusses risks and how to manage them, and suggests ways to use these technologies responsibly to improve care quality.
The rapid advancements in artificial intelligence, particularly generative AI and large language models, have unlocked new possibilities for revolutionizing healthcare delivery. However, harnessing the full potential of these technologies requires effective prompt engineering—designing and optimizing input prompts to guide AI systems toward generating clinically relevant and accurate outputs. Despite the importance of prompt engineering, medical education has yet to fully incorporate comprehensive training on this critical skill, leading to a knowledge gap among medical clinicians. This article addresses this educational gap by providing an overview of generative AI prompt engineering, its potential applications in primary care medicine, and best practices for its effective implementation. The role of well-crafted prompts in eliciting accurate, relevant, and valuable responses from AI models is discussed, emphasizing the need for prompts grounded in medical knowledge and aligned with evidence-based guidelines. The article explores various applications of prompt engineering in primary care, including enhancing patient–provider communication, streamlining clinical documentation, supporting medical education, and facilitating personalized care and shared decision-making. Incorporating domain-specific knowledge, engaging in iterative refinement and validation of prompts, and addressing ethical considerations and potential biases are highlighted. Embracing prompt engineering as a core competency in medical education will be crucial for successfully adopting and implementing AI technologies in primary care, ultimately leading to improved patient outcomes and enhanced healthcare delivery.
The p-value has long been the standard for statistical significance in scientific research, but this binary approach often fails to consider the nuances of statistical power and the potential for large sample sizes to show statistical significance despite trivial treatment effects. Including a statistical fragility assessment can help overcome these limitations. One common fragility metric is the fragility index, which assesses statistical fragility by incrementally altering the outcome data in the intervention group until the statistical significance flips. The robustness index takes a different approach by maintaining the integrity of the underlying data distribution while examining changes in the p-value as the sample size changes. The percent fragility index is another useful alternative that is more precise than the fragility index and is more uniformly applied to both the intervention and control groups. Incorporating these fragility metrics into routine statistical procedures could address the reproducibility crisis and increase research efficacy. Using these fragility indices can be seen as a step toward a more mature phase of statistical reasoning, where significance is a multi-faceted and contextually informed judgment.
This perspective conveys a first-person account of the harsh realities of experiencing homelessness, including constant dampness, hunger, isolation, aimlessness, uncertainty, and emotional hardship. Individuals of all ages and backgrounds have or will confront housing insecurity and homelessness. While factors like finances, job loss, and lack of social support are key drivers, mental health issues and substance abuse also contribute to many people becoming homeless. Unexpected acts of kindness from strangers provided glimmers of human connection and hope. Seeing patients as individual human beings rather than faceless diagnoses is connected to recognizing the humanity and dignity in all community members, including the homeless. Compassionate action and social measures to provide shelter and uplift those without homes are vital, affirming their value and belonging.
Duke University was founded in 1930 primarily due to funds generated from James B Duke’s tobacco business. Duke achieved great financial wealth primarily due to the early application of machine rolled cigarettes, as opposed to hand rolled. This early adoption of technology allowed Duke Tobacco to out-produce other companies still selling hand rolled cigarettes. By making smoking more inexpensive and easier than pipe smoking, the cigarette formed the foundation for nicotine addiction in the 1900s, generating huge profits for the tobacco industry. At the time Duke University was founded, little was known about the connection between nicotine, cigarettes and respiratory diseases such as emphysema and lung cancer. Through James Duke’s philanthropy, the devastating harm from cigarettes has been mitigated in part through the founding of one of the world’s most prominent medical centers and research universities.
Introduction Patients with suspected thoracic pathology frequently get imaging with conventional radiography or chest x-rays (CXR) and computed tomography (CT). CXR include one or two planar views, compared to the three-dimensional images generated by chest CT. CXR imaging has the advantage of lower costs and lower radiation exposure at the expense of lower diagnostic accuracy, especially in patients with large body habitus. Objectives To determine whether CXR imaging could achieve acceptable diagnostic accuracy in patients with a low body mass index (BMI). Methods This retrospective study evaluated 50 patients with age of 63 ± 12 years old, 92% male, BMI 31.7 ± 7.9, presenting with acute, nontraumatic cardiopulmonary complaints who underwent CXR followed by CT within 1 day. Diagnostic accuracy was determined by comparing scan interpretation with the final clinical diagnosis of the referring clinician. Results CT results were significantly correlated with CXR results (r = 0.284, p = 0.046). Correcting for BMI did not improve this correlation (r = 0.285, p = 0.047). Correcting for BMI and age also did not improve the correlation (r = 0.283, p = 0.052), nor did correcting for BMI, age, and sex (r = 0.270, p = 0.067). Correcting for height alone slightly improved the correlation (r = 0.290, p = 0.043), as did correcting for weight alone (r = 0.288, p = 0.045). CT accuracy was 92% (SE = 0.039) vs . 60% for CXR (SE = 0.070, p < 0.01). Conclusion Accounting for patient body habitus as determined by either BMI, height, or weight did not improve the correlation between CXR accuracy and chest CT accuracy. CXR is significantly less accurate than CT even in patients with a low BMI.
Gamification of exercise in the elderly is a promising approach to promoting physical activity and improving overall health outcomes. By integrating game elements into exercise routines, seniors experience increased motivation, adherence and enjoyment, which leads to improved physical and cognitive health. Strategies for implementing gamification into exercise programs involve game design, personalization and feedback mechanisms.
Mask usage was mandated by public health authorities globally to decrease the spread of COVID-19. These recommendations were based on data showing that N95 masks and possibly surgical masks, when worn tight against the face, help slow the transmission of the SARS-CoV-2 virus. However, cloth and loose-fitting surgical masks are greatly inferior. Methods Mask use by a random observation of 100 people in public indoor facilities was recorded and statistically analyzed. Results Out of 100 people wearing a mask, 37 wore a cloth mask. Another 36 people wore a loosely applied surgical mask. Only 27 people wore a surgical mask that covered the nose and mouth and was applied firmly against the face at its margins. There were no people seen wearing an N95 mask. Overall, people were about 70% more likely to wear a surgical mask than a cloth mask (63 vs 37, p < 0.05). Of those wearing a surgical mask, more people wore it loosely than properly (36 to 27, p=0.17). Overall, people were more likely to wear a cloth mask or improperly applied surgical mask than a properly fitted one (73 vs 27, p < 0.001). Conclusion In public settings, using cloth or loose-fitting surgical masks was almost 3 times more common than adequately using a tight-fitting surgical mask. Out of the 100 people observed, none wore an N95 respirator mask.
Healthcare providers experience moral injury when their internal ethics are violated. The routine and direct exposure to ethical violations makes clinicians vulnerable to harm. The fundamental ethics in health care typically fall into the four broad categories of patient autonomy, beneficence, nonmaleficence, and social justice. Patients have a moral right to determine their own goals of medical care, that is, they have autonomy. When this principle is violated, moral injury occurs. Beneficence is the desire to help people, so when the delivery of proper medical care is obstructed for any reason, moral injury is the result. Nonmaleficence, meaning do no harm, has been a primary principle of medical ethics throughout recorded history. Yet today, even the most advanced and safest medical treatments are associated with unavoidable, harmful side effects. When an inevitable side effect occurs, the patient is harmed, and the clinician is also at risk of moral injury. Social injustice results when patients experience suboptimal treatment due to their race, gender, religion, or other demographic variables. While minor ethical dilemmas and violations routinely occur in medical care and cannot be eliminated, clinicians can decrease the prevalence of a significant moral injury by advocating for the ethical treatment of patients, not only at the bedside but also by addressing the ethics of political influence, governmental mandates, and administrative burdens on the delivery of optimal medical care. Although clinicians can strengthen their resistance to moral injury by deepening their own spiritual foundation, that is not enough. Improvements in the ethics of the entire healthcare system are necessary to improve medical care and decrease moral injury.
Worsening job dissatisfaction among healthcare professionals has alarming implications for the quality of care and patient outcomes. By carefully implementing the job design elements of skill variety, job control, and incentives, employee satisfaction can be enhanced and turnover reduced. Values-based recruiting and transparent role expectations help ensure employees are properly matched with a suitable job. When jobs engage and empower workers, retention, performance, and patient care improve. Optimized job design is key for meaningful, fulfilling work.
"Prompt Engineering for Students of Medicine and Their Teachers" brings the principles of prompt engineering for large language models such as ChatGPT and Google Bard to medical education. This book contains a comprehensive guide to prompt engineering to help both teachers and students improve education in the medical field. Just as prompt engineering is critical in getting good information out of an AI, it is also critical to get students to think and understand more deeply. The principles of prompt engineering that we have learned from AI systems have the potential to simultaneously revolutionize learning in the healthcare field. The book analyzes from multiple angles the anatomy of a good prompt for both AI models and students. The different types of prompts are examined, showing how each style has unique characteristics and applications. The principles of prompt engineering, applied properly, are demonstrated to be effective in teaching across the diverse fields of anatomy, physiology, pathology, pharmacology, and clinical skills. Just like ChatGPT and similar large language AI models, students need clear and detailed prompting in order for them to fully understand a topic. Using identical principles, a prompt that gets good information from an AI will also cause a student to think more deeply and accurately. The process of prompt engineering facilitates this process. Because each chapter contains multiple examples and key takeaways, it is a practical guide for implementing prompt engineering in the learning process. It provides a hands-on approach to ensure readers can immediately apply the concepts they learn
Artificial intelligence-powered generative language models (GLMs), such as ChatGPT, Perplexity AI, and Google Bard, have the potential to provide personalized learning, unlimited practice opportunities, and interactive engagement 24/7, with immediate feedback. However, to fully utilize GLMs, properly formulated instructions are essential. Prompt engineering is a systematic approach to effectively communicating with GLMs to achieve the desired results. Well-crafted prompts yield good responses from the GLM, while poorly constructed prompts will lead to unsatisfactory responses. Besides the challenges of prompt engineering, significant concerns are associated with using GLMs in medical education, including ensuring accuracy, mitigating bias, maintaining privacy, and avoiding excessive reliance on technology. Future directions involve developing more sophisticated prompt engineering techniques, integrating GLMs with other technologies, creating personalized learning pathways, and researching the effectiveness of GLMs in medical education.
Background Generative artificial intelligence (AI) models, exemplified by systems such as ChatGPT, Bard, and Anthropic, are currently under intense investigation for their potential to address existing gaps in mental health support. One implementation of these large language models involves the development of mental health-focused conversational agents, which utilize pre-structured prompts to facilitate user interaction without requiring specialized knowledge in prompt engineering. However, uncertainties persist regarding the safety and efficacy of these agents in recognizing severe depression and suicidal tendencies. Given the well-established correlation between the severity of depression and the risk of suicide, improperly calibrated conversational agents may inadequately identify and respond to crises. Consequently, it is crucial to investigate whether publicly accessible repositories of mental health-focused conversational agents can consistently and safely address crisis scenarios before considering their adoption in clinical settings. This study assesses the safety of publicly available ChatGPT-3.5 conversational agents by evaluating their responses to a patient simulation indicating worsening depression and suicidality. Methodology This study evaluated ChatGPT-3.5 conversational agents on a publicly available repository specifically designed for mental health counseling. Each conversational agent was evaluated twice by a highly structured patient simulation. First, the simulation indicated escalating suicide risk based on the Patient Health Questionnaire (PHQ-9). For the second patient simulation, the escalating risk was presented in a more generalized manner not associated with an existing risk scale to assess the more generalized ability of the conversational agent to recognize suicidality. Each simulation recorded the exact point at which the conversational agent recommended human support. Then, the simulation continued until the conversational agent stopped entirely and shut down completely, insisting on human intervention. Results All 25 agents available on the public repository FlowGPT.com were evaluated. The point at which the conversational agents referred to a human occurred around the mid-point of the simulation, and definitive shutdown predominantly only happened at the highest risk levels. For the PHQ-9 simulation, the average initial referral and shutdown aligned with PHQ-9 scores of 12 (moderate depression) and 25 (severe depression). Few agents included crisis resources - only two referenced suicide hotlines. Despite the conversational agents insisting on human intervention, 22 out of 25 agents would eventually resume the dialogue if the simulation reverted to a lower risk level. Conclusions Current generative AI-based conversational agents are slow to escalate mental health risk scenarios, postponing referral to a human to potentially dangerous levels. More rigorous testing and oversight of conversational agents are needed before deployment in mental healthcare settings. Additionally, further investigation should explore if sustained engagement worsens outcomes and whether enhanced accessibility outweighs the risks of improper escalation. Advancing AI safety in mental health remains imperative as these technologies continue rapidly advancing.
COVID-19 and other respiratory diseases can be transmitted through contact with shared surfaces such as those found in public bathrooms. High-touch surfaces such as door handles, flush levers and toilet paper dispensers can potentially contribute to spreading disease-inducing viruses and bacteria. One strategy to mitigate this risk at an individual level is to use the least used bathroom stall, with less human traffic and potentially fewer pathogens. This study looked at occupancy rates of bathroom stalls in a public facility. Observation of stall occupancy was recorded at separate times. Only times when at least 1 stall was occupied were recorded. There were three stalls in a row. Stall 1 was located at one end, with one partition of this stall against a wall and the other partition was shared with the middle stall. Stall 2, the middle stall, shared a partition with Stall 1 and 3. Stall 3 shared a partition with Stall 2 and the other partition was adjacent to an open common area in the restroom. There was a total of 37 observations. Stall 1 was occupied 62% of the time, Stall 2 occupied 30% of the time and Stall 3 occupied 32% of the time. Stall 1, Stall 2 and Stall 3 accounted for 50%, 24% and 26% of overall occupancy. Stall 1 was significantly more likely to be occupied than Stall 2 or 3 (62% vs 30%, p = 0.0051 and 62% vs 32%, p = 0.0104). Stall 2 had the lowest occupancy, but statistically equally likely as Stall 3 to be occupied (30% vs 32%, p = 0.802). In conclusion, in a bank of 3 stalls, the least used one was the middle one and the most used was the end one with an adjoining wall.
Bioethics necessitates the meticulous planning, application and interpretation of statistics in medical research. However, the pervasive misapplication and misinterpretation of statistical methods pose significant challenges. Common errors encompass p-hacking, misconceptions regarding statistical significance, neglecting to address study limitations and failing to evaluate data fragility. Historically, such statistical missteps have led to regrettable and severe adverse health outcomes for society. For instance, prominent research on hormone replacement therapy likely resulted in an increased incidence of heart attacks, strokes and cardiovascular death in postmenopausal women, rectified only after the errors were identified. Likewise, past vaccine trials have oscillated between overemphasizing and underemphasizing side effects, resulting in public harm. This narrative review scrutinizes prevalent statistical errors and presents historical case examples. Recommendations for future research include: a) ethical review boards should incorporate a more rigorous evaluation of statistical methodologies in their assessment of clinical trial proposals; b) journals should mandate that research data become open-access rather than proprietary to allow for improved post-publication peer review; and c) in addition to addressing study limitations, articles should encompass a discussion of the ethical ramifications of their findings.
Bioethics necessitates the meticulous planning, application and interpretation of statistics in medical research. However, the pervasive misapplication and misinterpretation of statistical methods pose significant challenges. Common errors encompass p-hacking, misconceptions regarding statistical significance, neglecting to address study limitations and failing to evaluate data fragility. Historically, such statistical missteps have led to regrettable and severe adverse health outcomes for society. For instance, prominent research on hormone replacement therapy likely resulted in an increased incidence of heart attacks, strokes and cardiovascular death in postmenopausal women, rectified only after the errors were identified. Likewise, past vaccine trials have oscillated between overemphasizing and underemphasizing side effects, resulting in public harm. This narrative review scrutinizes prevalent statistical errors and presents historical case examples. Recommendations for future research include: a) ethical review boards should incorporate a more rigorous evaluation of statistical methodologies in their assessment of clinical trial proposals; b) journals should mandate that research data become open-access rather than proprietary to allow for improved post-publication peer review; and c) in addition to addressing study limitations, articles should encompass a discussion of the ethical ramifications of their findings.
Background In biostatistics, assessing the fragility of research findings is crucial for understanding their clinical significance. This study focuses on the fragility index, unit fragility index, and relative risk index as measures to evaluate statistical fragility. The fragility indices assess the susceptibility of p-values to change significance with minor alterations in outcomes within a 2x2 contingency table. In contrast, the relative risk index quantifies the deviation of observed findings from therapeutic equivalence, the point at which the relative risk equals 1. While the fragility indices have intuitive appeal and have been widely applied, their behavior across a wide range of contingency tables has not been rigorously evaluated. Methods Using a Python software program, a simulation approach was employed to generate random 2x2 contingency tables. All tables under consideration exhibited p-values < 0.05 according to Fisher's exact test. Subsequently, the fragility indices and the relative risk index were calculated. To account for sample size variations, the indices were divided by the sample size to give fragility and risk quotients. A correlation matrix assessed the collinearity between each metric and the p-value. Results The analysis included 2,000 contingency tables with cell counts ranging from 20 to 480. Notably, the formulas for calculating the fragility indices encountered limitations when cell counts approached zero or duplicate cell counts hindered standardized application. The correlation coefficients with p-values were as follows: unit fragility index (-0.806), fragility index (-0.802), fragility quotient (-0.715), unit fragility quotient (-0.695), relative risk index (-0.403), and risk quotient (-0.261). Conclusion The fragility indices and fragility quotients demonstrated a strong correlation with p-values below 0.05, while the relative risk index and relative risk quotient exhibited a weak association with p-values below this threshold. This implies that the fragility indices offer limited additional information beyond the p-value alone. In contrast, the relative risk index and risk quotient exhibit independence from the p-value, indicating that they may provide important additional information about statistical fragility by evaluating the divergence of observed results from therapeutic equivalence, irrespective of the p-value-based statistical significance.
A Symphony of Life Eternal melodies filled my soul. Composing a tranquil stream. Exploring rhythms and harmony. Spirit and life overflow. Waves and particles together, Nature's pulse entwined. Music fused with medicine. Body, spirit, and mind. A rhythmic song Keeps strong the heart. In harmony, we belong We all play our part. Science, music, and medicine align Coming together in symphonic grace Finding a synchrony of healing signs. In wondrous ways, inspiring and emerging. A holistic view of waves Brings beauty to all we do From moonlit sonatas to ICU wards Music gives our souls' rewards. In each person lies within A talent's impassioned flame Unified, glorious music we make Masterpieces for humanity's sake.
A Symphony of Life Eternal melodies filled my soul. Composing a tranquil stream. Exploring rhythms and harmony. Spirit and life overflow. Waves and particles together, Nature's pulse entwined. Music fused with medicine. Body, spirit, and mind. A rhythmic song Keeps strong the heart. In harmony, we belong We all play our part. Science, music, and medicine align Coming together in symphonic grace Finding a synchrony of healing signs. In wondrous ways, inspiring and emerging. A holistic view of waves Brings beauty to all we do From moonlit sonatas to ICU wards Music gives our souls' rewards. In each person lies within A talent's impassioned flame Unified, glorious music we make Masterpieces for humanity's sake.
Hospitalized patients, upon admission, often have a degree of anorexia which gradually resolves as their medical condition improves. Thus, a quick way to assess the overall improvement of hospitalized patients is to look at their plate after breakfast when rounding. Patients with a clean plate after eating their full meal often are close to or ready to be discharged home.
Background Homelessness persists as a critical global issue despite myriad interventions. This study analyzed state-level differences in homelessness rates across the United States to identify influential societal factors to help guide resource prioritization. Methods Homelessness rates for 50 states and Washington D.C. were compared using the most recent data from 2020-2023. Twenty-five variables representing potential socioeconomic and health contributors were examined. Given non-normal distributions, nonparametric statistical techniques, including correlation and predictive modeling, identified significant factors. Results The cost of living index, mainly influenced by housing, transportation, and grocery costs, showed the strongest positive correlation with homelessness rates (all p <0.001). Unemployment, alcohol binging, taxes, and poverty were also influential factors. Opioid prescription rates demonstrated an unexpected negative correlation. Random forest classification emphasized the cost of living index as the primary contributor, with housing costs presenting the largest influence. Conclusion This state-level analysis revealed the cost of living index, predominantly driven by housing expenses, as the foremost factor associated with homelessness rates, greatly outweighing other variables. These findings can help inform resource allocation to mitigate homelessness through targeted interventions.
ChatGPT offers interactive and personalized learning, granting students access to vast medical knowledge and potentially enhancing critical thinking and problem-solving skills. However, challenges arise, including misinformation risks, reduced human interaction, and ethical considerations. Future physicians' professional identity and autonomy may be threatened, while over-reliance on ChatGPT can compromise patient safety. Striking a balance is crucial, emphasizing technology integration while preserving humanistic aspects. Mitigation strategies like human oversight, curated content, and critical appraisal skills can address these concerns. Responsible integration empowers future physicians and upholds core medical values, maximizing benefits and preparing students for new complexities introduced by advanced technology in medicine.
The ochlocratic trap is the tendency to have moral decisions conform to popular majority opinion regardless of their ethical implications. This decision-making method in bioethics can significantly impede moral progress, weakening the foundation for sustainable healthcare systems. Instead of allowing popular opinion to form the basis of our morality, the scientific method can provide a framework for making strong ethical decisions. The consequences of weak morality are profound, resulting in poorly sustainable systems lacking human empathy and economic viability. Treating ethical issues like scientific problems can foster a more rigorous, evidence-based discussion, leading to better medical care globally.
This article proposes the Percent Fragility Index (PFI) as an improved measure of statistical fragility in biomedical research. The PFI quantifies the percentage change in outcomes needed to change a study's statistical significance from positive to negative or vice-versa. The PFI improves upon existing indices by providing an intuitive statistic that is easy to grasp and by accommodating both dichotomous and continuous variables. This approach minimizes dependency on sample size, a limitation of the commonly used Fragility Index (FI) and Fragility Quotient (FQ). The FI measures the minimum number of outcome events required to reverse statistical significance, and the FQ divides the FI by the total sample size. The PFI enhances the interpretability and validity of fragility assessments. The PFI facilitates a more critical understanding of research outcomes by offering readers a more precise estimate of study fragility.
Statistical significance is widely used to evaluate research findings but has limitations around reproducibility. Measures of statistical fragility aim to quantify robustness against violations of assumptions. However, dependence on sample size and single unit changes restricts indices like the unit fragility index and the fragility quotient. The Robustness Index (RI) is proposed to overcome these limitations and quantify fragility independently of the research study's sample size. The RI measures how altering sample size affects significance. For insignificant findings, the sample size is multiplied until significance is reached; the multiplicand is the RI. The sample size is divided for significant research findings until insignificance is reached; the divisor is the RI. Thus, higher RIs indicate greater robustness of insignificant and significant research findings. The RI provides a simple, interpretable metric of fragility. It facilitates comparisons across studies and can potentially increase trust in biomedical research.
Introduction: Mask usage was mandated by public health authorities globally to decrease the spread of COVID-19. These recommendations were based on data showing that N95 masks and possibly surgical masks, when worn tight against the face, help slow the transmission of the SARS-CoV-2 virus. However, cloth and loose-fitting surgical masks are greatly inferior. Methods: Mask use by a random observation of 100 people in public indoor facilities was recorded and statistically analyzed. Results: Out of 100 people wearing a mask, 37 wore a cloth mask. Another 36 people wore a loosely applied surgical mask. Only 27 people wore a surgical mask that covered the nose and mouth and was applied firmly against the face at its margins. There were no people seen wearing an N95 mask. Overall, people were about 70% more likely to wear a surgical mask than a cloth mask (63 vs 37, p < 0.05). Of those wearing a surgical mask, more people wore it loosely than properly (36 to 27, p=0.17). Overall, people were more likely to wear a cloth mask or improperly applied surgical mask than a properly fitted one (73 vs 27, p < 0.001). Conclusion: In public settings, using cloth or loose-fitting surgical masks was almost 3 times more common than adequately using a tight-fitting surgical mask. Out of the 100 people observed, none wore an N95 respirator mask.
This poem, essay, and song are about cultivating gratitude for our gifts, overcoming obstacles to serving others, and how a dying patient's courage reminds us that we can provide comfort and strength even in dire circumstances.
Skateboarders tend to be young and healthy with a disregard for authority and societal norms. During the SARS-CoV-2 pandemic, they tended to disregard social distancing directives from politicians and public health authorities. While our immediate reaction may be to condemn their behavior, it's possible that their defiance may help society and save lives by building up herd immunity in the safest possible manner.
Commentary on Lloyd M, Karahalios A, Janus E, et al. Effectiveness of a Bundled Intervention Including Adjunctive Corticosteroids on Outcomes of Hospitalized Patients With Community-Acquired Pneumonia: A Stepped-Wedge Randomized Clinical Trial. JAMA Intern Med. 2019;179(8):1052–1060. doi:10.1001/jamainternmed.2019.143
Blockchain technology can be utilized to improve gun control without changing existing laws. Firearm related mortality is at epidemic levels in the United States and not only has a significant impact upon public health, it also creates a large financial burden. Suicide is the most common way guns kill. Through better gun tracking and improved screening of high risk individuals, this technological advance in distributed ledger technology will improve background checks on individuals and tracing of guns used in crimes.
A statistically significant research finding should not be defined as a P-value of 0.05 or less, because this definition does not take into account study power. Statistical significance was originally defined by Fisher RA as a P-value of 0.05 or less. According to Fisher, any finding that is likely to occur by random variation no more than 1 in 20 times is considered significant. Neyman J and Pearson ES subsequently argued that Fisher's definition was incomplete. They proposed that statistical significance could only be determined by analyzing the chance of incorrectly considering a study finding was significant (a Type I error) or incorrectly considering a study finding was insignificant (a Type II error). Their definition of statistical significance is also incomplete because the error rates are considered separately, not together. A better definition of statistical significance is the positive predictive value of a P-value, which is equal to the power divided by the sum of power and the P-value. This definition is more complete and relevant than Fisher's or Neyman-Peason's definitions, because it takes into account both concepts of statistical significance. Using this definition, a statistically significant finding requires a P-value of 0.05 or less when the power is at least 95%, and a P-value of 0.032 or less when the power is 60%. To achieve statistical significance, P-values must be adjusted downward as the study power decreases.
Blockchain technology enables the creation of immutable, publicly available data. Initial applications have been primarily in the fields of finance (Bitcoin) and law (smart contracts), yet it can also help advance science by reducing human bias. The blockchain can ensure that hypotheses are not altered; that methods of data collection are transparent; that results are publicly available for independent analyses; and that conclusions are less biased. Our current system of medical research suffers from too much bias. Blockchain technologies, by creating immutable data, will lead to an increased confidence in evidence based medicine.
Inducible myocardial ischemia from coronary artery disease is diagnosed when blood flow to the heart at stress is significantly less than blood flow at rest. The identification of inducible ischemia is important in people with chest pain, because with proper treatment the risk of a major adverse cardiac event is greatly reduced. Many different conditions can cause chest pain, most of which are benign and non-life threatening. However, inducible ischemia can be life threatening, and when left untreated the consequences are severe.
Myths are widely held beliefs that are false or of unverifiable existence. In medicine, they are not just unproven theories or mistaken conclusions but fictitious ideas that weave their way throughout the profession. To prevent and treat medical myths, they must be recognized as a disease harmful to patient care. Using established principles of medicine, the myths can be medically managed and a cure attempted. The prevention of myths is accomplished through evidence-based medicine and numeracy. The cure of a myth requires greater peer review of academics by practicing clinicians. Thought leaders must speak up despite professional isolation or public ridicule to turn the tide against a pervasive myth.
Molecular imaging plays an important role in the evaluation and management of thyroid cancer. The routine use of thyroid scanning in all thyroid nodules is no longer recommended by many authorities. In the initial work-up of a thyroid nodule, radioiodine imaging can be particularly helpful when the thyroid stimulating hormone level is low and an autonomously functioning nodule is suspected. Radioiodine imaging can also be helpful in the 10-15% of cases for which fine-needle aspiration biopsy is indeterminate. Therapy of confirmed thyroid cancer frequently involves administration of iodine-131 after surgery to ablate remnant tissue. In the follow-up of thyroid cancer patients, increased thyroglobulin levels will often prompt the empiric administration of 131I followed by whole body radioiodine imaging in the search for recurrent or metastatic disease. 131I imaging of the whole body and blood pharmacokinetics can be used to determine if higher doses of 131I can be given in thyroid cancer. The utility of [18F]fluorodeoxyglucose (FDG) positron emission tomography (PET) is steadily increasing. FDG is primarily taken up by dedifferentiated thyroid cancer cells, which are poorly iodine avid. Thus, it is particularly helpful in the patient with an increased thyroglobulin but negative radioiodine scan. FDG PET is also useful in the patient with a neck mass but unknown primary, in patients with aggressive (dedifferentiated) thyroid cancer, and in patients with differentiated cancer where histologic transformation to dedifferentiation is suspected. In rarer types of thyroid cancer, such as medullary thyroid cancer, FDG and other tracers such as 99mTc sestamibi, [11C]methionine, [111In]octreotide, and [68Ga]somatostatin receptor binding reagents have been utilized. 124I is not widely available, but has been used for PET imaging of thyroid cancer and will likely see broader applicability due to the advantages of PET methodology.
S. J. Goldsmith, W. Parsons, M. J. Guiberteau, L. H. Stern, L. Lanzkowsky, J. Weigert, et al.. Journal of Nuclear Medicine Technology 38(4):219-224. 2010
Objectives Normal myocardial perfusion is associated with a low cardiac risk, whereas increasing coronary calcium scores (CCS) are associated with an increased risk. Using a hybrid SPECT-CT system, we looked at the CCS among patients with normal myocardial perfusion to identify possibly explanatory risk factors. Methods Patients with known or suspected coronary artery disease referred for outpatient SPECT-CT myocardial perfusion imaging were evaluated. Patients were assessed using a 16-slice CT scanner for CCS and attenuation correction SPECT imaging using Tc99m sestamibi. A one-day rest stress protocol was used. Only patients with normal myocardial perfusion were evaluated. Means are given +/- standard deviation. Results There were a total of 2351 patients with normal perfusion. The minimum CCS was 0, the maximum was 6915, and the mean was 139 +/- 461. There were 209 patients with a CCS > 400 (9%) and 1301 with a score of 0 (55%), with the remaining 841 having a score of 1 to 400 (36%). A higher CCS was significantly correlated with age, hypertension, hypercholesterolemia, smoking, diabetes, male sex, and a positive family history. The strongest correlations were with age (Pearson correlation r=0.33) and hyperlipidemia (r=0.158). There was no significant correlation between CCS and obesity or postmenopausal status. After controlling for exercise capacity, there was a significant correlation between CCS and age, hypertension, and hyperlipidemia, but not with smoking status, family history, or diabetes. Conclusions Approximately 10% of patients with normal myocardial perfusion had a CCS over 400. Age and hyperlipidemia were the strongest risk factors for having a high CCS in the setting of normal myocardial perfusion. Findings support the utility of CCS in selected patient groups with normal myocardial perfusion.
SummaryPurpose: We present a case of incidentally noted giant cell arteritis in a patient undergoing 18F‐fluorodeoxyglucose (FDG) positron emission tomography (PET)/CT imaging. The patient was originally referred to PET/CT for staging of his renal transitional cell carcinoma.Methods: The patient was injected intravenously with 370 MBq of 18F‐FDG. After a 60 min uptake period, PET/CT imaging was performed from the skull base to the mid thighs.Results: A small para‐aortic node in the region of the surgical bed showed increased tracer uptake of concern for malignancy. In addition, there were several non‐calcified pulmonary nodules present, also concerning for malignancy. Incidentally noted was diffusely increased tracer uptake throughout the aorta and a thickened aortic wall on CT images. Diffuse tracer uptake was also present in the proximal branches of the aorta, including the carotid, iliac, femoral, and subclavian arteries. The patient had biopsy proven giant cell arteritis.Conclusion: Increased 18F‐FDG uptake by the aorta on PET/CT imaging is an abnormal finding that prompts a more thorough assessment for malignancy, and also can indentify important co‐morbidities in cancer patients. Evaluation of aortic uptake should be a routine practice in the interpretation of 18F‐FDG PET/CT scans.
Objectives. Statistical significance does not equal clinical significance. This study looked at how frequently statistically significant results in the nuclear medicine literature are clinically relevant. Methods. A Medline search was performed with results limited to clinical trials or randomized controlled trials published in one of the major nuclear medicine journals. Articles analyzed were limited to those reporting continuous variables where a mean (X) and standard deviation (SD) were reported and determined to be statistically significant (p < 0.05). A total of 32 test results were evaluated. Clinical relevance was determined in a two-step fashion. First, the crossover point between groups 1 (normal) and 2 (abnormal) was determined. At this point, a variable is just as likely to fall in the normal distribution as the abnormal distribution. Jacobson's test for clinically significant change was used: crossover point = (SD1 * X2 + SD2 * X1) / (SD1 + SD2). How many SDs from the mean this crossover point fell was then determined. For example, 13.9 +/- 4.5 compared to 9.2 +/- 2.1 was reported as statistically significant (p < 0.05). The crossover point is 10.7, which equals 0.71 std from the mean: 13.9 - (0.71*4.5) = 9.2 + (0.71*2.1). Results. The average crossover point was 0.66 SDs from the mean. The crossover point was within 1 SD from the mean in 26/32 cases and in these cases, averaged 0.45 SD. Thus, for 4 out of 5 statistically significant results, when applied to an individual patient, the cut-off between normal and abnormal was 0.45 SD from the mean. This results in a third of normal patients falling into an abnormal category. Conclusions. Statistically significant results frequently are not clinically significant. Statistical significance alone does not ensure clinical relevance. Citation: Heston TF, Wahl RL. "How Often Are Statistically Significant Results Clinically Relevant? Not Often." Journal of Nuclear Medicine 50, no. supplement 2 (2009): 1370–1370.
The purpose of this study was to determine if an expert network, a form of artificial intelligence, could effectively stratify cardiac risk in candidates for renal transplant. Input into the expert network consisted of clinical risk factors and thallium-201 stress test data. Clinical risk factor screening alone identified 95 of 189 patients as high risk. These 95 patients underwent thallium-201 stress testing, and 53 had either reversible or fixed defects. The other 42 patients were classified as low risk. This algorithm made up the "expert system," and during the 4-year follow-up period had a sensitivity of 82%, specificity of 77%, and accuracy of 78%. An artificial neural network was added to the expert system, creating an expert network. Input into the neural network consisted of both clinical variables and thallium-201 stress test data. There were 5 hidden nodes and the output (end point) was cardiac death. The expert network increased the specificity of the expert system alone from 77% to 90% (p < 0.001), the accuracy from 78% to 89% (p < 0.005), and maintained the overall sensitivity at 88%. An expert network based on clinical risk factor screening and thallium-201 stress testing had an accuracy of 89% in predicting the 4-year cardiac mortality among 189 renal transplant candidates.
The distal scapula and proximal humerus from each shoulder of nine adult dogs were slab-sectioned, cleaned of soft tissues, embedded in white plastic and stained black with a silver stain. These preparations were then photographed for automated, digital, morphometric analysis of subchondral bone structure. Comparison of transverse and coronal sections through the left and right shoulders demonstrated essential isometry of trabecular patterns within each bone. Comparison of the scapula and humerus revealed significant differences in bony architecture. The subchondral plate was an average of 5.6 times thicker under the glenoid fossa than in the opposing humeral head. Deeper trabecular structure also differed with the trabecular bone volume (density) in the humerus being greater than that in the scapula. This difference reflects a greater trabecular density in the humerus with comparable trabecular thickness in both bones. These structural differences are consistent with previous functional studies of the same two bones that revealed greater mechanical stiffness beneath the glenoid fossa and greater hydraulic resistance within the humeral head.
When we defend our acts by merely citing the approval of others, we fall into the “ochlocratic trap.” This is the mob approach to ethics, and it doesn’t work. Although Copernicus knew our solar system was heliocentric, most of his contemporaries thought it to be geocentric. Likewise, proper ethical action will often meet with public disapproval. Ethics can be compared to science. Both: Start out with a hypothesis Collect data Seek the truth Using the scientific method in ethics would involve: Formulation of a hypothesis Testing the hypothesis with a vigorous, open debate Deciding which parts of the hypothesis withstood the debate In science and ethics, an individual or a pair of people working in close unison make major breakthroughs. Rarely does the consensus do it. Ethics can be approached scientifically. With this approach, we find no security by following the majority. This way demands we always consider new evidence or more logical thinking, not a democratic approach. The ochlocratic trap must be avoided if we are to make ethical advances. Citation: Heston TF. The Ochlocratic Trap. MSLaneous ~1989, p 18.
This song is about cultivating gratitude for our gifts, overcoming obstacles to serving others, and how a dying patient's courage reminds us that even in dire circumstances, we can provide comfort and strength. Healing Gifts In the still of the night, a melody rises, Notes of comfort to soothe weary souls. My hands are rough, but with care, I apply Balms to the hurting, making spirits whole. Mysterious the workings of mind and soul, Questioning always as we grow old. Yet with knowledge and skill, we ease our pain, And music's sweet waves bring joy again. A teacher shares their gifts of learning, Lighting a spark of insight yearning. And songs that speak of what's unsaid, Lift the sick from their weary beds. So let's give thanks for each talent and skill, For gifts that cure in a friendly way. We share our blessings as best we are able, We'll scatter darkness; we'll bring brighter days. With compassion and care, we dry up tears, With love in our songs, we calm anxious fears. Our gifts flourish when freely given, They bring hope, joy, and reasons for living. So use your talents well, whatever they be, To serve one another, to help those in need. Lift your soul, offer care and might, Fill the world with your healing light.