How to interpret clinical trial results for Long COVID
Even when results are “statistically significant,” the treatment might not actually work.
Heather Hogan / The Sick Times. Sources: Shutterstock, Canva Pro
Key points you should know:
There are currently no approved treatments for Long COVID. Experts say that the diversity of symptoms and severity makes it more difficult to run clinical trials.
Trials of new drugs and devices are stratified by phases. Only phase 3 trials are designed to prove that a treatment is effective.
When reading a study, experts say it’s important to note whether the primary outcome, which determines whether a trial succeeds, is both statistically significant and clinically meaningful. A statistically significant outcome does not necessarily mean an intervention will meaningfully improve one’s daily life.
To understand whether an off-label medication may benefit you, experts say it’s important to understand whether the studied population has similar symptoms and clinical presentation to you, and to work with a medical provider to ensure a treatment is trialed safely.
Simon Spichak also wrote a resource for evaluating clinical trial results that accompanies this story.
With no proven treatments or cures, many people with Long COVID are turning to off-label medications, supplements, and medical devices to treat their symptoms. There are numerous options — and some of them even have clinical trials underway or are reporting results.
But even if one study suggests a treatment might work, another might contradict it. And depending on who’s included in the study, a treatment might not work for everyone with Long COVID.
Long COVID’s heterogenous nature means vast variations in symptoms and severity that may also make it more difficult to study and design good treatment trials compared to “more defined entities like rheumatoid arthritis or asthma,” Nita Jain, a biotech founder and a patient representative on the National Institutes of Health’s RECOVER program, told The Sick Times.
So how do you separate the strong studies from the weak ones, and how do you spot the bunk?
For five years, I’ve covered Alzheimer’s and dementia trials, watching both promising drugs and those based on speculative or outright fraudulent data fail in late-stage trials. As more Long COVID trials start to share their results, I wanted to speak with experts and share the tricks we’ve learned about interpreting trials to help readers critically navigate the data coming.
What kind of study is it?
Every promising new treatment starts out as a set of basic laboratory experiments run by overworked, stressed-out graduate students. Before they’re ever tested in humans, supplements, drugs, and devices go through the preclinical gauntlet: testing in vitro on cells grown in petri dishes and then in various animal models. Most treatments that succeed at this stage don’t work out in human trials.
Since testing treatments in humans is expensive, researchers often turn to real-world data and electronic health records for clues. Studies that follow people over time, with no randomization or control group, are called observational studies. They come in different sizes: from small case studies and series to larger studies following thousands or even hundreds of thousands of individuals. In Long COVID, such studies have examined treatments like stellate ganglion block and Vitamin D.
But treatments that are linked to better health outcomes in observational and case studies seldom work out when tested rigorously — trials comparing the antiviral Paxlovid to placebo for Long COVID have so far failed to show efficacy, despite some observational trials suggesting it as a potential treatment.
“Treatments are not given randomly in real life,” explained Jeffrey S. Morris, director of biostatistics at the Perelman School of Medicine at the University of Pennsylvania. Studies that use different methods to adjust for these biases, and multiple different analyses to see if their results hold up, are stronger, he added.
To understand if a new drug for Long COVID works, scientists must rely on clinical trials. Morris described these as experiments that investigate how a drug is absorbed in the body, its safety, and its efficacy.
“At each step, there’s attrition,” said Morris, meaning that fewer and fewer drugs make it to each successive trial stage. Since it can cost hundreds of millions of dollars to develop and test drugs, companies will discontinue them if the results aren’t promising.
“At each step, there’s attrition,” said Morris, meaning that fewer and fewer drugs make it to each successive trial stage.
The earliest studies are phase 1 trials: small trials that test the safety of a new supplement, medication, or device. Phase 1 may be skipped if researchers are repurposing something that has already proven safe. These studies cannot tell researchers whether a drug works, even if they might hint that it’s promising.
The next stage, phase 2, helps provide a sign of whether the drug might work.
High-quality trials require randomizing participants to a treatment or placebo-controlled arm. For medical devices (or certain procedures), studies use a sham control that looks, sounds, and feels similar to the actual device but does not offer the treatment under study.
The best trials include double-blinding, in which neither the participants nor the researchers know who’s receiving the actual treatment. Many will even test how successful the treatment blinding is by surveying participants at the end of the study.
But even a successful phase 2 isn’t enough: Many drug trials that succeed at this stage ultimately fail in larger, more definitive trials.
Phase 3 is the final stage for new treatments. These are large studies conducted across multiple hospitals or research centers that pit a new drug or medical device against a placebo. “You don’t have the confidence to claim it’s effective if there hasn’t been a phase 3 trial” Morris said.
Lastly, a phase 4 trial may add more safety and effectiveness data or test whether an approved treatment or supplement could be repurposed.
You can typically find a study’s phase in the “methods” section or on the trial registry site ClinicalTrials.gov, which will also show you a record of any changes. It’s a red flag if the study is registered after testing has already begun or the protocol substantially changes after enrolling participants. It may mean that researchers are changing how success is measured in the trial, undermining its scientific integrity.
For example (see bottom right for stage information):
Example of a clinical trial record on clinicaltrials.gov, for a trial of the drug pembrolizumab. The study’s phase is highlighted with a red circle and arrow.
A crash course in statistics
Statistical tests don’t tell you if a treatment worked. They test a null hypothesis: that there is no real difference between the control and treatment groups.
To make this determination, researchers calculate a number called the p-value. The p-value tells you the chances you’d see these results if the treatment didn’t work as intended.
A common p-value threshold is 0.05. If the p-value falls below the threshold, the result is statistically significant, meaning there is a less than 5% chance that the changes you’re seeing between the groups is the result of random chance. When p is less than 0.05, you’ve got yourself a potentially successful trial.
But if you run enough statistical tests, researchers will find a positive outcome by pure chance alone even if the intervention is completely nonsensical. When researchers run an additional analysis that wasn’t preplanned, but present it as an important or positive result without caveats, that’s called p-hacking — and it’s widely considered unethical.
The infamous PACE trial reported that graded exercise therapy was effective for myalgic encephalomyelitis (ME) through another form of p-hacking where they changed their trial outcomes while the study was underway, which would make treatment appear effective.
As a result, regulators expect researchers to determine what hypotheses are being tested and how they will measure success before the study starts. The primary outcome is the main hypothesis being tested. The secondary and exploratory outcomes provide more information about the effects of a treatment, but don’t provide any definitive insights if the primary outcome fails.
Sometimes a study might fail on its primary outcome, but if there are well-defined subgroups where it appears to work, “you can do a repeat study in that subgroup, and maybe you can validate it,” said Morris. But he cautions these results might not hold up in further testing.
Take the REGAIN study on oxaloacetate for Long COVID. The primary outcome in the trial registry is the Chalder Fatigue Score, a fatigue survey, and it showed no difference between treatment and control groups. But the researchers concluded that the trial’s results were promising based on other outcomes.
Helen Brownlie, an ME patient-researcher, cautioned that researchers may overinterpret their findings: “Do not simply accept the authors’ summary and interpretation.”
Scientific sleuths and patient communities sometimes spot statistical problems that aren’t obvious to nonexperts. PubPeer is a post-publication review platform where researchers may describe potential issues with studies. Science for ME (or S4ME) is another forum that discusses and flags potential issues with trials and study design.
Example of a primary outcome measure as listed on clinicaltrials.gov, for the REGAIN study on oxaloacetate for Long COVID.
Check for dropouts. Do a larger number of people in the treatment group opt not to finish the study compared to the placebo group? Do the researchers explain why? High dropout rates could be a sign of severe side effects.
Even when statistics show a significant difference, the impact of a treatment on people’s daily lives might be so minuscule that it’s irrelevant. For example, a recent Long COVID clinical trial reported that the antidepressant fluvoxamine led to a statistically significant reduction in the primary outcome at 60 days. Compared to the placebo group, the treatment group showed a small improvement in self-reported fatigue. But in trials using similar fatigue scales for other, better-defined conditions like chronic hepatitis C infection, lupus, and multiple sclerosis, such small improvements are not considered clinically meaningful, or not noticeable in changing people’s lives.
Who is being tested and how are outcomes measured?
For individuals considering whether a new treatment might be promising for them, understanding who was tested in the trial is important. Experts recommend reading the inclusion and exclusion criteria of the trial, which describe this in detail.
Long COVID has many broad and specific definitions, which can make it difficult to test whether a treatment works, Jain said. Some trials may test only a narrow, unrepresentative slice of people with Long COVID. Most trials still exclude people with severe forms of the disease.
For instance, some studies that claim to study Long COVID actually use participants who were hospitalized for SARS-CoV-2 and represent a distinct population, said Alba Azola, a Johns Hopkins physician who treats people with Long COVID and ME.
Defining subgroups before starting the study allows researchers to test specific questions and determine whether a biological mechanism might be directly causing symptoms for some individuals. On the other hand, if the primary outcome measure being used to judge the effectiveness of a treatment is a general survey, like a fatigue or quality-of-life measurement, without any biological measures, it isn’t very informative about underlying mechanisms, Jain said.
Without subgrouping beforehand, even the negative results of trials become less informative. The preliminary results of the RECOVER-AUTONOMIC ivabradine trial suggested no improvement in quality-of-life measures, “but you have a lot of different subtypes of POTS in this population,” said Jain. However, some drugs, like migraine medications, are often approved solely on the basis of patient-reported outcomes.
Though the DePaul Symptom Questionnaire (DSQ) is often used to measure post-exertional malaise (PEM) in trials, some ME experts have emphasized that it’s a screening tool and not an outcome measure. It may signal that someone has PEM, but without clinical validation, it could lead to false positives. Other questionnaires like FUNCAP, which were developed with the input of people with ME, might better reflect symptoms and functional ability.
Without subgrouping beforehand, even the negative results of trials become less informative.
Other details that matter
The most robust trials are run across multiple clinical centers. Morris looks for variability among the results across different centers, which could provide clues about how well study inclusion criteria are applied across the board and any potential biases that may have influenced the results. He’s more confident about studies that show equivalent outcomes across multiple trial sites.
Sometimes the dose chosen in the trial might not be tolerable for some individuals with Long COVID, said Jain. She took part in the baricitinib REVERSE-LC trial, and though she’s happy that the drug was being tested, the dose she was assigned to landed her in the hospital.
Funding sources also play an important role. Sometimes, researchers receive funding from pharmaceutical or medical device companies. While that does not mean the research is invalid, Zeest Khan, an anesthesiologist with Long COVID, is cautious. Sometimes these companies promote these products without disclosing this conflict of interest, and she said they may exaggerate the findings without transparency.
“For people with Long COVID who have cognitive injury, it is sometimes impossible to dig deeper than the headline,” said Khan.
One obvious red flag is if a treatment is described in a way that’s too good to be true, Khan said —for example, “Everyone who takes this drug improves.”
For people with Long COVID who have cognitive injury, it is sometimes impossible to dig deeper than the headline.Zeest Khan, anesthesiologist with Long COVID
Navigating the uncertainty
When evaluating the evidence, Morris said it’s important not to “cherry-pick one paper and ignore the rest.” You need to look at a lot of studies critically, recognizing the strengths of weaker studies and the weaknesses of the strong ones, to help you come to a good conclusion.
Khan hopes people are able to find a medical practitioner they can work with to “tailor your treatment for your specific needs.”
She also recommends adding in one new treatment at a time so that you and your doctor can monitor changes and have a better idea of whether it works. In addition, it’s important to ask doctors about what side effects might be expected and consider the financial burden, accessibility, and time to travel for a treatment.
Simon Spichak is a Toronto-based science and health writer with a MSc in neuroscience. His work has been published in Being Patient, National Geographic, MIT Tech Review, The Guardian’s Scientific Observer, The New York Times, and other outlets. He was a recipient of the 2025 National Press Foundation’s Rare Disease Reporting Fellowship. He is the founder of a low-cost online therapy clinic called Resolvve and runs a newsletter about underreported health and disability issues in Canada.
Check out Simon’s resource for evaluating clinical trial results here.
All articles by The Sick Times are available for other outlets to republish free of charge. We request that you credit us and link back to our website.
The Sick Times is dedicated to independent Long COVID journalism, without denial, minimizing, or gaslighting.
By donating to this nonprofit publication, you:
Help us produce more unique news and commentary stories like this one.
Support free news access for everyone impacted by Long COVID, regardless of their financial situation.
Keep this crisis in the spotlight, as mainstream media wants to put the pandemic in the past tense.
Make a monthly tax-deductible donation:
Not ready to give monthly? A one-time donation of any amount also has a profound impact.
More science stories