Evaluating study results: statistical inference, hypothesis testing and significance

Part of a comprehensive series showing how to evaluate a clinical study or research paper. This article focuses on types of potential error and statistical power.
Dark purple background with Sum greek symbol, calculator, syringe and p=0.013, graph with steps and bell curve, number 6

By the end of this article, you should be able to:

  • Describe the process of statistical inference, the importance of statistical power and the factors affecting it, as well as the interpretation of P values;
  • Discuss the potential errors in statistical testing and in interpreting study findings: Type I error (i.e. alpha) and Type II error (i.e. beta).

This article is the part of a comprehensive series exploring how to evaluate clinical studies when addressing information needs using a five-step process:

  1. Identifying study type or design;
  2. Appraising the journal, authors and study purpose;
  3. Critiquing the methods used;
  4. Understanding basic statistical tests.
  5. Analysing study data and results, and the discussion section;

This article explores the fourth step: understanding basic statistical tests. It is recommended that you read it in conjunction with:

The previous article covered measures of central tendency and variability commonly used to describe study data, as well as confidence intervals used to estimate the corresponding values in the target population. The confidence interval (CI) is one part of statistical inference, which is the overall process used to draw conclusions about the underlying population from the evidence (i.e. data) obtained in the smaller study sample. In this article, hypothesis testing is discussed as a component of statistical inference along with the importance of statistical power, interpreting results such as probability or P values from statistical tests, potential errors in statistical testing and statistical versus clinical significance.

Hypothesis testing

Statistical inference incorporates ‘hypothesis testing’, which is the process of determining whether the data gathered support the study’s hypothesis. In a clinical study examining therapy efficacy, hypothesis testing helps determine the likelihood of whether an observed treatment effect might be explained by chance variation or some factor other than the therapy. Hypothesis testing is generally used in a study — either to reject or fail to reject the null hypothesis, based upon the results of statistical testing. In the article ‘Assessing authorship, study purpose, and journals when evaluating clinical studies’, we reviewed the null hypothesis. In a controlled experimental clinical trial, the null hypothesis assumes there is no difference between the treatments studied or no difference between comparisons made before and after therapy within the same study group. Thus, statistical hypothesis testing is used either to reject, or fail to reject (i.e. accept), the null hypothesis. If not clearly stated in a study, the null hypothesis is determined from the study’s objective.

There are several important concepts in hypothesis testing:

  • Probability (P) values;
  • Type I error, alpha (α);
  • Type II error, beta (β);
  • Statistical power;
  • Factors influencing power.

This article will review each of these and describe how they interrelate.

P values

Statistical tests are used to generate a probability or P value, which is used to test the null hypothesis. The P value provides the likelihood (i.e. probability) that the study results seen would have occurred if the null hypothesis were true. Another way to think of the P value is that — assuming there is no actual treatment effect — the P value is the probability of observing a result at least as large as the one found in the study because of random or chance variability. Thus, if the P value is very small, it is very unlikely that the null hypothesis is true, and there is a very small likelihood that the difference observed would have resulted from some random factor. Meaning, as the P value becomes smaller, the treatment will be more likely responsible for the difference seen, and the null hypothesis should be rejected, which is less likely to be true. By contrast, the larger the P value, the greater the likelihood that the null hypothesis is true and the greater the likelihood that random or chance variability produced the effect seen.

Suppose a study measures the difference in glucose concentrations between a drug and placebo and reports a P value of 0.00001 for that difference. Can a very small P value prove that the drug caused the difference in glucose concentrations seen? The answer is no. A P value provides the likelihood that the null hypothesis was true for that finding. The P value in this example indicates that the likelihood that the null hypothesis was true for the difference in glucose concentrations is very small — only 1/100,000 (0.00001) — so the likelihood that random or chance variability produced that finding was also very small. Keep in mind that while certainly a small number, there is still a very slight possibility that the effect did not result from the therapy (i.e. null hypothesis was really true).

To apply a study’s findings to patients, clinicians must decide whether to fail to reject (i.e. accept) or reject the null hypothesis in a study. However, P values are on a continuum that can range between 0 and 1. Is there a commonly accepted ‘cut-off’ for the P value below which the null hypothesis should be rejected and above which it should be accepted? Yes, a value of 0.05, which is referred to as the level of significance or alpha, is generally used for this cut-off in clinical studies. When P < 0.05, it is concluded that the risk of the null hypothesis being true (any difference is a chance occurrence and not an actual treatment effect) is acceptably small (less than 5/100 or 1/20), so the finding is termed ‘statistically significant’. When P ≥ 0.05, which is the likelihood that the null hypothesis is true, is considered unacceptably large, the finding is not statistically significant.

Key points about the P values

  • The cut-off for statistical significance (alpha) is usually set at 0.05;
  • When P < 0.05 the following are concluded:
    • The finding is statistically significant;The null hypothesis is likely false (i.e. reject it);
    • A real treatment effect was found.
  • When P ≥ 0.05 the following are concluded:
    • The finding is not statistically significant;
    • The null hypothesis is likely true (i.e. do not reject it, accept it);
    • A real treatment effect was not found.
  • If a study does not find a statistically significant difference between treatments, it does not automatically mean that the treatments are the same (i.e. equivalent), since other important differences might exist between them. Conclude only that the study did not find a significant difference in the outcome measured;
  • P values do not indicate the clinical importance of a difference found. They simply provide the probability that the null hypothesis is true, and the difference seen might have resulted from random/chance variability. If one P value reported for a measure in a study is 0.02 and another P value in that study is 0.002, it does not mean that the smaller P value is 10x ‘better’ than the other. It means that the likelihood that the null hypothesis is true is 10x less for the outcome measure with P = 0.002 compared to the measure with P = 0.02.

Worked example 1: statistical significance

A study compared the efficacy of oral mesalazine (n = 28 patients) with topical mesalazine (n = 31 patients) for the treatment of distal ulcerative colitis. Following 2 weeks of therapy with either agent, the clinical response rate was 43% with oral mesalazine versus 58% with topical mesalazine (P = 0.003).

Is this therapy difference statistically significant?

Yes. P = 0.003 indicates that the risk the null hypothesis is true, or the probability that the difference in clinical response rates might have resulted from random variability, is only 3/1,000. This is acceptably low since it is less than the cut-off of 0.05. The study difference in response rates between topical and oral mesalazine is concluded to be statistically significant.

P values provide a guide for interpreting whether a study’s findings were likely owed to the therapy. As stated earlier, a statistically significant finding (P < 0.05) might still have resulted from something other than the treatment. Conversely, a non-statistically significant finding (P ≥ 0.05) does not prove that the treatment had no effect, since the study might have failed to identify a real treatment difference. So, it is possible that the conclusion made from a reported P value is in error or wrong. Examples of these incorrect conclusions are:

  • A finding is concluded to be statistically significant and resulted from the treatment given, when the treatment was not really responsible;
  • A finding is concluded to not be statistically significant when the treatment actually caused the difference seen (i.e. a real treatment effect was missed).

These two types of erroneous conclusions are termed ‘Type I error’ and ‘Type II error’, respectively.

Type I error

We mentioned that the level of significance, which is referred to as alpha (a), sets the cut-off for concluding whether a finding is statistically significant, with alpha usually set at P =0.05. A Type I error is defined as rejecting the null hypothesis when it is really true (i.e. a false positive effect). Thus, it is only possible when P is less than the cut-off of 0.05 (i.e. when we reject the null hypothesis). Alpha, then, establishes the cut-off for determining Type I error probability. Think of a Type I error as wrongly concluding that a therapy produced a significant effect when the observed difference probably resulted from random or chance variability.

The P value calculated from a statistical test is used to determine statistical significance based on whether it is above or below the specified level of significance alpha.

Suppose a study comparing two antihypertensive drugs finds a difference of 8 mmHg in diastolic blood pressure between treatments (P = 0.03). Since P < 0.05, the difference between treatments is:

  • Statistically significant;
  • Concluded to represent a real treatment effect.

Is there any way to identify if a Type I error occurred with this conclusion? No, we will never know if the blood pressure difference was truly from the treatment or if it resulted from random or chance variability and if a Type I (i.e. false positive) error was made. All we can say is that the risk of Type I error is acceptably low when P < 0.05.

Key points about Type I error

  • Studies should state the level of significance (i.e. alpha) they are using, usually 0.05. If you do not see this clearly stated in the text, the alpha used will generally be in the power calculation provided in the study (refer to discussion of Type II error and power);
  • The greater the number of statistical comparisons made in a study, the greater the probability that at least one comparison will be statistically significant owing to chance alone (i.e. a Type I error) when no true differences exist. For example, if 100 different comparisons are made with α = 0.05 — assuming all null hypotheses are true — approximately 5 statistically significant results would be expected by chance. Think of running through a golf course holding a metal golf club during a thunderstorm. The more times you do this, the more likely it is that you will be hit by lightning purely by chance. You may never be hit by lightning, but the likelihood of it increases with each dash through the golf course. Sometimes when investigators are performing many statistical comparisons, they will set the alpha level for the study lower, such as 0.01 or 0.025. Since this means the findings will only be statistically significant when the P is less than 0.01 or 0.025, reducing the probability of Type I error compared to the standard cut-off of 0.05;
  • Be cautious when interpreting the findings from studies that perform many comparisons. Statistical procedures can be used to reduce the risk of Type I error from multiple comparisons and should be performed when needed.

Type II error

The opposite of a Type I error is a Type II error. Type II error occurs when it is wrongly concluded that there was no treatment effect or that any difference seen in an outcome measure was owed to chance and not from the therapy. We fail to reject (i.e. accept) the null hypothesis when it is false and should be rejected. Type II error is only possible when P ≥ 0.05, or greater than the selected alpha cut-off, since that is when the null hypothesis is accepted. Type II error can also be thought of as missing a real treatment effect because the study failed to find a statistically significant difference (i.e. false negative). Beta (β) is defined as the probability of a Type II error, or the probability of accepting the null hypothesis when it is false and should be rejected.

How can one determine the likelihood of Type II error for a study finding with a P > 0.05? Statistical power helps us answer this question. In a clinical study, power is the likelihood that a study will identify a specified treatment effect or difference as statistically significant when an actual treatment effect is present. In other words, power is the probability of correctly rejecting a false null hypothesis. Mathematically, power = 1 − beta. Thus, another way of thinking of power is that it is the likelihood of not making a Type II error.

What is the acceptable cut-off value for power — and beta — in a study to minimise the likelihood of a Type II error? A power of ≥ 80% is desired. Since power = 1 – beta, if the desired power is 80% (0.8) or greater, the desired beta is 20% (0.2) or less. As power increases, the likelihood of Type II error decreases so a greater power lowers the probability of a Type II error.

Key points about Type II error

  • Too small a sample size is usually the main reason for inadequate power for a specified study analysis. Prior to beginning a study, the investigators should calculate the number of patients needed to have a power of at least 80%. The power calculation should be reported for the reader;
  • The best way to increase the statistical power in a study is to increase the number of study patients (i.e. sample size);
  • Check the effect size used for the power calculation to see if it is a reasonable value to use for a clinically significant effect. Sometimes investigators will select too large an effect size in their power calculation just to make sure that the power is at least 80%. If this happens, the study might have insufficient power to detect smaller but still clinically important differences as statistically significant, thereby increasing the Type II error risk;
  • Power is usually calculated for only a specific outcome measure(s), which is generally the primary outcome measure in a study. Check to see what measure(s) the power apples to.;
  • Even when the power is > 80% at the start of a study, watch for things that could decrease the power by the end of the study, such as:
    • Drop-outs. This refers to subjects who do not complete a study. They can decrease the study’s sample size if the data from dropouts are not included in the analyses of results, such as when the per protocol data handling method is used;
    • Subgroup analyses. This refers to analyses of smaller groups within the study to determine if certain subject characteristics might affect the findings. For example, the investigators might want to see if the therapies used produced different results in the female patients or only in patients who had diabetes or another condition. Since these subgroups would have less persons in the analyses compared to analysis of the entire study sample, sample size is reduced in the subgroup analyses.
  • If a study does not report any power, the risk of Type II error for non-statistically significant findings (e.g. P > 0.05) is unknown and this risk should be considered.

The Table below summarises P value interpretation and types of error.

Table: Summary of P values and types of error

Calculating power

How is power calculated? There are equations and formulas available to calculate power, which is beyond the scope of this discussion. However, it is important for clinicians to understand four main factors affecting the calculation of power and how they affect power. These factors are summarised below:

  • Sample size (i.e. the number of patients in the study). Sample size should ideally be calculated prior to beginning the study to ensure that sufficient power is present:
    • As sample size increases, power increases;
    • As sample size decreases, power decreases;
    • This is the easiest factor for investigators to adjust to increase the power for a study measure.
  • Effect size. The size of the difference in the outcome measure that, if present, one would like to identify as statistically significant:
    • For a given number of study patients (i.e. sample size), as the effect size used in the power calculation increases, the calculated power increases. As the effect size used in the power calculation decreases, the resulting power decreases. For example, in a study of 100 patients, it would be easier to detect a treatment difference (i.e. effect size) of 15 mmHg in diastolic blood pressure than a smaller 5 mmHg diastolic blood pressure treatment difference. The risk of missing the larger difference as statistically significant (i.e. making a Type II error) would be less than with the small difference; thus, the power would be greater for detecting the larger difference;
    • If a very small effect size is chosen for the power calculation, more patients will need to be included in the study to have adequate statistical power compared to a large effect size. Keep in mind that outcome measures have certain inherent variability. For example, a person’s blood pressure (BP) might vary by a small amount — sometimes higher, other times lower — depending on the time of day taken. There is always a certain amount of background ‘noise’ or variability that must be overcome to detect a real treatment effect. If there are only a few patients in a study, this ‘noise’ might make small effects harder to identify. When many study patients are included, slightly lower or higher values from person to person tend to balance out, so small but real treatment differences can be easier to identify;
    • The effect size used in a power calculation should usually be the smallest difference that the investigators consider to be clinically important.
  • Alpha (i.e. risk of Type I error). An α = 0.05 is usually selected as the acceptable cut-off for Type I error in the power calculation;
  • Variability of the outcome measure in the population, which is estimated by the SD: 
    • This refers to the inherent variability of the outcome measure in the population. For example, if serum potassium is the outcome measure, what is the normal SD (e.g. variability and spread of data) of serum potassium measurements in a large group of subjects?
    • The value for the SD used in a study’s power calculation is identified from previous studies involving that outcome measure. The investigators have no control over this value. As a result, many times investigators do not include the SD used when providing their power calculation in a published study;
    • If the baseline SD in the underlying population is large, a larger number of patients will need to be enrolled in the study to identify a significant treatment effect.

Investigators should ideally calculate power before the study begins since they need this calculation to determine the number of patients (i.e. sample size) to enrol for a desired power. The investigators should also include the values used in their power calculation. When present, the power calculation is usually found in a published study’s methods section.

Worked example 2: Type of error

Complete the following sentences by filling in the blanks with the correct words:

  1. When P < 0.05, a ________ (Type I, Type II) error is possible. With this type of error, one would ________ (reject, fail to reject) the null hypothesis when it is really ______ (true, false). This can also be thought of as a false- _______________ (positive, negative) finding.
  2. When P ≥ 0.05, a ________ (Type I, Type II) error is possible. With this type of error, one would ________ (reject, fail to reject) the null hypothesis when it is really ______ (true, false). The likelihood of this type of error is referred to as _______ (alpha, beta). This can also be thought of as a false- _______________ (positive, negative) finding.

Answers

  1. Type I, reject, true, positive;
  2. Type II, fail to reject, false, beta, negative.

Worked example 3: Statistical significance

Investigators compared a new drug to placebo for the prevention of headaches. The power of the study was not reported. A total of 30 patients who experienced frequent headaches were enrolled, with 15 patients assigned to receive placebo and 15 patients assigned to receive the new drug. Following 10 weeks of therapy, there was a 40% reduction in headache frequency with the new drug compared to placebo (P = 0.07). It was concluded that the new drug did not appear to be superior to placebo for headache prophylaxis.

Was the reduction in headache frequency statistically significant? Is a Type II error a possibility for the finding in this study?

Since P is greater than 0.05, the reduction in headache frequency with the new drug compared to placebo is not statistically significant. This means that the null hypothesis of no difference between study treatments is accepted (i.e. failed to be rejected).

Yes, Type II error is a possibility for non-statistically significant findings. Note that a reduction of 40% in headache frequency appears to be a fairly substantial treatment effect, even though it was not statistically significant. Concluding there was no difference between the new drug and placebo might be a wrong conclusion here (i.e. a Type II error). Power provides the likelihood of not making a Type II error (β). This study did not report power, so we do not know if it was acceptable (80% or higher). Thus, the actual likelihood of Type II error is unknown. Since there were a rather small number of study patients (i.e. a small sample size), the power could be low. A larger study should be performed.

Worked example 4: sample size and significance

A study examined the cure rates from two different antibiotics (i.e. Drug A and Drug B) used to treat pneumonia. A sample size of 200 patients was needed for a power of 80% to detect a difference of 5% in the cure rates as statistically significant. A total of 205 patients were enrolled in the study. Once the study was completed, the investigators decided to examine whether the antibiotics produced different results (i.e. cure rates) in the patients who had diabetes (i.e. 60 patients in the Drug A group and 55 patients in the Drug B group). The analyses in patients with diabetes found that the cure rate with Drug A was 96% versus a cure rate of 86% with Drug B (P = 0.14).

Was Type II error a possibility for the cure rate comparison in patients with diabetes?

Yes. P = 0.14 is greater than 0.05, so the antibiotic cure rate difference was not statistically significant. The null hypothesis of no difference between drugs is accepted, so Type II error is possible. A power of 80% was reported for the cure rate analysis, but this was calculated for a sample size of 200 patients. In the subgroup analysis of patients with diabetes, the total sample size was only 115 patients (60+55). This means the power would be less than the desired 80% minimum, which increases the value of beta (i.e. Type II error risk) above the desired maximum of 0.2 or 20% (1 – power = beta). It is also very important when evaluating the possibility of Type II error to look at the actual treatment effect seen. There was a 10% difference in cure rates in the diabetes analysis. This appears to be a potentially important difference that might have been erroneously ‘missed’ as statistically significant (i.e. false negative).

Statistical significance versus clinical significance (i.e. importance)

There is a difference between statistical significance and clinical significance. Statistical significance is present when the P value is less than the designated alpha cut-off, usually 0.05, and the null hypothesis is rejected. The likelihood that chance alone was responsible for the effect seen is acceptably low. However, even if a difference between treatments in a study is not likely a chance finding, does that automatically mean that the difference is large enough to be clinically important? Not necessarily. A study with many patients may find a fairly small treatment difference to be ‘statistically significant’, which is unlikely to have resulted from chance, but too small to have a meaningful effect in patient care or clinical practice. For example, suppose a large study finds a mean difference in diastolic BP of 2.3 mmHg between two antihypertensive treatments to be statistically significant (P < 0.05). Would a difference of 2.3 mmHg in diastolic BP be clinically important for most patients? Probably not. When a study finding is statistically significant, always look at the actual treatment effect or difference to determine whether it would be clinically significant or important in practice.

Worked example 5: clinical significance

A study compared the efficacy of a new drug to zolpidem for insomnia therapy. 420 patients were enrolled in the study and randomly assigned to receive the new drug (210 patients) or zolpidem (210 patients). After two weeks of therapy, the new drug was found to reduce sleep latency (i.e. time needed to fall asleep) by a mean = 2.3 minutes (P = 0.028).

Is this finding statistically and clinically significant?

The mean difference of 2.3 minutes is statistically significant since the P value is less than 0.05. However, the actual difference of only about 2 minutes in the time to fall asleep is not likely to be clinically important for most individuals given a normal sleep latency of about 10–20 minutes.

Key points about statistical and clinical significance

  • How should we interpret a ‘statistically significant’ finding?
    • ‘Statistical significance’ simply means that a study’s finding was not likely to result from chance. The null hypothesis of no treatment difference is rejected;
    • A finding is statistically significant when P < 0.05 (sometimes a smaller P value might be used, but a cut-off of 0.05 is most common).
  • Can a study finding be concluded to be clinically significant if it is not statistically significant?
    • No. There first needs to be statistical significance to conclude that a study finding is clinically important. Without statistical significance, the finding has a higher than acceptable likelihood that the null hypothesis is true and was not a real treatment effect. Thus, no definitive conclusions about clinical significance can be made in the absence of statistical significance;
    • If a study finds a difference large enough to be of clinical importance, but it is not statistically significant, consider whether a Type II error occurred — meaning, a real treatment effect was missed. Most likely this type of error would result from insufficient power in a study and too small a sample size. When a study finds a large treatment difference that is not statistically significant, especially when power is not reported or is too low, conclude that further study is needed.
  • Can a study finding be statistically significant but not clinically significant?
    • Yes. With a large number of patients, a small treatment effect might be statistically significant, but the size of the effect is too small to be of clinical importance in actual practice.
  • Can one determine if a study finding is statistically significant if there is a CI reported but no P value is given?
    • Yes. If a CI includes the value that indicates no treatment effect or difference, the finding will not be statistically significant. For a difference in a specific outcome measure (e.g. BP following therapy, lab values before and after therapy, comparison of success rates between treatments, etc.), a value of 0 would equal no treatment effect. For example, suppose a study reported cure rates from Drug A and Drug B as 72% and 68%, respectfully, with a difference between cure rates of 4% (95% CI = –2% to 10%). The CI for the difference in cure rates is interpreted as: 95% confidence that in the population, the cure rate with Drug A could be from 2% less than, to 10% greater than, the cure rate for Drug B. Since the value of 0 (no difference in cure rate between Drugs A and B) is within this CI range, both drugs could have equal population cure rates. Based on the CI, the difference in cure rates would not be statistically significant — one does not need to look at the P value to determine this.

Key points

  • The P value provides the likelihood (i.e. probability) that the study’s result would have occurred if the null hypothesis were true. Or it gives the probability of observing a result at least as large as the one in the study from random or chance variability if there was no actual treatment effect;
  • P values do not indicate the clinical importance of a finding. One cannot conclude that a measure of P = 0.0003 is more ‘important’ than a measure with P = 0.03. Rather, the smaller P value simply indicates the null hypothesis of no treatment effect is less likely to be true;
  • Alpha (i.e. level of significance) provides the cut-off for statistical significance and the probability of a Type I error when the null hypothesis is true. It is usually set at 0.05 for power and statistical calculations, with P < 0.05 considered statistically significant for that alpha;
  • Beta is the likelihood of making a Type II error and is acceptable when less than or equal to 0.2 (20%);
  • Power is defined as the likelihood of not making a type II error — it is calculated as 1 − beta and is acceptable when ≥80% (0.8);
  • The easiest factor that investigators can change to increase a study’s power is the sample size. Increasing the sample size will increase power. Anything that decreases the sample size during the study, such as smaller subgroup comparisons within the study and patient dropouts that are not included in the data analyses (i.e. per protocol data handling), will decrease the power calculated at the start of the study;
  • A finding cannot be concluded to be clinically significant unless it is first statistically significant. However, a finding can be statistically significant but too small to be of clinical importance.

How to apply to practice

  • The smaller the P value, the more likely it is that the treatment was actually responsible for the finding;
  • If P < 0.05 — unless the investigators indicate a different alpha — the finding is statistically significant and the null hypothesis is rejected;
  • If P ≥ 0.05, the finding is not statistically significant (i.e. the null hypothesis is accepted) and consider whether a Type II error might have occurred;
  • Look for a reported power of at least 80% in a study. If the power is not adequate, it is most likely because not enough patients were included;
  • For statistically significant findings, look at the actual size of the treatment effect. If large enough to be of clinical importance, the finding is clinically significant;
  • If a non-statistically significant finding looks large enough to be of clinical importance, consider whether a Type II error might have occurred (i.e. missing a real treatment effect). Make sure the power for that outcome measure was at least 80%. If the power was < 80% (at start of study or through decreases in sample size during study) or not reported (missing or calculated for a different outcome measure), Type II error is a potential problem.

Acknowledgements

This article was adapted from Drug Information and Literature Evaluation, Second Edition, previously published by Pharmaceutical Press.

A full list of resources and materials used to prepare the book can be accessed from the bibliography page.

Last updated
Citation
The Pharmaceutical Journal, PJ September 2026, Vol 317, No 8013;317(8013)::DOI:10.1211/PJ.2026.1.428151

    Please leave a comment 

    You might also be interested in…