Understanding risk and clinical utility, dropouts, data handling and the discussion section when evaluating studies

Part of a comprehensive series showing how to evaluate a clinical study or research paper. This article focuses on measures of risk, data handling and the key points of a study’s discussion section.
Lighter purple background with syringe, black and white silhouettes, microscope, computer screen and keyboard, DNA helix and number 7

By the end of this article, you should be able to:

  • Describe, calculate and interpret important measures of risk, risk reduction or increase, and clinical utility, including odds ratio, relative risk, relative risk reduction or increase, absolute risk reduction or increase, and number needed to treat or harm;
  • Discuss the importance and impact of patient dropouts on the clinical applicability of a study’s findings and the advantages/disadvantages of the methods used to handle missing data;
  • Identify the key points that should be included in a study’s discussion and conclusions.

This article is the part of a comprehensive series exploring how to evaluate clinical studies when addressing information needs using a five-step process:

  1. Identifying study type or design;
  2. Appraising the journal, authors and study purpose;
  3. Critiquing the methods used;
  4. Understanding basic statistical tests.
  5. Analysing study data and results, and the discussion section;

This article explores the fourth step: understanding basic statistical tests. It is recommended that you read it in conjunction with:

Clinical studies often look at the likelihood of an event occurring with therapy, either an adverse effect or a beneficial outcome such as prevention of a stroke, blood clot, myocardial infarction or relapse of a condition. When comparing two therapies (i.e. treatment versus placebo and treatment versus active control), clinicians would like to know the extent to which the treatment might reduce the likelihood or risk of an adverse event so they can best apply the results to patients. There are several measures that can be used to accomplish this, which are referred to broadly as measures of risk, risk reduction and clinical utility:

The specific measures covered in this article are:

  • Odds ratio (OR);
  • Relative risk or risk ratio (RR);
  • Relative risk reduction (RRR);
  • Relative risk increase (RRI);
  • Absolute risk reduction (ARR);
  • Absolute risk increase (ARI);
  • Number needed to treat (NNT)
  • Number needed to harm (NNH).

In addition, handling of patient/subject dropouts in a study will be discussed, including how the method used for handling patient dropouts can affect the likelihood of error. Finally, considerations when reading a study’s discussion section and conclusions will be reviewed.

Odds and odds ratio

Most of us are familiar with the term ‘odds’, particularly related to gambling or games such as the lottery. The odds of an event occurring is calculated as the number of times the event occurs divided by the number of times it does not occur. For example, the odds of rolling the number 6 with 1 throw of a die is the number of times a 6 would occur on the die divided by the number of times a 6 would not occur on the die. In this case it would be the 1 time a 6 would occur divided by the 5 times a 6 would not occur = 1/5 or 1 to 5. The odds ratio (OR), which is often reported in clinical studies, is a ratio of the odds of an event occurring in one treatment group divided by the odds of the event occurring in the other group. Usually, the numerator of the OR is the odds in the treatment group — or in the cases for a case–control study — and the denominator is the odds in the control group (note: sometimes a study might reverse these, so it is important to identify which group is in the numerator and denominator since it will affect the interpretation). Thus, the OR is interpreted as follows:

  • OR < 1 (the odds of the event occurring in the treatment group is less than the odds in the control group, assuming treatment is in the numerator);
  • OR = 1 (the odds of the event occurring in the treatment group is equal to the odds in the control group; both the numerator and denominator are equal);
  • OR > 1 (the odds of the event occurring in the treatment group is greater than the odds in the control group, assuming treatment is in the numerator).

For example, suppose a study examined the efficacy of a new drug for migraine prophylaxis compared with propranolol (i.e. control). A total of 350 patients were randomly assigned to receive either the new drug (150 patients) or propranolol (200 patients). The patients were followed over 2 weeks, and the outcome measure was migraine development. A total of 30 patients (20%) in the new drug group and 50 patients (25%) in the propranolol group had a migraine over the 2-week study period. What is the OR for migraine development with the new drug compared with propranolol?

The odds of a migraine developing with the new drug therapy are less than with propranolol (OR < 1). Since the OR is actually 0.76, the odds of a migraine with the new drug are 76% of the odds with propranolol.

Risk and relative risk

Now let us talk about risk. The risk of an event occurring is calculated as the number of times the event occurs divided by the total number of persons in the involved (i.e. exposed) group. Note that the difference between calculating risk or odds is in the denominator. Using the previous die example, the risk of rolling the number 6 with one throw of a die is the number of times a 6 will occur/ total number of numbers present on die, which in this case would be 1 time a 6 would occur/total of 6 numbers on the die = 1/6 or 0.17. The relative risk (RR), which is often reported in clinical studies, is a ratio of the risk of an event occurring in one treatment group divided by the risk of the event occurring in the other group. The RR is also referred to as the risk ratio. Usually, the numerator of the RR is the risk in the treatment group, and the denominator is the risk in the control group — sometimes a study might reverse these, so it is important to identify which group is in the numerator and denominator since it will affect the interpretation. Thus, the RR is interpreted as follows:

  • RR < 1 (i.e. the risk of the event occurring in the treatment group is less than the risk in the control group, assuming treatment is in the numerator);
  • RR = 1 (i.e. the risk of the event occurring in the treatment group is equal to the risk in the control group — both the numerator and denominator are equal);
  • RR > 1 (i.e. the risk of the event occurring in the treatment group is greater than the risk in the control group, assuming treatment is in the numerator).

Let us use the same study example shown earlier: 350 patients were randomly assigned to receive either the new drug (150 patients) or propranolol (200 patients). The patients were followed over 2 weeks, and the outcome measure was migraine development. A total of 30 patients (20%) in the new drug group, and 50 patients (25%) in the propranolol group had a migraine over the 2-week study period. What is the RR for migraine development with the new drug compared to propranolol?

The risk of a migraine developing with the new drug therapy is less than with propranolol (RR is <1). Since the RR is 0.8, the risk of migraine with the new therapy is 80% of the risk with propranolol.

Key points about OR and RR

What is the difference between the OR and RR? Are they both saying the same thing? Although they are both measures for comparing the likelihood (i.e. odds or risk) of an event occurring — or not occurring — between two groups, they are not saying the same thing:

  • The OR gives the relative odds of the event occurring, and the RR is looking at the probability (i.e. risk) of an event occurring;
  • There can be large differences between both measures depending on the extent to which the event has occurred (see below);
  • For case–control study designs, only the OR makes sense. Why? Think back to the design of a case–control study (see ‘Identifying and evaluating research study types’). With a case–control study, the cases are patients who already have the event or outcome, and the controls are another group of similar patients who lack the event or outcome. To calculate risk, one needs to know the total number exposed to the therapy (i.e. treated) for the denominator. That is unknown for a case-control design;
  • Either an OR or RR can be used for designs other than the case–control, but the results need to be interpreted correctly.

For example, suppose that for every 100 patients who go to their doctor with a sudden onset of blurred vision, 60 patients had a stroke. The ‘risk’ that the blurred-vision patients had a stroke is 60/100 = 0.6 or 60%. However, the ‘odds’ that these persons had a stroke is 60/40 = 1.5, which is a fairly big difference from 0.6. The 1.5 is saying the odds of the blurred-vision patients having had a stroke is 1.5 times the odds that they did not.

Now let us make the stroke rates much smaller. Now, for every 100 patients who go to their doctor with a sudden onset of blurred vision, 20 patients had a stroke. The ‘risk’ that the blurred-vision patients had a stroke is 20/100 = 0.2 or 20%. The ‘odds’ of these persons having had a stroke is 20/80 = 0.25, which is much closer in value to the risk of 0.2.

The more frequently the event occurs (i.e. the larger the actual occurrence rate), the larger the difference will be between a calculated OR and RR. There is one study design (i.e. case–control) that should report the OR. For the other study designs, either can be appropriately used so it is important to know how to interpret results.

Worked example 1: relative risk (RR)

A study compared apixaban to warfarin control for reducing the risk of recurrent thrombosis in patients who previously had deep vein thrombosis. The following was reported for the thrombosis risk with apixaban compared to warfarin: RR = 0.85 (95% confidence interval (CI) = 0.65 to 1.05).

Is the RR finding statistically significant?

No. If a CI includes the value that indicates no treatment effect or difference, the finding will not be statistically significant. The value that indicates no treatment difference for a RR or an odds ratio (OR) is 1. When RR or OR = 1, the risk or odds in the treatment group is the same as the risk or odds in the control group. The CI shown includes 1, so the RR is not statistically significant.

In summary, when a treatment difference is reported in a study and the value of 0 is in the CI range, the reported difference is not statistically significant. When a study reports a RR or OR, and the value of 1 is in the CI range, the RR or OR is not statistically significant.

Measures of risk reduction: relative risk reduction and absolute risk reduction

Another commonly used measure of risk is the relative risk reduction (RRR) or the extent of the reduction in relative risk when the treatment event rate is less than the control rate. An analogous way to think of this is as follows: suppose you go shopping and note that the regular price of a shirt/blouse you would like to buy is $25, but the sale price is $18. You just love a good sale! What percentage reduction are you getting with this sale? In other words, relative to $25, how much smaller (i.e. proportionally and in percentage) is the $18 price?

The ‘RRR’ can be calculated in two ways: (1) 1 − RR; or (2) (control rate — treatment rate)/control rate. In the migraine example, the RRR for migraine development with the new drug compared to propranolol is 1 − 0.8 (RR) = 0.2 or 20%. For the shirt/blouse example, the second equation might be simpler to use and would be ($25 [regular price] – $18 [new price]) / $25 [regular price] = 0.28 or a 28% reduction (i.e. sale).

Since shoppers like to see the percentage sale reduction they are receiving when making a purchase, the RRR has value for clinicians because it provides an indication of the proportionate reduction in risk that patients might experience. For example, if a patient is at high risk for experiencing a serious event or complication and a drug could reduce the risk of developing this event or complication by 75%, that could be very clinically important for the patient.

Since RRR is a proportion (i.e. percentage), it can be deceiving because it can have the same value for large or small actual event rates. Consider the following two examples:

  • Ulcers occur in 5.3% of patients taking drug A and in 2.3% of patients taking drug B. What is the RRR for ulcer development with drug B compared to drug A?  RRR = (5.3% − 2.3%)/5.3% = 0.57 = 57% (RRR is usually expressed as a percentage);
  • Ulcers occur in 53% of patients taking drug A and in 23% of patients taking drug B. What is the RRR for ulcer development with drug B compared to drug A?  RRR = (53% − 23%)/53% = 0.57 = 57%.

The RRR is identical in both instances, even though the actual difference between the ulcer incidence rates is only 3% for the first example but 30% in the second example. This illustrates the problem with the RRR: it does not differentiate between small and large actual event rates.

Is there a measure that provides the actual difference in the event rates between treatments? Yes, the absolute risk reduction (ARR) gives the actual difference in rates when the treatment rate is less than the control. Simply subtract the % event rate with the treatment from the % event rate in the control group (ARR = % event rate with control – % event rate with treatment). For the migraine example in which a migraine occurred in 20% of patients taking the new drug and 25% of patients taking propranolol, the ARR = 25% − 20% = 5%. It is always good to check the ARR (easy to calculate) to determine actual treatment effects, particularly when a RRR (proportional reduction) is reported.

RRR vs ARR

Sometimes a study will provide a ‘risk reduction’ without indicating whether it is a RRR or an ARR. Is there a reason why investigators might want to do this? Yes, there can be. Suppose one or more of the investigators are employed by the drug’s manufacturer or there are other conflicts of interest present (refer to potential conflicts of interest in ‘Assessing authorship, study purpose, and journals when evaluating clinical studies’). They might want to present the findings in a manner that looks the ‘best’ to readers. For example, suppose a pharmaceutical manufacturer has a drug to treat diabetes and a study is performed to compare the efficacy of their drug to a commonly used control drug for diabetes.

The manufacturer provided the funding for the study and several of the investigators either worked for the company, owned company stock or received honoraria from the company. In the study, the manufacturer’s drug caused mild hypoglycaemia in 1.5% of patients compared to 3% with the control. The investigators reported a hypoglycaemia ‘risk reduction’ of 50% for their drug. While this appears very impressive at first glance, the 50% is an RRR, with only a 1.5% ARR. With mild hypoglycaemia, the small actual risk reduction would be of minimal clinical importance. On the other hand, suppose the event was death and there was a 1.5% death rate in the treatment group vs. a 3% death rate with control. In this case, a RRR = 50% could be clinically important even with a small ARR =1.5%.

In summary:

  • Always check to see whether a risk reduction reported in a study is a RRR or an ARR;
  • If a study emphasises the RRR, determine if one or more potential conflicts of interest are present that might have influenced the way the results were presented;
  • Consider the clinical importance of the risk reduction. A small ARR could still be clinically significant for a very serious event.

Number needed to treat

Not all patients are going to benefit from any given therapy to the same extent. In the migraine example given earlier with the new drug and propranolol, a migraine occurred in 20% of the new drug group compared to 25% of the propranolol patients during the study period. Since each of these therapies might have its own unique adverse effects, clinicians also need to weigh the benefits versus risks of one therapy against the other. Thus, it would be helpful to consider how many patients would need to be treated with the new drug, instead of propranolol, to prevent one additional migraine headache over the 2-week period. The number needed to treat (NNT) tells us that. The NNT is the number of patients that need to receive the treatment instead of the control to prevent 1 additional adverse event or bad outcome over the time period of the study. It is easily calculated as NNT = 1/ARR.

In the migraine example, the ARR with the new drug compared to propranolol is: 25% (propranolol) – 20% (new drug) = 5% = 0.05 (always make sure to use the decimal instead of the % value to calculate NNT; otherwise, you can end up with a NNT of <1 or a fraction of a patient!). The NNT = 1/0.05 = 20. This is interpreted as 20 patients need to be treated with the new drug instead of propranolol to prevent one additional migraine episode over a 2-week period.

The NNT allows a clinician to consider the possible risks patients would be exposed to from a treatment before seeing a benefit. The NNT is only expressed as a whole number and ranges in value from 1 (i.e. the best possible NNT) to a potentially very large number. The smaller the NNT, the better. For example, a NNT = 2 for a drug therapy means that only 2 patients need to be treated with the drug instead of control to prevent 1 additional adverse outcome. If NNT=200, it means that 200 patients need to be treated with the drug to prevent 1 additional adverse outcome compared to control. With the small NNT of 2, only 1 of two patients would be exposed to possible drug-related adverse effects without receiving benefit. In contrast, with the large NNT of 200, 199 patients would be exposed to possible drug-related adverse effects without receiving benefit. If these adverse effects are potentially serious or very bothersome, the benefit vs. risk of drug use should be weighed.

Key points about NNT

  • If the NNT is not a whole number, always round up any decimals to the nearest whole number. For example, suppose the ARR for ulcer development with drug A (4.3% ulcer incidence) compared to drug B (6.9% ulcer incidence) is 2.6% (6.9% − 4.3% = 2.6%), and the calculated NNT = 1/.026 = 38.46. The NNT here would be 39 patients (rounded up);
  • The smaller the difference in event rates between study groups (i.e. ARR), the larger the NNT will be. The following two study examples comparing drug A and drug B illustrate this point:
    • Study 1. ARR = drug A (10% thrombosis incidence) – drug B (3.3% thrombosis incidence) = 6.7%;
    • Study 2. ARR = drug A (22% thrombosis incidence) – drug B (7.2% thrombosis incidence) = 14.8%.
  • The NNT in Study 1 would be 1/0.067 = 14.93 = 15, and the NNT in Study 2 would be 1/.148 = 6.76 = 7;
  • The smaller the NNT, the smaller the number of patients that will need to be treated (over the duration of time in the study) to prevent one additional adverse event. Let’s take the previous example but suppose this time the ARRs are as follows: Study 1 = 6.7% and Study 2 = 0.148%. The NNT in Study 1 is 15 but the NNT in Study 2 is now 1/0.00148 = 675.68 = 676. This makes intuitive sense, in that the smaller the treatment difference, more patients would need to be treated with one drug instead of another to demonstrate a potential advantage;
  • There is no ‘cut-off’ for how small the NNT should be before it is considered important or ‘significant’. Always consider the benefits and risks of the treatments to determine whether a certain NNT justifies treatment. For example, if NNT = 500 but the outcome (that would be prevented) is serious and any adverse effects from therapy are mild, therapy might still be justified. On the other hand, if the NNT = 20 but the adverse outcome to be prevented is very minor and potential adverse effects from therapy are serious, therapy might not be justified;
  • The NNT calculated in one study can be difficult to apply or transfer to another study if the treatment duration, underlying patient characteristics, study designs used, or outcome measures vary substantially across studies.

Worked example 2: odds ratio (OR)

A study compared the incidence of severe hypoglycaemia in patients with diabetes who received either a new oral drug to treat diabetes (238 patients) or a control drug (295 patients). At the end of a year, 65 patients (27.3%) who received the new oral drug and 103 patients (34.9%) who received the control developed at least one severe episode of hypoglycaemia.

What is the OR for development of severe hypoglycaemia with the new oral drug compared to control?

The OR = odds of severe hypoglycaemia with the new oral drug / odds of severe hypoglycaemia with control = 65 (# patients with hypoglycaemia with new oral drug)/173 (# patients without hypoglycaemia with new oral drug) /103 (# patients with hypoglycaemia with control)/192 (# patients without hypoglycaemia with control) = 0.376/0.536 = 0.70. Since 0.7 is less than 1, it means that the odds of severe hypoglycaemia with the new oral drug is 0.7 times the odds with control.

What is the risk ratio (RR) for development of severe hypoglycaemia with the new oral drug compared to control?

The RR = risk of severe hypoglycaemia with the new oral drug / risk of severe hypoglycaemia with control = 65 (# patients with hypoglycaemia with new oral drug)/238 (# patients in oral drug group) / 103 (# patients with hypoglycaemia with control)/295 (# patients in control group) = 0.273/0.349 = 0.78. Since 0.78 is less than 1, it means that the risk of severe hypoglycaemia with the new oral drug is 0.78 times (or 78% of) the risk with control.

What is the relative risk reduction (RRR) for development of severe hypoglycaemia with the new oral drug compared to control?

The RRR = 1 – RR = 1 – 0.78 = 0.22 = 22% (RRR is usually expressed as a percentage). Calculating it using the other equation, RRR = 34.9% – 27.3% / 34.9% = 7.6% / 34.9% = 0.218 = 22%.

What is the number needed to treat (NNT) for the development of severe hypoglycaemia with the new oral drug compared to control? Interpret precisely what the NNT value means.

Absolute risk reduction (ARR) = 34.9% − 27.3% = 7.6%. NNT = 1/ARR = 1/0.076 = 13.16 = 14. This NNT means that 14 patients (round decimal up) need to be treated with the new oral drug instead of control for a year (the study duration) to prevent 1 case of severe hypoglycaemia.

Measures of risk increase: RRI and ARI

Suppose investigators are studying a new therapy and find that it is worse, not better, than the control. Let us take the migraine example from earlier, but this time the percentages are reversed: A total of 350 patients were randomly assigned to receive either propranolol (150 patients) or the new drug (200 patients). The patients were followed over 2 weeks, and the outcome measure was migraine development. A total of 30 patients (20%) in the propranolol group and 50 patients (25%) in the new drug group had a migraine over the 2-week study period. In this case, there is not an absolute risk reduction with the new drug compared to propranolol. Rather, the risk of migraine is increased with the new drug compared to control. In this case, we use a term called absolute risk increase (ARI) to describe the risk of migraine for the new drug compared to control. It is determined in a similar manner as ARR but with the percentages reversed: ARI = % event rate with treatment – % event rate with control. In this example, the ARI = 25% – 20% = 5%.

Number needed to harm

Instead of calculating the NNT, we now calculate an analogous number needed to harm (NNH). The NNH = 1/ARI and is defined as the number that needs to be treated with one therapy instead of the other to cause 1 additional adverse outcome over the time period studied. In our example here the NNH is 1/.05 = 20. This means that 20 patients would need to be treated with the new drug instead of propranolol to cause 1 additional migraine episode. Some have advocated avoiding NNH since its meaning can be misunderstood by readers (Higgins et al.). If used, recognise that a NNH value does not mean that ‘X’ number of patients with have the adverse outcome. It should be interpreted as stated above.

How should the NNH be rounded? Unlike the NNT that is always rounded up, the most cautious and conservative way — to avoid underestimating the harm — is to round the NNH down to the nearest whole number. For example, suppose a calculated NNH = 32.6. If it is rounded up, it means that 33 people need to be treated with the therapy to cause 1 harmful event compared to control. If it is rounded down, it means that only 32 people need to be treated with the therapy to cause 1 harmful event. Rounding it down is therefore the more cautious approach; however, many will still round the NNH up, similar to the NNT.

Opposite to the NNT, the larger the NNH, the better. For example, a NNH = 2 for a drug therapy means that only 2 patients need to be treated with the drug instead of control to cause 1 adverse outcome. If NNH=200, it means that 200 patients need to be treated with the drug to cause 1 additional adverse outcome compared to control. With the small NNH of 2, 1 of two patients would be harmed from therapy. In contrast, with the large NNH of 200, 1 would be harmed from therapy but 199 patients would not.

The Table below provides a summary of the equations to use for calculating the measures of odds, risk, risk reduction and increase, and clinical utility.

Table: Summary of risk, odds, clinical usefulness calculations

Key points about NNH

  • There is no ‘cut-off’ for how large the NNH should be before it is considered important or ‘significant’. It depends on the potential benefits vs. risks of the therapy. If the benefit from a therapy is substantial and the adverse event is minor, a smaller NNH might still be acceptable;
  • Any time more adverse events occur with a treatment compared to the control, the ARI should be used to describe the actual increase in event rates with the treatment. To keep from being confused, consider the actual study findings. It makes no logical sense to report an ARR for a therapy with more adverse events than the control. Also, the ARR should not be a negative number. If an ARR is calculated and the result is negative, you are likely looking at a risk increase with treatment and should be using the ARI;
  • When an ARI is reported, calculate an NNH instead of the NNT;
  • The larger the NNH the better. With a large NNH, many patients would need to receive the therapy to cause 1 adverse event compared to control; thus, most of the patients would not experience the adverse event from therapy.

Dropouts and data handling

‘Evaluating study methods: treatment considerations, outcomes, variables and measurements’ described the use of outcome measurements to determine treatment effects. In an ideal study, every patient enrolled would complete the study and the investigators would be able to obtain all the necessary data from every patient. Unfortunately, there usually is no such thing as an ‘ideal study’. Patients will often quit a study (i.e. drop out) for a variety of reasons, some of which might have nothing to do with the study itself. For example, patients might move out of the area or simply grow tired of showing up for all the necessary follow-up visits. In other cases, patients might develop adverse effects or might not be receiving sufficient benefit and decide to quit the study. The question arises, how should the study data be analysed when not all the patients complete the data collection points or the study?

There are two commonly used methods for handling the data from dropouts in a study:

  • Intent-to-treat (i.e. intention-to-treat);
  • Per protocol (i.e. exclusion of subjects).

Intent-to-treat

With intent-to-treat the data from all patients randomised into a treatment group are analysed in the study regardless of whether they completed the study. An advantage of using this data-handling method is that it can better represent clinical practice in that some patients in real life will quit taking their therapy or not adhere to the therapy as directed. If a study includes these types of patients, the efficacy observed would mimic actual practice. Intention to treat also minimises potential bias resulting from how and when patients might drop out of a study. In addition, since data from all patients enrolled are analysed at study completion, dropouts will not affect the statistical power with an intent-to-treat analysis. Some studies will use a modified intent-to-treat data handling method.

With this method, all patients who meet a certain criterion are analysed, such as all patients who received at least one dose of the study intervention or had data collected from at least one time point of measurement. While frequently used, this approach is not strictly considered ‘intent-to-treat’. Caution is needed with this approach since definitions of the minimum criteria to employ can vary among studies and removal of any patients might lead to bias introduction.

With an intent-to-treat analysis, how does a study include data from patients who dropped out and are no longer available for subsequent outcome measurements? There are different approaches that studies can use to include measures from patients no longer in a study.

These often include the following:

  • Last observation carried forward — the last measure the study obtained from a patient is used for subsequent measurements after they drop out;
  • Some studies that have as an outcome measure ‘success’ or ‘failure’ will assume that a patient has ‘failed’ therapy after dropping out;
  • Imputation of data — missing data points are filled in by using various procedures or equations that estimate what the values would be if the patients remained in the study.

Per protocol

With the per protocol method, only those patients who complete the study protocol as specified would be included in the data analyses. This method excludes those patients who dropped out before the end of the study for any reason or those who deviated from the protocol (e.g. skipping follow-up measurements).

A proposed advantage of the per protocol data-handling method compared to intent-to-treat is that per protocol can better determine the actual efficacy of the therapy, especially if a significant number of patients drop out owing to reasons unrelated to the therapy (e.g. lost to follow-up). For example, suppose 100 patients are enrolled in a study and receive drug A. A total of 35 patients dropped out during the study for a variety of reasons. Overall, it was found that 45 of the remaining patients had a favourable response to drug A. What would be the efficacy if the intent-to-treat data handling method were used? The efficacy would be 45/100 (data from all patients are included regardless of whether they dropped out) or 45%. What would the efficacy be if per protocol was used for the data-handling method? The efficacy would be 45/65 (data from only those patients who completed the study were included in the analysis) or 69%. Note that drug A efficacy is higher with the per protocol method because the denominator is smaller when the dropouts are removed.

When to use each

The intent-to-treat data handling method is preferred over per protocol for clinical trials. With per protocol data handling, it is possible that bias might be introduced by who drops out of the study (e.g. patients with certain characteristics might quit the study to a greater extent than patients who lack these characteristics). The benefits of random assignment might be reduced if patients drop out of the study in a non-random manner.

Studies will often use both the intention-to-treat and per protocol data-handling methods to analyse study findings when they have a considerable number of dropouts. This allows readers to determine the effect that each data handling method has on efficacy rates and allows the advantages of each method to be considered.

The per protocol data handling method will decrease sample size with dropouts since those patients will be excluded from the analyses. Consider the potential effect on statistical power when per protocol data handling is used. When there are several dropouts removed from the analyses, sample size might be reduced below the number needed for at least 80% power, which could increase the risk of Type II error for non-statistically significant findings.

Worked example 3: dropouts

A 1-year study compared Fosamax (200 patients) with Actonel (200 patients) for preventing bone fractures in patients with osteoporosis. It was determined that the power of the study to detect a 5% difference in the incidence of bone fractures, with 400 total patients and an α = 0.05, was 80%. Sixty patients (30%) dropped out of the Fosamax group and 10 patients (5%) dropped out of the Actonel group. The number of patients with new bone fractures was determined in both treatment groups.

What was the effect size used for the power analysis?

The effect size — ideally the smallest difference between treatments that would be clinically important to detect as significant — is the 5% difference in the incidence of bone fractures.

Which data handling method, intention-to-treat or per protocol, would reduce the statistical power for comparing the bone fracture rates between treatment groups?

The per protocol method would reduce the statistical power since a total of 70 patients dropped out of the study, and these patients would not be included in the bone fracture analysis.

Discussion section

Busy clinicians will often skim a published study to learn its key features and findings. If of interest, they will then read the study in more detail. One part of the publication that clinicians will usually read is the discussion section. This section should ideally summarise and analyse all the important findings of the study and compare the results to what previous studies or other relevant literature on the subject had reported.

Unfortunately, the study’s authors can sometimes use this section to focus on only their positive results while glossing over possible therapy adverse effects. They might neglect to mention important limitations of their work that readers should be aware of. The authors might also inappropriately extrapolate their findings to situations outside the objectives or scope of the study, or to a population that was not defined by the study’s eligibility criteria. The article ‘Assessing authorship, study purpose, and journals when evaluating clinical studies’ discussed the potential conflicts of interest for a study’s investigators. The discussion section of a published study is one of the areas in which bias can be introduced by the investigators.

Readers should look for the following points when analysing the discussion section of a published study:

  • Summary of the major findings and their importance;
  • Comparison of findings with those from other studies, including how the findings support or conflict with prior work, with possible explanations given for differences or inconsistencies;
  • Explanation of the meaning of the results found, considering alternative explanations;
  • Acknowledgement of any study limitations;
  • Suggestions for future research needed in the subject area;
  • A final summary or conclusion that clearly states the study’s conclusions and the potential generalisability and clinical applicability of the study findings, without over-extrapolating results beyond the scope of the study or overinterpreting or inflating the importance of the results.

Key points

  • The OR, RR, RRR, ARR and NNT are measures used to describe the risk, risk reduction and clinical usefulness of therapy;
  • The OR gives the ratio of the odds of an event occurring, and the RR gives the ratio of the risk of an event occurring. Both measures have the same numerator: the number of persons in the treatment group that have an event. They differ in the denominator: the denominator for the OR is the number of persons in the treatment group that do not have the event, while the denominator for the RR is the total number of persons in the treatment group;
  • Only the OR should be reported for a case–control study;
  • The ARR provides the actual difference (i.e. reduction) in event occurrence between treatment groups. The RRR provides the relative (i.e. proportional) reduction in event occurrence between treatment groups.
  • The NNT is calculated as 1/ARR. It should always be rounded up to the nearest whole number.
  • If a treatment has a greater incidence of an adverse event than the control group, an ARI would be reported for that treatment. An NNH can be determined from the ARI, calculated as 1/ARI. To be conservative, the NNH can be rounded down to the nearest whole number but is often still rounded up, similar to an NNT;
  • Always look at the data handling method used in a study; the intent-to-treat method is preferred to minimise the possibility of bias and better mimic actual practice;
  • If there are a significant number of dropouts in the study, the study should best analyse their findings using both the intent-to-treat and per protocol methods;
  • The discussion section of a published study should summarise its important results. Results should be discussed in light of the findings from other relevant studies in the literature, including similarities and any important or unexpected differences. The strengths and weaknesses, how results could be applied to clinical practice and suggestions for future research should also be provided.

How to apply to practice

  • When a study reports an OR or RR, always look to see whether the treatment or control data are in the numerator or denominator. When the treatment group is in the numerator and the control group is in the denominator (i.e. the usual situation), the odds or risk of the event occurring is less in the treatment group when the value is <1 and greater in the treatment group when the value is >1. If the control group is in the numerator and the treatment group is in the denominator, the odds or risk of the event occurring is less in the control group when the value is <1 and greater in the control group when the value is >1;
  • If a study reports only a ‘risk reduction’, determine whether it was a RRR or an ARR. If a RRR is reported, always look at the ARR to see if the actual treatment difference is of possible clinical importance;
  • Look for, or calculate, the NNT (= 1/ARR) to determine the possible risks patients would be exposed to from a treatment before seeing the benefit. The larger the NNT, the larger the number of patients that would be exposed to possible adverse effects without necessarily receiving a treatment benefit;
  • If the intent-to-treat data handling method is used to analyse the results in a study with a fairly large number of patient dropouts, the efficacy of the drug might appear lower than it actually is;
  • If the per protocol data handling method is used to analyse data with a fairly large number of patient dropouts, fewer patients will be included in the analyses so the power could be adversely affected. Make sure the power would still be acceptable.
  • Carefully read the discussion section of a published study to determine if biased statements were made, adverse effects were minimised or inappropriate extrapolations were stated. This might be more likely when there are potential conflicts of interest for the study’s investigators.

Acknowledgements

This article was adapted from Drug Information and Literature Evaluation, Second Edition, previously published by Pharmaceutical Press.

A full list of resources and materials used to prepare the book can be accessed from the bibliography page.

Last updated
Citation
The Pharmaceutical Journal, PJ September 2026, Vol 317, No 8013;317(8013)::DOI:10.1211/PJ.2026.1.428154

    Please leave a comment 

    You might also be interested in…