
Charlotte Gurr
After reading this article, you should be able to:
- Describe adherence and how it can influence a study’s findings;
- Explain the importance of outcome measure selection in a study and the concepts of validity, reliability, sensitivity and specificity related to the outcome measures used;
- Describe the dependent and independent variables in a study and possible confounding variables;
- Describe the scales (i.e. levels) of measurement of study data and the importance of differentiating among them.
This article is the part of a comprehensive series exploring how to evaluate clinical studies when addressing information needs using a five-step process:
- Identifying study type or design;
- Appraising the journal, authors and study purpose;
- Critiquing the methods used;
- Analysing study data and results, and the discussion section;
- Understanding basic statistical tests.
This article explores the fourth step: critiquing the methods used. It is recommended that you read it in conjunction with the article ‘Evaluating clinical trial eligibility criteria: recruitment methods, informed consent, control and blinding‘.
To access other articles in the collection visit the hub page on The Pharmaceutical Journal.
This article reviews some important considerations when examining the methods used in a study, with the focus on drug treatments, adherence, outcomes, variables, and measurements. It is highly unlikely there will ever be the ‘perfect’ study. The key is identifying a study’s weaknesses and limitations and then assessing whether they are important enough to invalidate the findings, or whether they simply limit how and to whom the results could be applied.
Treatment considerations
There are several aspects to consider when reviewing the drug(s) being used in a clinical study, including:
- Dosage and dosage forms;
- Dosing frequency;
- Route of administration;
- Duration of therapy;
- Drug concentrations obtained;
- Use of any concurrent (i.e. concomitant) non-study medications;
- Adverse effects;
- Therapy adherence (i.e. compliance).
The dosages of the experimental drug and any active control in a study should be appropriate for the drugs used. If there are established dosing ranges, the drugs should best be dosed at comparable ends of the ranges (e.g. both dosed at the lower end or both dosed at the higher end). In general, established drugs should be dosed similarly to how they would be dosed in practice: in a single, fixed dose or individually adjusted doses, as most appropriate.
The dosing frequency should be consistent with the known pharmacokinetics of the study drugs. If a drug has an established therapeutic concentration range, the study should ensure that concentrations are measured in patients and are within that range. Concentrations should be measured at appropriate times of the day and at steady state for chronic therapy. The duration of therapy should be long enough to allow for the full therapeutic effect to occur and for adequate time to assess the benefit on the medical condition.
Box 1: Non-study medication use
If concurrent non-study medication might affect the study outcomes, determine whether:
- Comparable quantities were taken among the patients in each study group;
- The overall amounts taken were large enough to impact the findings.
For example, suppose a study comparing a new antihistamine to placebo does not exclude decongestant use. Since decongestants could affect certain allergy symptoms, their use could alter some of the study findings. In this study, the investigators should record the amount of decongestants taken in both the antihistamine and placebo groups. The reader should ask:
- Were decongestants used to a large enough extent in both groups to cause significant reductions in nasal congestion symptoms?
- If decongestants were taken by more patients in one of the groups, could they have made the antihistamine appear more (or less) efficacious?
A study may or may not allow patients to take non-study medications concurrently. When outside medications are known to interact with the study drugs or will otherwise affect the condition being studied, they will usually be excluded as part of the exclusion criteria. If concurrent, non-study medications are allowed during a trial, consider whether they could interact with the study drugs or affect the disease state or symptoms studied.
Adverse effects should always be considered when determining the benefits versus risks of a therapy. Mild as well as serious adverse effects can be of concern and affect the clinical usefulness of a drug. For example, a drug can be very efficacious but cause many minor, but annoying, side or adverse effects. Another drug might be very efficacious and cause only a few adverse effects, but the effects are serious. Both situations could result in patients dropping out of a study or being non-adherent (non-compliant) with their therapy. As previously mentioned, adverse effects are also a source of unblinding in a study.
Patient adherence can be an issue for both the procedures used in a study (e.g. completing a diary of symptoms or adverse effects, filling out diet/meal cards) as well as the drug therapy. A study will not be able to obtain complete, accurate data if patients do not complete needed diaries or records. Likewise, adherence with a therapeutic regimen is very important in clinical studies and in practice. A therapy cannot be expected to work if the patient does not take it. It is important that a study always considers adherence for any needed patient-supplied diaries/records and the drug therapy and assesses the extent to which it occurred. Otherwise, the findings might reflect varying adherence rates and not actual efficacy differences.
Compliance bias occurs when differences in adherence to a study’s treatment or intervention affects the outcomes. For example, suppose two treatments are being compared in a randomised study, and it is found that patients in one group were twice as likely to take their assigned treatment as directed. If treatment efficacy is significantly greater in the highly adherent group, it is possible that compliance bias, and not an actual efficacy difference between treatments, was responsible. When compliance bias is present, it can really confound a study’s findings since non-adherence to a study intervention could be a ‘marker’ for other unhealthy behaviours.
A study should measure and compare the adherence rates to the therapy (and other necessary study procedures) in each patient group to determine if compliance bias might have occurred. Adherence rates should be similar to rule out possible compliance bias.
Box 2: Non-adherence as a confounding factor in studies
A study compared the mortality risk in patients who were asked to undergo a variety of cancer screening tests (i.e. screening group) with control group patients who received usual care. It was found that non-adherent patients in the screening group who underwent none of the recommended screening tests had greater mortality from non-cancer-related causes than the patients who completed all the recommended screens. It was unlikely that cancer screening tests protected patients from non-cancer-related deaths.
What could be an explanation for the association observed?
It is possible that non-adherence might reflect other behaviours that could increase mortality. In general, adherent study patients might be more likely to engage in regular exercise, eat healthier diets, or receive vaccinations than non-adherent patients. Owing to the potential impact of compliance/adherence on a study’s findings, studies should consider how any differences found in treatment adherence might influence their results1.
It is easy to measure adherence when patients are closely monitored (e.g. inpatient studies, IV medications administered by nurses) or when patient participation in the intervention can be easily determined (e.g. attendance at counselling sessions or scheduled office/clinic visits, undergoing surgical procedures). For adherence with study procedures, diaries or other records can be reviewed and missing entries quantitated. Measuring adherence to drug therapy can be more difficult and less reliable, particularly in outpatient settings in which the patients are responsible for taking all their drug doses. Since there is no perfect method for assessing drug therapy adherence, it is best for a study to use more than one approach.
Methods for determining adherence
The methods more commonly used to determine adherence to drug therapy include:
- Pill counts — quick and easy but can be unreliable (i.e. inflated) since doses can be lost or non-taken doses thrown out;
- Review of pharmacy refill records — easy but obtaining a prescription does not indicate whether patients actually took the drug as they were supposed to;
- Use of electronic caps or devices that record the times the drug’s container or package were opened — can be a good method when available, but its use could also be circumvented;
- Asking patients directly or have patients keep a diary/record of doses taken — often used but can be inaccurate;
- Measuring drug concentrations or adding an inert ‘marker’ to a drug product that can be measured — can indicate if a drug was taken, but not adherence with all doses or with the dosing schedule since missed doses could be taken shortly before a scheduled visit;
- Measuring a drug’s physiologic effect on a non-primary study outcome (e.g. certain lab changes, antibacterial activity present in urine) — interpreting this information can be difficult since patients might respond differently to these variables and it does not guarantee that patients took the drug regimen the way it was prescribed.
Worked example 1
Two pain medications, Drug A and Drug Y, were compared for the treatment of patients with severe back pain. After three months of therapy, 89% of patients receiving Drug A reported ‘complete’ pain relief compared to only 60% of Drug Y patients. Bitter taste was reported by 80% of patients receiving Drug Y; other side effects included mild nausea and headache in 4–5% of patients taking either drug. Using pill counts by investigators and diaries in which the patients were asked to record their doses taken, it was found that 96% of patients receiving Drug A took most (at least 90%) of their doses compared to 65% of Drug Y patients..
Was adherence appropriately measured in this study? Was compliance bias a potential problem?
There are several methods that can be used to determine the degree of patient adherence in a study; none of them is 100% accurate so a combination of methods is best used. This study used two acceptable methods (e.g. pill counts, patient diaries). Compliance bias occurs when different rates of patient adherence in the treatment groups lead to differing efficacy rates. The much lower degree of drug adherence by the Drug Y patients could explain the finding of less efficacy of Drug Y compared to Drug A.
Outcomes and variables
A study’s objective should specify the overall outcome(s) of interest (e.g. hypertension control, diabetes control, smoking cessation). In their methods, investigators should clearly specify their primary outcome(s) of interest and any secondary outcomes (i.e. of interest but not the main concern or focus of the study). Studies should also indicate a desired end point where appropriate, meaning the point at which an outcome will be successfully met or the study hypothesis supported. For example, in a study comparing the efficacy of two antihypertensive drugs, the primary outcome might be diastolic blood pressure changes and an appropriate end point might be a diastolic blood pressure of <80 mmHg. Secondary outcomes might include systolic blood pressure, renal function, serum cholesterol, or other metabolic changes that could result from the therapy.
Variables are related to the outcomes and refer to a study characteristic that can assume different values. Two important study variables are the dependent and independent variables. The independent (i.e. explanatory) variable(s) in a study affect the value of the dependent (i.e. response) variables. In a clinical study, type of treatment is generally one independent variable. The dependent variables in a study are those that change in value because of the independent variable; in a clinical study, the outcome measures are the dependent variables since their values would be altered by exposure to the therapy (i.e. control or treatment) administered.
A third type of variable, the confounding variable, can also play an important role in studies. It is a factor that, by itself, can affect or interfere with the study’s outcome measures in addition to the exposure or therapy being studied. For example, suppose a study compared a new antiviral drug to placebo to prevent influenza. Several study patients had recently been vaccinated with the influenza vaccine. The vaccine would be a confounding variable that could affect the determination of drug efficacy. As another example, a study could report that coffee drinkers are more likely to develop lung cancer than non-coffee drinkers but ignore the contribution of smoking, a confounding variable that might be more common in coffee drinkers. Investigators, and readers, need to consider all potentially confounding variables in a study. These variables need to be accounted for during analyses because they can interfere with the study results and their interpretation.
Measurements
The tests or procedures used to measure changes in the desired outcomes (i.e. dependent variables) should be appropriate for the intended purpose. Ideally, they should measure only the effect of interest without allowing interference from unrelated factors.
Studies should select the best test or combination of tests to measure an outcome, such as using a ‘gold standard’ test when appropriate and feasible for the situation and study setting.
For a study’s measurements to be meaningful and truly indicate the extent to which the outcomes have been achieved, the tests or procedures used should be: (1) valid; (2) reliable; (3) sensitive; and (4) specific.
Important characteristics for measurements include:
- Validity — the extent to which a measure is truly determining what is desired or what should be measured;
- Reliability (e.g. reproducibility, consistency) — the extent to which a measure of the same outcome provides similar results when used at different times or by different individuals.
- Sensitivity — the degree to which a measure can accurately identify as positive those who have a condition or characteristic (i.e. no false negatives); or, the ability to which a measure can identify or detect the presence of a condition or characteristic when present.
- Specificity — the degree to which a measure can accurately identify as negative those who do not have a condition or characteristic (i.e. no false positives); or the extent to which a measure can accurately detect only the condition or characteristic of interest.
Worked example 2
Suppose the Arthritis Quality of Life Scale can measure accurately only the changes in quality of life that result from arthritis but not from any other related medical conditions. It cannot detect small changes in quality of life, however; it only detects fairly large changes.
Based on the information provided here, which of the following characteristics, validity, reliability, sensitivity, or specificity, does this survey instrument appear to possess?
Specificity, since it measures quality of life changes only due to arthritis and not other related conditions. Meaning, false positive results (i.e. the test is positive in a patient without arthritis) are unlikely. The Arthritis Scale also seems to be valid if it accurately measures the quality of life. However, one cannot determine the reliability of this instrument given the information provided. For scales, surveys or questionnaires that are developed by investigators, it is very important that their validity, which includes reliability, be determined (i.e. the scale, survey, or questionnaire should be validated) prior to use in a study. Sensitivity seems to be lacking in that the scale misses small changes in quality of life. Meaning, there might be false negative results (i.e. the test is negative in a patient whose quality of life improved).
Worked example 3
A new, inexpensive home nasal swab test for COVID-19 is marketed. Compared to a currently used PCR test for COVID-19, the nasal swab test has a greater likelihood for false positive results. However, both tests are 98% accurate at identifying asymptomatic or mild COVID-19 when present. Select the correct words to fill in the blanks based on the information provided.
Compared to the PCR test, the nasal swab test has less ______, but equal _____.
- Specificity, sensitivity
- Reliability, specificity
- Sensitivity, reliability
- Sensitivity, specificity
- Specificity, reliability
- Reliability, sensitivity.
The correct answer is a.
The nasal swab test is more likely to cause false positive results, which means the test can produce a positive result for COVID-19 even when the person does not have it. Therefore, specificity, the extent to which a person tests negative when they do not have a condition, is low. Since both types of tests have high positive rates when a person has the condition, the sensitivity is equally high for both (i.e. low number of false negatives — it is very unlikely that a person with COVID-19 will be missed with this test).
There are several types of validity including internal, external, face (i.e. the extent to which a test appears to be measuring what it is supposed to measure), content (i.e. the extent to which a test measures all aspects of the desired condition or behaviour), construct (i.e. the extent to which a test actually measures the knowledge/skills/behaviours that should be measured), and criterion (i.e. the extent to which a measure agrees with an established standard for what is being measured) validity. Two of the major ones to consider when evaluating a study are internal and external validity. Internal validity is the extent to which a study’s findings were appropriate and correct – was the relationship found between the intervention and outcomes accurate? If a study has internal validity, the extent to which its findings can be applied or generalised to patients and settings outside the study is referred to as external validity. For example, if a study enrolled patients who did not adequately represent the population of interest (e.g. had milder or more severe disease, were older or younger, etc.) or used a treatment dosage not normally used in practice, the external validity will be limited.
Box 3: Key points when assessing study methods
- The stronger a study’s design, methods and analyses, the greater the internal validity. Randomised, controlled experimental studies have stronger internal validity than other types of trials. However, any important weaknesses or flaws will reduce internal validity, even with a ‘gold standard’ study design;
- External validity (i.e. generalisability) is an important consideration for clinicians who wish to apply a study’s results to their patients;
- A variety of factors can affect external validity, including the types of patients enrolled, how they were selected, the study setting (e.g. inpatient versus outpatient), and even the timing of the study (e.g. studying antihistamines at a time of the year when certain allergens might be less prevalent);
- A measure can be reliable without being valid, but to be valid it must also be reliable;
- A study should provide any needed training and use clear, standardised directions, protocols, and definitions to ensure that an outcome measure is appropriately used across study locations and by different investigators and patients;
- Patients should also be handled similarly, with the exception of the treatments used, to help ensure that the study methods or measures themselves did not lead to altered patient behaviours that might impact the findings. The Hawthorne effect is when patients alter their performance, behaviours, or attitudes simply due to being observed or given attention in a study, not from the intervention given. How can the Hawthorne effect be minimised in a study? It can be reduced by handling patients in each study group as similarly as possible, with the exception of the treatment/intervention given;
- For example, if one study intervention required patients to be seen by investigators each week in the clinic for follow-up visits while the other intervention only required patients to be seen once monthly, the patients being seen more frequently might alter their behaviours because they are being observed more often. Suppose these patients were in a study comparing two antihypertensive drugs. Patients being observed more closely might decide to reduce their salt intake and exercise to a greater extent, simply because of being observed more closely. This behaviour change could make their drug appear more efficacious, when reduced sodium and greater exercise alone were responsible for the difference.
Measurements produce numerical data that can have different meanings. For example, the number 5 could indicate the number of patients with a certain characteristic, a value in a rating scale that means ‘extremely positive’, or a serum potassium concentration. Levels or scales of measurement differentiate these meanings. The level/scale of measurement of specific data is very important because it is one of the factors that determine the type of statistical test to use for the data analysis. The three levels of measurement are nominal, ordinal, and continuous, with continuous data further divided into interval and ratio levels. Since the same types of statistical tests are generally used for interval and ratio-level data found in clinical studies, the term ‘continuous data’ will be used hereafter to include both.
Definitions of levels/scales of measurement
Nominal (categorical): Data that lack numerical qualities but can be placed in mutually exclusive categories. Nominal data that can assume only one of two values is also referred to as dichotomous. Some examples of dichotomous data include cure/no cure, live/die, and disease present/absent. Other examples of nominal data include presence or absence of adverse effects or outcomes, such as the proportion/ percentage of patients with a headache, those who were cured, or those whose hypertension was controlled.
Ordinal: Data that can be rank-ordered on a scale; one value is more or less than another, but any assigned numbers do not have exact differences between them (e.g. the difference between a 1 and a 2 on an ordinal scale is not exactly the same as the difference between a 3 and a 4); examples include:
- Opinions ranked using a scale of 5 = strongly agree, 4 = agree, 3 = neutral, 2 = disagree, 1 = strongly disagree;
- Ranking of quality of life as 4 = excellent, 3 = good, 2 = fair, 1 = poor;
- Rating of the severity of a condition as severe, moderate, mild, absent.
Continuous: Data that can assume an unlimited number of numerical values within a range (e.g. blood pressures, FEV1 measures, electrolyte and other concentrations, etc.), with equal distances between numbers; the types include:
- Interval: Continuous data that lack a true zero point; cannot be presented as a ratio; examples include Fahrenheit or Celsius temperatures and pH values;
- Ratio: Continuous data that have a true zero point; it is the most common type of continuous data used for clinical research; examples include height, weight, and drug concentrations.
Note that continuous or ordinal data can be converted to nominal level in a study for data analyses. Thus, closely consider not only the level of the data gathered but how the data are classified and used for data analyses. Examples include:
- A study comparing antihypertensive drugs, in addition to analysing the actual differences in systolic or diastolic blood pressure (BP) readings (continuous-level data), might also compare the proportions of patients with a normal BP (defined as a BP < 120/80 mmHg) following therapy. The proportions with normal BP (i.e. either normal or elevated) are categorical and nominal level;
- A study measures patients’ ratings of quality of life (QOL) using a 4-point survey from 4 = excellent to 1 = poor (ordinal level data). Investigators also compare the overall QOL change from baseline to end of therapy defined as either Beneficial (change of at least 1 point) or Not beneficial (no change or worsening). The overall QOL change data are nominal level.
Worked example 4
List the level/scale of measurement for each of the following types of study data:
- Thyroxine serum concentrations following thyroid hormone replacement;
- Severity of neuropathic pain (3 = severe, 2 = moderate, 1 = mild, 0 = absent) following therapy with gabapentin or placebo;
- Number of osteoporosis patients who experienced a fracture during treatment with either alendronate (3 of 29; 10.3%) or risidronate (5 of 35; 14.3%);
- Haemoglobin A1c concentrations at baseline and following metformin therapy;
- Patients with diabetes who experience an episode of hypoglycemia (blood glucose concentration < 65 mg%) following therapy with pioglitazone or dapagliflozin.
Answers
- Continuous — there are an infinite number of possible values, with equal distances between numbers;
- Ordinal — the severity can be rank-ordered, but the difference between a pain score of 1 and 2 is subjective and not the same as the difference between 2 and 3;
- Nominal — the patients can be placed into mutually exclusive categories: experienced a fracture or did not experience a fracture;
- Continuous — there are an infinite number of possible concentration values, with equal distances between numbers;
- Nominal — although blood glucose concentrations themselves are continuous level, when the numbers or proportions of patients who experience hypoglycaemia are compared, this comparison becomes nominal level (patients either had hypoglycaemia or they did not; they can be placed in mutually exclusive categories).
As with other parts of a study, errors or biases can occur with measurements. These errors (i.e. biases) have been categorised as random error or systematic error. Random error, as the name suggests, occurs in a non-predictable or non-reproducible manner and are generally unavoidable. They owe to normal fluctuations around the true value (both directions, slightly higher or lower) in measurement devices or in the measurement readings by investigators. The best way to minimise any effect from random error is to ensure that there is an adequate sample size (i.e. increase the number of study subjects) so there will be sufficient measurements to ‘cluster’ around the actual value when an average is taken. In contrast, systematic error is a consistent, repeatable and reproducible error in which measures are skewed in the same direction from the true value.
Systematic error generally results from problems with the measurement instruments or tools used, or it can result when poor study design or methods are used. The best way to minimise systematic error is to ensure measurement instruments and techniques are accurate and a strong study design and methods are used (e.g. double blinding, randomisation, etc.).
Summary and key points
- All drugs being studied should be dosed and administered appropriately;
- If concurrent non-study medications are allowed, the use of any that might affect the study treatments or results should be quantitated and compared between groups;
- Adherence should always be considered in a study and reported. More than one method for determining therapy adherence should best be used;
- Primary and secondary outcomes should be clearly defined and appropriate for the study objectives;
- Study measurements should be valid, reliable, specific, and sensitive;
- The level/scale of measurement of data is an important factor that determines the type of statistical test to use.
How to apply to practice
Be sure to check:
- The dosing and administration regimens used to ensure they are appropriate and representative of what would be used in clinical practice;
- Whether concurrent non-study medications could have significantly affected the findings;
- Whether adverse effects led to significant patient drop-out or nonadherence, were frequent or serious enough to affect the clinical usefulness of the drug, or might have led to unblinding;
- The adherence rates for each study group. If significant adherence differences are present among study groups, ask whether compliance bias occurred and if these differences might have contributed to the results seen (e.g. did the group with much lower therapy adherence also have a lower drug response?);
- Whether the validity and reliability were determined for any surveys or questionnaires that the investigators developed themselves;
- If the outcome measures used were appropriate and optimal given the study objectives;
- Whether any limitations or weaknesses in a study or its design could reduce the internal validity, thereby affecting the external validity;
- The scale or level of measurement of the data obtained and analysed. Watch for differences in how the patients were handled that might affect the study results independently of the treatments used (i.e. Hawthorne effect).
The methods section is one of the most important parts of a published study. A study must be appropriately designed to fulfil its hypothesis and objectives, and the methods section specifies precisely how the study was conducted. Serious weaknesses in a study can invalidate its findings. Several important study design considerations are covered in this and the previous article. Readers need to be familiar with these considerations, including the use of controls, randomisation and blinding, treatments used, and measurements — among others — in order to analyse and draw valid conclusions from clinical studies.
Self-assessment questions
Question 1
Select the type of validity that refers to the ability to extrapolate a study’s results to other patients who were not part of the study.
A: Internal validity
B: External validity
C: Face validity
D: Content validity
E: Criterion validity
Question 2
A randomised single-blind study compares a new triptan (T) drug to sumatriptan (S) to treat migraine headaches. A total of 54 patients are assigned to receive either T or S. When a migraine is experienced, patients take a dose of their assigned drug and record the severity of pain over the next 24 hours. The outcome measures include the time to headache relief and the average pain severity over the 24-hour period. Patients are allowed to take non-steroidal anti-inflammatory drugs (NSAIDs) as needed if they do not experience headache relief within two hours after their T or S dose.
What do the terms ‘single-blind’ and ‘randomised’ indicate in this study? Would the use of NSAIDS be a potential problem in this study?
Question 3
What measurement term refers to the ability of a study to accurately find true and correct results within the study based on having a strong design and methods used?
Question 4
A study of methotrexate for treating rheumatoid arthritis measured the degree of joint erosion and joint space narrowing on X-rays as two of the primary outcomes. Assume that the x-rays can measure even very small changes in the degree of joint erosion and joint space narrowing. Which measurement attribute describes this ability to detect very small changes in the outcome measure?
A: Reliability
B: Sensitivity
C: Specificity
D: Validity
Question 5
A double-blind, randomised study examined the use of a new medication to treat patients suffering from psoriasis. A total of 28 study sites were used at various locations in the United States and Canada. One of the outcome measures included the rating of psoriasis severity by investigators using the following rating scale that was developed by the investigators: 4 = severe; 3 = moderately severe; 2 = moderate; 1 = mild; 0 = absent. The study found that, following use of the medication, there was a slight but significant reduction in the rating of psoriasis severity. What are the potential problems with the rating scale data from this study?
Answer guidance
Question 1
B: External validity.
The extent to which its findings can be applied or generalised to patients and settings outside the study is referred to as external validity.
Question 2
‘Single-blind’ means that the patients (most likely) are not aware of whether they are taking T or S but the investigators know the drug each patient was assigned. ‘Double-blind’ is preferred to reduce the risk of bias.
‘Randomised’ means that each patient had an equal chance of being assigned to receive either T or S therapy. Randomisation is very important to minimise the likelihood that the patients in the study groups are different with regard to factors (both known and unknown) that might influence the results. Studies should state the method used for randomisation so the reader can determine if it was appropriately performed. Since randomisation does not guarantee that the patients in the study groups will be comparable to each other, the investigators should still compare post-randomisation baseline characteristics between groups to see if any significant differences are present owing to chance.
The use of concurrent NSAIDS could be a problem since these drugs can help relieve headache pain. If several doses are taken, the NSAIDS could reduce migraine pain severity and interfere with the patient’s assessment of pain relief from T or S therapy. It is important in this study that the investigators quantitate and compare the use of NSAIDS in both treatment groups. If one group takes many more NSAIDS than the other and reports greater pain relief, it will be difficult to separate the NSAID effect from the drug treatment effect.
Question 3
Internal validity.
Internal validity needs to be present before a study can have external validity (i.e. the ability to generalise its findings outside of the study).
Question 4
B: Sensitivity.
A test that is sensitive means that it can detect the presence of a condition or characteristic when present. In this example, if joint erosion or joint space narrowing occurs, even to a small degree, the X-ray is able to detect it.
Question 5
The rating scale could lack validity and reliability. Several psoriasis rating scales have been used and validated in practice, including the standardised Psoriasis Area and Severity Index (PASI). Refer to the study by Robinson et al. (2012) for more information about psoriasis assessment. For any rating scales used in a study, clear definitions for each rating level should be provided to all study investigators across the investigative sites to help ensure consistency of the results. For example, the extent to which redness, thickness and scaling were present, as well as the percentage of skin area involved, are important considerations for evaluating psoriasis and should have been specified for each rating scale level.
- 1.Pierre-Victor D, Pinsky PF. Association of Nonadherence to Cancer Screening Examinations With Mortality From Unrelated Causes. JAMA Intern Med. 2019;179(2):196. doi:10.1001/jamainternmed.2018.5982


