# INBDE practice questions: FK10 — Research and biostatistics

Ten original INBDE practice questions on FK10 — Research and biostatistics, each answered on this page with a rationale and a source.

Last updated: 2026-08-10.

## Question 1

A trial in 24 patients compares two restorative materials and reports no significant difference in two-year failure, p = 0.31, with a confidence interval spanning a large advantage for either material. A colleague concludes the two are equivalent and interchangeable. What is the correct reading?

- A. The trial could not tell them apart; equivalence needs a pre-specified margin
- B. The materials are equivalent, since the p-value sits far above the 0.05 threshold
- C. The wide interval shows a real difference that the p-value failed to detect
- D. The trial proves neither material is better, which is what equivalence means

**Answer A:** The trial could not tell them apart; equivalence needs a pre-specified margin

Power is the chance of detecting a real effect of a given size, and studies are designed for 80% or 90% power; a small trial reporting "no significant difference" has not shown the treatments are the same, only that it could not tell. The confidence interval here proves the point: it spans a large advantage in both directions, which is the signature of an imprecise study. Showing that a treatment is not meaningfully worse requires an equivalence or non-inferiority design with a margin set in advance. B rests the whole conclusion on a threshold, which the ASA statement warns against. C reads imprecision as a hidden finding, when an interval crossing no effect supports no such claim. D restates the same fallacy in different words.

**Common trap:** Treating "not statistically significant" as proof of no difference in a study too small to find one.

Source: [ASA Statement on Statistical Significance and P-Values](https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf)

## Question 2

A 44-year-old man asks whether a chlorhexidine rinse will help his gingivitis. The trial his dentist finds reports a relative risk of 0.86 for the outcome he cares about, with a 95% confidence interval of 0.68 to 1.09 and p = 0.21. How should she read that interval?

- A. Significant, because the estimate of 0.86 sits below the no-effect value of 1
- B. Not significant, because a ratio's confidence interval must exclude 0 to count
- C. Not significant: the interval crosses 1, the no-effect value for a ratio
- D. Uninterpretable without the p-value, which alone decides whether a result is significant

**Answer C:** Not significant: the interval crosses 1, the no-effect value for a ratio

A confidence interval carries direction, size, and precision at once. Which value means "no effect" follows from the measure: 0 for a difference, 1 for a ratio. This interval runs from 0.68 to 1.09 and therefore includes 1, so the result is not significant at the 5% level, and its width says the study was imprecise about how large any benefit might be. A stops at the point estimate and ignores the range around it, which is the whole information the interval adds. B applies the right rule to the wrong measure — 0 is the no-effect value for a difference such as a risk difference, not for a risk ratio. D inverts the hierarchy: current oral-health reporting guidance asks for estimates in clinically meaningful units with confidence intervals rather than reliance on p-values.

**Common trap:** Memorising "the interval must not cross zero" and applying it to a ratio.

Source: [OHStat Guidelines](https://www.iadr.org/about/news-reports/press-releases/ohstat-guidelines-reporting-observational-studies-and-clinical)

## Question 3

A two-year trial in high-caries-risk adults reports that 25% of the control group developed a new root lesion, compared with 15% of those enrolled in a varnish and recall programme. A 61-year-old patient asks how many people like him must join the programme to prevent one lesion.

- A. Four patients, computed as one divided by the control event rate of 25%
- B. Three patients, computed from the 40% relative risk reduction the programme achieved
- C. Seven patients, computed from the average of the control and treated event rates
- D. Ten patients, computed from the absolute risk reduction of ten percentage points

**Answer D:** Ten patients, computed from the absolute risk reduction of ten percentage points

The absolute risk reduction is the control risk minus the treated risk: 25% − 15% = 10 percentage points. The number needed to treat is 1 ÷ ARR = 1 ÷ 0.10 = 10, always rounded up because a fraction of a patient cannot be treated. A divides by the control risk alone, which answers no question and ignores the treated arm entirely. B divides by the relative risk reduction — here RR = 15 ÷ 25 = 0.60, so RRR = 40% — and the relative measure hides the baseline that NNT exists to expose; the same 40% reduction on a much lower baseline would yield an enormously larger NNT. C invents an arithmetic with no basis in the definitions. Relative numbers make treatments sound impressive; absolute numbers tell this patient what to expect.

**Common trap:** Building the NNT from the relative risk reduction instead of the absolute one.

Source: [Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (2024)](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-14)

## Question 4

A manufacturer's leaflet in the reception area states that a home-use product "reduces caries by 60%." The mother of a 9-year-old with no lesions and a two-year baseline risk of roughly 0.2% asks what that figure means for her son. What should the dentist tell her?

- A. His risk falls by 60 percentage points, which is a large and worthwhile reduction
- B. Absolutely by about 0.12 percentage points; the number needed to treat is 834
- C. The relative figure applies equally at any baseline, so the benefit is identical
- D. Nothing can be said until an odds ratio for his exact age group is published

**Answer B:** Absolutely by about 0.12 percentage points; the number needed to treat is 834

A 60% relative risk reduction sounds identical whether the baseline risk is 20% or 0.2%, which is exactly why relative numbers are the advertising numbers. On a 20% baseline the absolute risk reduction is 12 percentage points and the number needed to treat is 9; on a 0.2% baseline the same relative reduction gives 0.12 percentage points and an NNT of 834. A confuses a relative reduction with an absolute one and inflates a tiny benefit into an impossible one. C states the specific error the calculation refutes: the relative figure travels, but what it is worth does not. D withholds an answer that the baseline risk and the relative reduction already supply, and asks for the one measure most likely to be misread when outcomes are common.

**Common trap:** Quoting a relative reduction to a low-risk patient without converting it to what he can actually expect.

Source: [Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (2024)](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-14)

## Question 5

In a two-year sealant trial, 200 of 1,000 control children and 80 of 1,000 sealed children develop a new occlusal lesion. The paper reports an odds ratio of 0.35, and a study-group member concludes that sealing cuts a child's risk by 65%. How should that be read?

- A. The odds ratio overstates it; the relative risk here is 0.40, a 60% reduction
- B. The odds ratio understates it, because odds always sit closer to 1 than risks
- C. Both are correct, since an odds ratio equals a relative risk in randomized trials
- D. Neither applies, because a two-arm trial can only report absolute risk differences

**Answer A:** The odds ratio overstates it; the relative risk here is 0.40, a 60% reduction

From the counts, control risk is 200 ÷ 1,000 = 20% and sealant risk is 80 ÷ 1,000 = 8%, so the relative risk is 8 ÷ 20 = 0.40, a 60% relative reduction. The odds are 200 ÷ 800 = 0.25 and 80 ÷ 920 = 0.087, giving an odds ratio of 0.35. The odds ratio always sits further from 1 than the risk ratio, so reading 0.35 as a 65% risk reduction overstates the benefit. B reverses that relationship. C asserts an equivalence that holds only when the outcome is rare, and how rare is rare enough is a judgement rather than a fixed threshold — a 20% outcome is plainly not rare. D is false: a trial follows people at risk with known denominators, so risk, relative risk, absolute risk reduction, and NNT are all available.

**Common trap:** Reading an odds ratio as if it were a relative risk when the outcome is common.

Source: [Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (2024)](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-14)

## Question 6

A 34-year-old woman has a shadowed occlusal fissure on a bitewing. Her dentist wants an additional test whose negative result would let her leave the site unrestored and monitor it with confidence at recall. Which property of the candidate test should she weigh first?

- A. High specificity, since a specific test's negative result rules the disease out
- B. High accuracy, since it summarises both kinds of error in a single figure
- C. High sensitivity, since a sensitive test's negative result rules the disease out
- D. High positive predictive value, since it describes what a result means here

**Answer C:** High sensitivity, since a sensitive test's negative result rules the disease out

Sensitivity is TP ÷ (TP + FN), so a highly sensitive test produces few false negatives and a negative result can be trusted to rule disease out — the SnNout rule, which is precisely what she needs before deciding not to intervene. A applies the companion mnemonic backwards: specificity is TN ÷ (TN + FP), and a highly specific test rules disease in when positive, SpPin. B reaches for the weakest number on the table; in a low-prevalence setting a test scores high accuracy simply by calling everything negative, which is useless for this decision. D names a real and useful quantity, but the positive predictive value describes what a positive result means, and it swings with prevalence — her decision hangs on a negative.

**Common trap:** Reaching for the test with the best headline accuracy when the decision depends on trusting a negative.

Source: [STARD 2015](https://www.equator-network.org/reporting-guidelines/stard/)

## Question 7

A caries-detection aid with 90% sensitivity and 90% specificity was validated where 200 of every 1,000 sites were truly diseased, giving a positive predictive value of 69%. A dentist now applies it to a low-risk adult recall population in which roughly 20 of 1,000 sites are diseased. What should she expect?

- A. The same 69% positive predictive value, because the test itself has not changed
- B. A higher positive predictive value, because healthy populations produce fewer errors
- C. Lower sensitivity, roughly 16%, because the disease is rarer in this population
- D. About 16% positive predictive value — roughly five false alarms per true find

**Answer D:** About 16% positive predictive value — roughly five false alarms per true find

Run the same test at 2% prevalence over 1,000 sites: 18 true positives, 2 false negatives, 882 true negatives, and 98 false positives, so the positive predictive value is 18 ÷ 116 = 16% while the negative predictive value rises to 99.8%. Five flags in six are false, purely because the population changed. A treats a row-wise quantity as though it were a column-wise one; sensitivity and specificity belong to the test, predictive values to the population. B has the direction backwards: rarer disease makes a positive result less trustworthy, not more. C moves the wrong number — sensitivity is computed among the diseased only, so it does not shift with prevalence. Confirm any flagged site clinically before intervening.

**Common trap:** Carrying a predictive value from the validation study into a population with a different disease frequency.

Source: [STARD 2015](https://www.equator-network.org/reporting-guidelines/stard/)

## Question 8

A dentist compares two published adjuncts for detecting proximal lesions and wants a summary of test performance that will not shift when she moves between her private recall list and a community clinic. One adjunct reports 90% sensitivity and 90% specificity. Which number should she use, and what is its value?

- A. Likelihood ratios: LR+ is 9 and LR− is 0.11 for this test
- B. Predictive values, which combine sensitivity and specificity into numbers free of prevalence
- C. Accuracy, which counts every correct classification the test makes across the sample
- D. Likelihood ratios: LR+ is 0.11 and LR− is 9 for this test

**Answer A:** Likelihood ratios: LR+ is 9 and LR− is 0.11 for this test

LR+ = 0.90 ÷ 0.10 = 9 and LR− = 0.10 ÷ 0.90 = 0.11. Prevalence appears nowhere in either formula, which is why likelihood ratios avoid the dependence that makes predictive values travel; an LR+ above about 10 meaningfully raises disease probability and an LR− below about 0.1 meaningfully lowers it. B is exactly backwards: predictive values are the quantities that move with prevalence, which is the problem she is trying to escape. C offers the weakest number on the table, since in a low-prevalence population high accuracy is achieved by calling everything negative. D reverses the two ratios; the positive ratio must exceed 1 and the negative ratio must fall below it, or a positive result would argue against disease.

**Common trap:** Assuming that any single summary number describes the test rather than the population it was used in.

Source: [STARD 2015](https://www.equator-network.org/reporting-guidelines/stard/)

## Question 9

A device study reports 98% sensitivity for detecting proximal caries. Reading the methods, a dentist finds that only sites the device flagged were taken forward to the reference standard, while unflagged sites were recorded as sound without further verification. Which bias most directly inflates the reported sensitivity?

- A. Recall bias, because the examiners remembered which sites the device had flagged
- B. Verification bias: only test-positive sites received the reference standard, which inflates sensitivity
- C. Publication bias, because a device study with a null result would not appear
- D. Attrition bias, because the unflagged sites dropped out before the reference standard

**Answer B:** Verification bias: only test-positive sites received the reference standard, which inflates sensitivity

Sensitivity is TP ÷ (TP + FN), and false negatives can only be counted if test-negative sites are also verified. When the reference standard is applied only to flagged sites, missed lesions are silently reclassified as true negatives, the denominator loses its false negatives, and sensitivity is inflated — verification, or work-up, bias. The diagnostic-accuracy reporting standard exists partly to make participant selection and the reference standard visible for this reason. A misplaces recall bias, which concerns exposure remembered after an outcome and belongs to case-control studies. C describes which studies reach print, not how this one was conducted. D borrows a term for participants who differ from those who stay; nobody dropped out here, because the protocol never verified them in the first place.

**Common trap:** Accepting a spectacular sensitivity without asking who received the reference standard.

Source: [STARD 2015](https://www.equator-network.org/reporting-guidelines/stard/)

## Question 10

A 58-year-old man has twelve sound restorations, no current carious lesions, and a DMFT of 14. Reviewing the community survey these data feed, a colleague concludes that this population is suffering severe active caries. How should the dentist correct that reading?

- A. The DMFT is wrong and should be recalculated using only the decayed component
- B. A DMFT of 14 does confirm high current disease, since the index counts diseased teeth
- C. M and F record treatment history, so a high DMFT can mean no active disease
- D. A high DMFT with no lesions means the examiner was not calibrated to WHO criteria

**Answer C:** M and F record treatment history, so a high DMFT can mean no active disease

DMFT counts Decayed, Missing, and Filled Teeth, with dmft for primary teeth and DMFS or dmfs for surfaces. Only the D component reflects current disease; M and F are treatment history and never revert, which is why caries experience climbs with age even where new-lesion rates are falling, and why this man scores 14 with nothing active. A discards the index rather than reading it — the decayed count is a different measure, not a correction. B misreads the definition in the direction the index most often misleads. D blames calibration for a result the index produces by design; the WHO manual's standardised diagnostic criteria and index age groups exist to make surveys comparable across populations, not to prevent treatment history from accumulating.

**Common trap:** Reading a cumulative lifetime index as a measure of disease present today.

Source: [WHO Oral Health Surveys: Basic Methods, 5th ed](https://www.who.int/publications/i/item/9789241548649)

## Next step

[Take the free INBDE practice test](https://dentovio.com/inbde/free-practice-test)

Official reference: [JCNDE — Integrated National Board Dental Examination](https://jcnde.ada.org/inbde). Original exam-style questions written for study, never recalled exam content. Independent educational preparation, not clinical advice, and not affiliated with or endorsed by the Joint Commission on National Dental Examinations.
