Complete, plain-language reading

Why Gene Studies Use Such Tiny P-Values

This reading contains every idea and every piece of evidence needed for today's decision. The research links at the end are optional.

1

Why this matters

Large searches are powerful only when their many chances to be wrong are counted.

2

The question you are trying to answer

What happens when thousands of sensors each can false-alarm?

3

Begin with the idea you already earned

Effect size and confidence interval show magnitude and precision; a p-value measures model compatibility, not truth.

4

Study the analogy before the biology

Thousands of airport alarms require a stricter screen and a second check
  1. What happens when thousands of sensors each can false-alarm?
  2. Why set the alert rule before scanning?
  3. What does the second scanner add?
5

Turn the analogy into three rules

Rule 1: Count how many tests are being run.
Rule 2: Predefine a correction or threshold suited to the analysis.
Rule 3: Replicate the signal in independent data.

Limit: Genome signals are not threats, and statistical correction does not determine biological importance.

6

Map those rules onto the biology

Control false positives in a common-variant GWAS
Many sensorsMillions of variant tests
Strict alert ruleMultiple-testing threshold
Second scannerIndependent replication

When a study runs many tests, some small p-values can appear by chance. Multiple-testing methods reduce that false-positive risk.

A threshold near 5 x 10^-8 is widely used for many common-variant GWAS analyses. Different designs may need different corrections.

Replication tests the signal in new participants. Researchers also check data quality, ancestry structure, effect direction, and biological follow-up.

7

Read Mateo's labeled case evidence

EXP08-E1

Testing one million variants at p below 0.05 would produce many chance hits under the null.

A usual single-test threshold is not adequate for a genome-wide scan.

EXP08-E2

For many common-variant GWAS analyses, p below 5 x 10^-8 is a conventional threshold.

The convention is not universal for every genetic analysis.

EXP08-E3

The candidate peak repeats in an independent cohort with the same direction of effect.

Replication lowers concern that the first signal was sample-specific chance or error.

8

Make the concrete decision

You are the GWAS quality-control lead.

A peak has p equal to 2 x 10^-6 in the discovery sample and no replication result.

  1. Label it suggestive and require the planned threshold, quality checks, and replication.
  2. Declare it significant because it is below 0.05.
  3. Use 5 x 10^-8 as a magic rule for every genetic test.

Choose the report language and explain how multiplicity and replication change the claim.

Claim ceiling: You may apply the supplied common-variant GWAS convention. You may not treat that number as universal or proof of mechanism.

9

Write the 10-year takeaway

Many tests inflate false positives, so thresholds and independent replication must be planned before results are seen.

  • Why is p below 0.05 too loose for one million tests?
  • What does replication add beyond a smaller p-value?
10

Glossary in plain English

Labeled illustration: multiple testing
multiple testing

Running many statistical tests at once, which raises the chance of a false positive and requires a stricter cutoff to stay reliable.

Labeled illustration: false positive
false positive

A test result that signals a condition is present when it actually is not, a kind of error that can lead to needless worry or treatment.

Labeled illustration: Bonferroni correction
Bonferroni correction

A math adjustment that makes the cutoff for significance stricter when you run many tests, so chance results are not mistaken for real ones.

Labeled illustration: genome-wide significance
genome-wide significance

A very strict p-value cutoff (about 5 in 100 million) used in genome-wide studies so that testing millions of spots does not produce false hits.

Labeled illustration: false discovery rate
false discovery rate

The expected share of results called significant that are actually false alarms, a key check when many tests are run at once.

Labeled illustration: replication
replication

Repeating a study to see if the result holds; a key test of whether a finding is real.