Standard Deviation Fraud Tester
Are these standard deviations too similar?
Real data are noisy, and so are the standard deviations computed from them. When a paper's SDs barely move from one condition to the next, the numbers may not have come from real samples at all. This tool checks how often chance alone would produce SDs that similar.
With 15 to 20 people per cell, an SD typically lands within about 17% of its true value, so the top row is ordinary. The bottom row almost never happens by chance.
Enter the results table
Paste one row per condition, straight from a spreadsheet or typed by hand. Columns: study, condition, mean, sd, n. The condition column is optional. Each study needs at least two conditions.
Tabs, commas, or semicolons all work. The first line must be the column names.
Same seed, same answer.
Results
| Study | Conditions | N | Pooled SD | SD of SDs | Typical by chance | p |
|---|---|---|---|---|---|---|
| All studies combined (Fisher) | ||||||
"Typical by chance" is the median SD of SDs across the simulated studies. p is the share of simulated studies whose SDs were at least as similar as the ones entered.
How it works
The logic of the test
A sample's standard deviation is itself an estimate, and it varies from sample to sample. Under normality its standard error is roughly
SE(SD) ≈ σ / √(2n)
so with 15 people per cell, an SD wanders about 18% around its true value. Even if every condition in an experiment had exactly the same population SD, the observed SDs should differ noticeably.
People inventing numbers attend to the means, which carry the hypothesis, and write down SDs that look reasonable. Those SDs come out too uniform, because nobody intuits how much sampling noise there should be.
What the simulation does
For each study, the tool assumes every condition shares one population SD, set to the pooled SD of the reported cells. It then simulates the study thousands of times at the reported sample sizes, computes each cell's SD, and records how spread out those SDs are (the SD of the SDs). The p-value is the share of simulated studies whose SDs are at least as similar as the reported ones.
This null is generous to the authors. It is the setup that produces the least heterogeneity, since real manipulations often change spread as well as level, and floor and ceiling effects squeeze cells near the ends of a scale. Genuine data usually come out more heterogeneous than the null, not less.
For a single item on a whole-number scale, each cell is simulated from a rounded normal distribution tuned to match that cell's mean and the pooled SD. On a bounded scale, means matter: a cell near the floor can't spread out as far. For continuous data the means don't matter at all, because the sampling distribution of an SD under normality doesn't depend on the mean.
Studies are combined with Fisher's method. Because the smallest simulated p is 1/(B+1), the combined p is somewhat conservative.
Reading the result, and what it can't tell you
A small p is a reason to ask for the raw data, not proof of fabrication. That's how it worked in practice: Simonsohn flagged the anomalies from summary statistics, then obtained the raw data, which settled the question.
A p near 1 means the SDs differ more than the equal-SD model expects. That's common in real data and is no cause for concern here.
Studies with only two conditions have little power, since two SDs are often close by chance. The test is most informative with more cells per study and when the same pattern recurs across studies.
Reported SDs are usually rounded to two decimals. That adds a trivial amount of noise and doesn't meaningfully affect the result.
Source
Simonsohn, U. (2013). Just post it: The lesson from two cases of fabricated data detected by statistics alone. Psychological Science, 24(10), 1875–1888. doi:10.1177/0956797613480366