Standard Deviation Fraud Tester

Are these standard deviations too similar?

Real data are noisy, and so are the standard deviations computed from them. When a paper's SDs barely move from one condition to the next, the numbers may not have come from real samples at all. This tool checks how often chance alone would produce SDs that similar.

Four real samples, same population 1.081.371.191.46 Four cells in a fabricated table 1.271.261.281.27

With 15 to 20 people per cell, an SD typically lands within about 17% of its true value, so the top row is ordinary. The bottom row almost never happens by chance.

Enter the results table

Paste one row per condition, straight from a spreadsheet or typed by hand. Columns: study, condition, mean, sd, n. The condition column is optional. Each study needs at least two conditions.

Tabs, commas, or semicolons all work. The first line must be the column names.

How the outcome was measured
From to

Same seed, same answer.

How it works

The logic of the test

A sample's standard deviation is itself an estimate, and it varies from sample to sample. Under normality its standard error is roughly

SE(SD) ≈ σ / √(2n)

so with 15 people per cell, an SD wanders about 18% around its true value. Even if every condition in an experiment had exactly the same population SD, the observed SDs should differ noticeably.

People inventing numbers attend to the means, which carry the hypothesis, and write down SDs that look reasonable. Those SDs come out too uniform, because nobody intuits how much sampling noise there should be.

What the simulation does

For each study, the tool assumes every condition shares one population SD, set to the pooled SD of the reported cells. It then simulates the study thousands of times at the reported sample sizes, computes each cell's SD, and records how spread out those SDs are (the SD of the SDs). The p-value is the share of simulated studies whose SDs are at least as similar as the reported ones.

This null is generous to the authors. It is the setup that produces the least heterogeneity, since real manipulations often change spread as well as level, and floor and ceiling effects squeeze cells near the ends of a scale. Genuine data usually come out more heterogeneous than the null, not less.

For a single item on a whole-number scale, each cell is simulated from a rounded normal distribution tuned to match that cell's mean and the pooled SD. On a bounded scale, means matter: a cell near the floor can't spread out as far. For continuous data the means don't matter at all, because the sampling distribution of an SD under normality doesn't depend on the mean.

Studies are combined with Fisher's method. Because the smallest simulated p is 1/(B+1), the combined p is somewhat conservative.

Reading the result, and what it can't tell you

A small p is a reason to ask for the raw data, not proof of fabrication. That's how it worked in practice: Simonsohn flagged the anomalies from summary statistics, then obtained the raw data, which settled the question.

A p near 1 means the SDs differ more than the equal-SD model expects. That's common in real data and is no cause for concern here.

Studies with only two conditions have little power, since two SDs are often close by chance. The test is most informative with more cells per study and when the same pattern recurs across studies.

Reported SDs are usually rounded to two decimals. That adds a trivial amount of noise and doesn't meaningfully affect the result.

Source

Simonsohn, U. (2013). Just post it: The lesson from two cases of fabricated data detected by statistics alone. Psychological Science, 24(10), 1875–1888. doi:10.1177/0956797613480366