AI-Detector.wiki
ANONYMOUS AGGREGATE RESULTS

AI Detector Benchmark™

How frequently do recorded AI detectors agree? This live benchmark aggregates verdicts saved through Detector History™ without publishing the writing samples themselves.

0Recorded tests
0History projects
0Detectors observed
0Comparable tests

Overall agreement

Unanimous verdict

0%

0 of 0 comparable tests had every recorded detector return the same verdict.

Majority, but not unanimous

0%

0 tests had more than half of recorded detectors agree, while at least one differed.

No majority

0%

0 tests had no verdict category supported by more than half of recorded detectors.

Method

A comparable test contains at least two recorded detector verdicts. “Agreement” means the same categorical verdict—Human, Mixed/Uncertain, or AI Characteristics—not identical probability scores.

Detector verdict distribution

DetectorTestsHumanMixedAI

No benchmark data has been recorded yet.

Pairwise agreement

How often two detectors returned the same categorical verdict when both were recorded on the same test. Small samples should not be interpreted as stable performance estimates.

Pairwise data will appear after tests include at least two detector results.

Benchmark limitations. This is an observational benchmark of user-recorded detector verdicts, not a controlled accuracy study. It does not establish which detector is correct, whether text was written by a human or AI, or the accuracy of any provider. The sample can be self-selected and incomplete. Public benchmark calculations use detector names and categorical verdicts; writing samples and free-text notes are not displayed here.
Run your own DIY comparison →