AI Detector Benchmark™
How frequently do recorded AI detectors agree? This live benchmark aggregates verdicts saved through Detector History™ without publishing the writing samples themselves.
Overall agreement
Unanimous verdict
0%0 of 0 comparable tests had every recorded detector return the same verdict.
Majority, but not unanimous
0%0 tests had more than half of recorded detectors agree, while at least one differed.
No majority
0%0 tests had no verdict category supported by more than half of recorded detectors.
Method
A comparable test contains at least two recorded detector verdicts. “Agreement” means the same categorical verdict—Human, Mixed/Uncertain, or AI Characteristics—not identical probability scores.
Detector verdict distribution
No benchmark data has been recorded yet.
Pairwise agreement
How often two detectors returned the same categorical verdict when both were recorded on the same test. Small samples should not be interpreted as stable performance estimates.
Pairwise data will appear after tests include at least two detector results.