Research
What detection can and cannot tell you
We publish our evaluation approach because a detection product that hides its error profile cannot be used responsibly.
Error profile first
Precision and recall are reported together with the passage lengths and content types they were measured on.
Known failure modes
Short passages, translated text, heavily edited AI output and writing by non-native speakers all degrade reliability.
No single threshold
A likelihood range and a confidence level are more honest than a pass/fail line, so we never ship one.
Human-in-the-loop by design
The product is evaluated as a review aid, not as an autonomous classifier.
Adversarial pressure
Paraphrasing tools and humanizers are part of the evaluation set, not an afterthought.
Independent scrutiny
We support external audits of the methodology by customers and researchers.
No detection model is connected to this build, so no accuracy figures are published here. Any number shown elsewhere in the product is illustrative placeholder data.
Evaluation plan
What we will publish when the engine ships
A model card describing training and evaluation data, measured precision and recall by content type and length, false-positive rates for at-risk writer groups, and a changelog for every model version deployed.