Detector-first products assess provenance; humanizer-first products sell rewriting followed by another scan.
Do not promise the impossible.
Build a quality editor with an explainable style-signal index. Do not sell “bypassing every detector”: science cannot support that guarantee, while search performance depends more on a text's value, accuracy, and originality.
Results depend on language, domain, length, generator, detector version, and the way a text was edited.
There is no single scientific “AI percentage.” In the MVP, this is an uncalibrated 8–98 feature index on a 100-point display, not an authorship probability.
Google prohibits scaled, low-value content created to manipulate rankings, regardless of whether a human or AI produced it.
The market already sells two different jobs.
Features and prices come from vendors' official pages as of the memo date. Marketing accuracy claims are not treated as independent validation.
| Product | Focus | Detector | Editor | API | Monthly entry price |
|---|---|---|---|---|---|
| Originality.ai ↗ | Detector-first | Yes | No | Enterprise | $14.95/mo |
| GPTZero ↗ | Detector-first | Yes | No | Yes | Check at checkout |
| Copyleaks ↗ | Detector-first | Yes | No | Enterprise | $16.99/mo |
| Winston AI ↗ | Detector-first | Yes | No | Separate plan | $18/mo |
| QuillBot ↗ | Writing suite | Yes | Yes | No | $17.95/mo |
| Undetectable AI ↗ | Humanizer-first | Yes | Yes | Yes | $9.99/mo |
| StealthWriter ↗ | Humanizer-first | Yes | Yes | Not found | $20/mo |
“API not found” means that no public documentation was found on the official domain during this research; it does not prove that no private integration exists.
A detector measures a signal. It does not establish authorship.
Controlled tests can show high accuracy. Generalization to a new language, genre, model, or paraphrased text is a separate test.
11 models, 8 domains, 11 attacks, and 4 decoding strategies. Detectors lost robustness under attacks, new generators, and different generation settings.
RAID · ACL 2024 ↗This was the performance of OpenAI's public classifier on an English-language challenge set. It was retired in July 2023 because of its low accuracy.
OpenAI · 2023 ↗The NIST pilot result applies to a narrow summarization task, so it is not a market-wide assessment, but it demonstrates the limits of universality.
NIST AI 700-1 · 2025 ↗Seven detectors evaluated on 91 TOEFL essays written by non-native English speakers. This is a limited 2023 sample, not an assessment of contemporary Russian-language text.
Liang et al. · Patterns 2023 ↗Additional evidence: M4 ↗ documents poor generalization to new domains and LLMs; Sadasivan et al. ↗ show that paraphrasing can change detector results dramatically.
An explainable heuristic before calibration.
Estimate the strength of observable style patterns and suggest editorial checks. The index does not establish who wrote the text.
score = clamp(8 + Σ(normalized_feature × weight), 8, 98); thresholds: <36 low, 36–65 medium, ≥66 highThe weights are an MVP product heuristic. They have not yet been trained or calibrated on a labeled corpus.
Uncertainty
- <80 words: ±30 points
- 80–199 words: ±21 points
- ≥200 words: ±14 points
The interval reflects only the heuristic's length-based instability; it is not a statistical confidence interval until calibration is complete.
What the editor checks
- numbers, URLs, and required terms;
- the change in the same style signals;
- model flags requiring human review.
Checking explicit invariants does not establish factual or semantic equivalence across the entire text.
Adaptive editing v2
- Give the editor only the active, explainable signals found in the source text.
- Check the result locally: index, numbers, URLs, required terms, and acceptable length change.
- For balanced/substantial texts of up to 6,000 characters, run at most one refinement pass when the score remains above 35 or guardrails require review.
- Choose the second version only when guardrails pass and the score improves by at least 4 points; surface free-form review flags for manual review.
Reach the low range of the style index while preserving numbers, URLs, required terms, and acceptable length bounds; zero is not a product objective.
When the index can be described as a risk estimate.
Until this protocol is complete, the product must label the result as an “uncalibrated style index.”
- 01
Build a labeled RU/EN corpus covering human, raw AI, mixed, and professionally edited AI text; split it by sources, authors, prompts, domains, and generators.
- 02
Freeze an independent holdout set containing new models and domains; prevent author and template overlap between train and test sets.
- 03
Train or adjust weights only on the training set; on the holdout set, measure AUROC, precision/recall at a predefined FPR, Brier score, and ECE.
- 04
Publish a confusion matrix separately by language, length, genre, and editing type; return insufficient evidence for weak segments.
- 05
Compare numbers, URLs, required terms, semantics, and readability before and after rewriting; never use detector score as the sole optimization target.
Detector score ≠ ranking signal.
Google officially evaluates usefulness and quality, not the method of creation itself. Scaled, unoriginal content produced to manipulate search rankings may be treated as scaled content abuse, regardless of AI involvement.