Your Text Failed the AI Detector: An Experiment with Russian-Language Texts
Anthropic
Moonshot AI
OpenAI
DeepSeek
Сбер
Яндекс
Text.ru
GPTZero
An experiment tested four AI detectors against texts generated by six large language models and human-written texts in Russian. The results show that detectors struggle with short texts and have 'blind spots' for certain models. The article discusses the principles behind AI detection and the controversy surrounding its accuracy.
The article from Habr's AI hub explores the rise of AI-generated content and the proliferation of AI detectors. It notes a Graphite study showing that in November 2024, AI-generated texts outnumbered human-written ones, with an updated study in spring 2026 claiming parity. The author describes how typical detectors work, citing Yandex's Neurodetector and Text.ru's AI detector as examples, and mentions that OpenAI shut down its classifier in 2023 due to low accuracy. Sber's GigaCheck, based on a fine-tuned Mistral model, claims 94.7% accuracy. The experiment used four detectors (Yandex Neurodetector, GigaCheck, Text.ru AI detector, GPTZero) and six LLMs (Claude Opus 5, Kimi K3, GPT-5.6 Sol, GLM 5.2, DeepSeek V4 Pro, GigaChat 3.5 Ultra). Three human-written texts on pangrams, published before 2021, served as control. The models generated texts of similar lengths, and all texts were checked once by each detector. Results were normalized to AI/HUMAN categories. The author acknowledges limitations: small sample, single run, and the date of July 27, 2026. Findings indicate that longer texts are classified more accurately, while short texts cause more errors. Notably, Yandex Neurodetector misclassified three Claude Opus 5 texts as human, and Text.ru mistakenly flagged GPT-5.6 Sol texts as human, suggesting 'blind spots'. The article refrains from recommending any detector.
- Сокращения
- LLM = Large Language Model — большая языковая модель
Source: Habr — хаб ИИ —
original
