Can ChatGPT Tell If Text Was Written by AI? How Accurate AI Detectors Really Are
It is a natural experiment. You have a long piece of writing - a contractor’s blog draft, a job applicant’s cover letter, a 3,000-word article an agency just delivered - and you paste it into ChatGPT with one question:
“Was this written by AI?”
You will get an answer. This post is about what that answer is actually made of, how much weight it can carry, and what the better question is.
What ChatGPT is really doing when you ask
ChatGPT has no way to check for a hidden signature in text. OpenAI’s own provenance tooling covers images and audio, not text, and the vendor text watermarks that do exist - Google’s in Gemini and, since August 2026, Anthropic’s in newer Claude models - can only be read by a detector holding that vendor’s key. A chatbot reading your paste has no such key.
So when it answers, it is doing the same thing a human editor does: forming an impression from the writing. A careful model will tell you as much. Ask it and you will typically get something shaped like this:
Assessment: moderately likely AI-assisted
Signals
- unusually uniform sentence structure
- repeated transition patterns (“Moreover,” “In addition,” “Ultimately,”)
- generic elaboration that restates the heading rather than adding a fact
- consistently polished grammar with no idiosyncratic phrasing
- rhetorical structures that repeat (triads, “not X but Y” reversals)
- some sections that read noticeably more human than others
And a verdict on a scale like strong AI characteristics / ambiguous / strong human characteristics.
That is a useful, honest read. What a well-behaved model will not do is tell you “87% chance AI wrote this,” because that number would imply a precision the method does not have. If it does give you a percentage, treat the number as decoration on top of the qualitative signals, not as a measurement.
The most valuable version of the exercise is not “AI or human?” but “show me, section by section, what makes each part read as machine-smoothed or human-lumpy.” That is a legitimate editing tool. The binary verdict is not.
What “accuracy” means for AI detectors, and why the numbers mislead
Dedicated detector tools (GPTZero, Originality, Turnitin’s AI indicator, Copyleaks, and a dozen more) do a calibrated version of the same statistical read. They score features such as:
| Feature | What it measures | Why unedited AI text scores “high” |
|---|---|---|
| Perplexity | How predictable each next word is | Models pick likely words by design |
| Burstiness | Variation in sentence length and rhythm | Model output tends toward even pacing |
| Lexical spread | Range and rarity of vocabulary | Defaults to common, safe words |
| Structural repetition | Recurring sentence and paragraph shapes | Models reuse patterns that “worked” |
| Uniformity across sections | Whether voice drifts | Single-pass output does not drift |
The trouble is not that these features are meaningless. It is that they describe smooth writing, and plenty of human writing is smooth, while plenty of AI writing has been roughed up. Consider the track record:
- OpenAI’s own classifier, launched January 2023, was retired that July for low accuracy. At launch it caught about 26% of AI-written text and mislabelled about 9% of human-written text as AI.
- Non-native writers get flagged. A Stanford study ran essays by non-native English speakers through popular detectors; a majority were classified as AI-generated, while essays by native speakers were mostly classified correctly. Simpler vocabulary and regular structure look “smooth.”
- Paraphrasing breaks detection. Researchers showed that running AI text through a paraphraser collapsed detection rates across a range of tools, and separately argued that reliable detection may be theoretically out of reach as models improve. Even before editing, ordinary prompt engineering - “write this in a more varied, personal style” - can push output past detectors.
- Institutions have backed away. Several universities disabled AI-detection features in 2023 and 2024 over false-positive concerns, and vendors themselves publish caveats advising against using scores as sole evidence.
When a vendor advertises “99% accuracy,” ask: measured on what? Usually on clean, unedited model output versus clean human text, in the vendor’s own benchmark. Real documents in 2026 are rarely either.
Why length helps but does not settle it
More text is genuinely better evidence. Across 3,000 words you can look for consistency of voice, vocabulary distribution, whether transitions repeat, and - especially - whether one section suddenly reads differently from the rest. Short texts, a two-line email or a social post, give a detector almost nothing to work with; scores on short text should be ignored outright.
But length cannot resolve the central ambiguity, because of how writing actually happens now:
AI draft → human edits → AI tightens → human edits
is nearly indistinguishable from
human draft → AI proofread
Both carry both fingerprints. And a model given a specific voice to write in - real vocabulary, real sentence-length preferences, real examples of past writing - produces prose that no longer matches the “smooth default” a detector was trained on. Detection assumes a stereotypical AI style; the stereotype is a property of lazy prompting, not of AI.
Where a detector verdict should never be the deciding evidence
Anywhere the consequence is real:
- Hiring and performance. Rejecting a candidate or marking down an employee on a detector score means acting on a method with documented bias against non-native writers.
- Academic discipline and plagiarism accusations. The false-positive rate that is tolerable for a spam filter is not tolerable when it ends someone’s semester.
- Legal and contractual disputes. “The tool said 92%” does not survive cross-examination once the tool’s own caveats are read into the record.
- Client and vendor disputes. If a client accuses your team of “just using ChatGPT,” a competing detector score is not a defense. A brief, a review trail, and a named approver are.
For those situations the responsible output is qualitative and specific: which passages show which signals, and how confident that makes you. Not a number.
The better questions (especially if you publish for a living)
If your interest is not policing but publishing - you run marketing for a company and want to know whether your content is good enough - then “was this written by AI” is the wrong question, and it is worth being deliberate about swapping it out. Search engines have said since 2023 that they judge helpfulness, not production method, and readers do not run detectors. Here is what they actually respond to:
- Is it specific to this reader? Their industry, their problem, their vocabulary, an example they recognize. Generic elaboration is what both detectors and humans notice first - and it is a briefing problem, not an AI problem (why AI content sounds generic).
- Does it sound like us? A written brand voice that a model can actually follow - words we use, words we ban, how formal, how long the sentences run - does more to make content read as “human” than any humanizer tool.
- Is every claim correct and every promise one we can keep? That is a review job: language, voice, accuracy, and compliance, checked before publishing rather than after a complaint.
- Who approved it? A named person on the approve button is what makes AI-assisted work defensible - to a client, to a regulator, to yourself.
Content that clears those four does not need to worry about a detector. Content that fails them reads as generic whether or not a tool ever flags it.
A practical note on “humanizer” tools: rewriting AI text purely to beat detectors optimizes for the wrong target. It usually makes prose worse - odd word swaps, broken rhythm - to satisfy a scorer no reader consults. Spend the effort on the brief and the voice instead; the detector score follows.
Frequently asked questions
Can I ask ChatGPT whether text was written by AI?
Yes, and it will give you a reasoned impression with signals. It cannot prove authorship - there is no hidden tag to check - and any percentage it offers should be read as decoration, not measurement.
How accurate are AI detectors?
Inconsistent enough that OpenAI retired its own; documented false positives against non-native writers; substantial accuracy loss after paraphrasing or human editing. Vendor accuracy claims are measured on clean benchmark text, not real edited documents.
Does longer text help?
It supports more defensible directional judgments and section-by-section analysis, but it cannot separate “AI draft, human edited” from “human draft, AI proofread.”
What should I ask instead?
Whether the content is specific, on-voice, correct, and approved by a person. Those questions have real answers and decide whether the content works.
The bottom line
ChatGPT can give you a thoughtful, signal-by-signal opinion on whether a long piece reads as AI-assisted. It cannot prove it, and neither can the detectors that put a percentage on the same guess. Use the qualitative read as an editing aid; never use the score as evidence in anything that matters to a person’s livelihood. And if the reason you were asking is that you publish content and want it to land, stop asking whether a machine touched it and start asking whether it is specific, on-brand, correct, and signed off.
That is the workflow Marqeable is built around: agents draft campaigns from a real brief and your brand voice, specialist review checks language, voice, and accuracy before you see the draft, and nothing sends without your approval.
Marqeable runs your campaigns, answers every visitor, text, and email in seconds, and turns them into booked jobs and meetings - even at 9pm on a Saturday. We’re in private beta with a small early cohort. Get early access
