Close Menu
  • Home
  • News
  • Security
  • Privacy
  • Cybercrime
    • Threat Groups
    • Ransomware
    • Explainers
    • Stealer Logs
  • AI
  • OSINT
  • Tools
    • Ransomtracker
    • Stealercheck
    • FortiBleed Checker
  • Newsletter
  • About Us
Facebook X (Twitter) Instagram Threads
Ransomnews
  • Home
  • News
  • Security
  • Privacy
  • Cybercrime
    • Threat Groups
    • Ransomware
    • Explainers
    • Stealer Logs
  • AI
  • OSINT
  • Tools
    • Ransomtracker
    • Stealercheck
    • FortiBleed Checker
  • Newsletter
  • About Us
Facebook X (Twitter) LinkedIn
Ransomnews
AI

“Delve” is dead: AI writing tells expire in 18 months

Martynas VareikisBy Martynas VareikisAugust 21, 2026Updated:August 22, 2026No Comments13 Mins Read51 Views
Share Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
Delve is dead, the dash took its place. AI lexical markers peak at 0.853 per 1,000 words in 2024 and fall to 0.384 by 2026, while dash usage rises to 0.877, across 21,442 arXiv abstracts.
Share
Facebook Twitter LinkedIn Pinterest Email Copy Link

The word “delve” has almost vanished from computer science research. In our sample of arXiv security and machine learning abstracts, it peaked in 2024 and then fell 94% by 2026. Nine other markers people use to spot AI writing collapsed alongside it. In the same abstracts, over the same two years, dash usage doubled and reached a 14-year high. Ransomnews measured 21,442 abstracts, 3.9 million words, from 2013 to 2026. The tells did not go away. They changed shift.

That last sentence is the whole story, and it has a cost attached. If the surface features that betray machine-generated text turn over roughly every eighteen months, every screening rule built on them decays on the same schedule. Word lists assembled in 2024 are now checking for a signature that current models mostly stopped emitting. Some of them are checking for the opposite of the truth.

What we set out to test, and what we actually found

This started as a much duller question. The em dash had become the internet’s favourite AI accusation through 2025, and we wanted to know whether publishers had responded by dropping it. So we built a corpus and started counting.

The answer to that question turned out to be no, mostly, which is a fine result but not much of an article. The interesting thing was sitting in the control data. Academic writing in AI-adjacent fields was moving, and it was moving in a direction nobody had predicted, at exactly the moment the famous vocabulary tells were dying.

So we narrowed the study to arXiv and rebuilt it properly around three categories:

  • cs.CR, cryptography and security. The field this publication covers.
  • cs.LG, machine learning. Maximum exposure to AI drafting tools, by definition.
  • cs.DS, data structures and algorithms. The control. Same academic register, same submission pipeline, same review culture, but the furthest from the habit of asking a chatbot to tidy your abstract.

Every year is sampled the same way: January, April and July, drawn at random from the complete pool of submissions in each of those months. That sounds fussy. It is the single most important decision in the study, and getting it wrong the first time cost us a headline. More on that below, because the mistake is instructive and other people will make it.

Which words gave the game away, and when they stopped

The marker list is not ours. It comes from Dmitry Kobak and colleagues at Tübingen, whose Science Advances paper analysed 15 million PubMed abstracts and found that at least 13.5% of 2024 biomedical abstracts carried traces of LLM processing. Their method is elegant: instead of guessing which text is machine-written, they measured which words suddenly appeared far more often than history predicted. “Delve” was the standout.

We tracked ten of those markers through our own corpus. In cs.CR and cs.LG combined, the marker rate sat at 0.100 occurrences per 1,000 words in 2021. It reached 0.222 in 2023. Then 2024 arrived and it hit 0.853, roughly four times the previous year in a single jump.

And then it fell. 0.664 in 2025. 0.384 in 2026, a 55% drop from peak in two years.

MarkerPeak yearFall from peak by 2026
delve202494%
meticulous202491%
intricate202485%
showcase202475%
realm201973%
pivotal202466%
underscore202562%
nuanced202548%
seamless202518%
harness2026still rising
Ten lexical markers in cs.CR and cs.LG abstracts. Ransomnews analysis of arXiv data, 2013 to 2026.

Nine of ten are in retreat. The one still climbing, “harness”, is the one nobody has made a joke about yet.

We are not first to the reversal. Mingmeng Geng and Roberto Trotta found the same thing across academic writing more broadly, with “delve” dropping sharply not long after it was publicly named as a ChatGPT giveaway in early 2024. They gave the pattern a name that deserves wider use: human-LLM coevolution. Writers take the machine’s draft and then scrub the words they know will out them. Awareness of a tell destroys the tell.

The dash went the other way

Here is where it gets strange. Dash usage in those same abstracts is close to a mirror image of the marker curve.

In cs.CR, the rate held at 0.341 dashes per 1,000 words across 2019 to 2024. Across 2025 and 2026 it reached 0.664, a rise of 94.5% (p = 0.0002 on a permutation test). cs.LG rose harder still, 136.8% (p = 0.0002).

cs.DS, the theory control, went from 0.390 to 0.431. That is a 10.5% wobble with a p value of 0.55, which in plain terms means it did not move at all. The mathematicians kept writing the way they have always written.

// THE TELLS SWAP PLACES, 21,442 ARXIV ABSTRACTS Occurrences per 1,000 words. cs.CR and cs.LG combined. cs.DS is the theory control. Every year sampled the same way: January, April and July, drawn at random from the full month. 0.000.250.500.751.001314151617181920212223242526markers peak0.8530.877 dashes0.384 markers0.492 control AI lexical markersdash usagecs.DS control
YearAI lexical markersDash usagecs.DS control
20130.0460.4630.349
20160.0760.3850.476
20190.1390.3430.376
20210.1000.3090.411
20220.1580.3200.468
20230.2220.3210.327
20240.8530.2750.499
20250.6640.5070.364
20260.3840.8770.492
Occurrences per 1,000 words. Markers and dash usage are cs.CR and cs.LG combined.

The raw counts keep this honest. The 2026 cs.CR figure rests on 103 dashes across 131,166 words. Fifty-six of 657 abstracts contain at least one, against roughly 3% of abstracts in the baseline years. It is not one weird paper dragging the average: the ten dashiest abstracts hold 30% of the total in 2026, down from 69% in 2021. The habit spread out.

Does this prove LLMs are writing security papers?

No. And the test that should have proved it came back pointing the wrong way, which is the part of this we would happily have found differently.

The logic is simple enough. If LLM drafting is pushing the dash rate up, then the abstracts carrying the lexical signature should also carry more dashes. Two symptoms, one cause, same documents. So we split the 2024 to 2026 cs.CR abstracts on exactly that line.

  • Abstracts with no marker vocabulary: 0.586 dashes per 1,000 words, across 1,715 papers.
  • Abstracts with marker vocabulary: 0.346 dashes per 1,000 words, across 250 papers.

That is 41% fewer, not more. The p value is 0.154, so formally this is a null rather than a reversal, but it is nowhere near the positive association the theory needs. The two supposed tells do not travel together.

We chased the obvious alternative too. Papers about language models grew from 0.3% of cs.CR in 2021 to 40.3% in 2026, so the rise could simply be a change of subject matter. It is not. Papers with no language-model content rose 55.6% on their own (p = 0.005), and by 2026 they were using dashes more heavily than the language-model papers, 0.968 against 0.521.

So we have a field-level pattern that fits the LLM story and a document-level pattern that refuses to. Our reading: something real changed in how AI-adjacent researchers write, the timing is hard to ignore, and the mechanism is not established. Anyone who tells you this is proof of machine authorship is going further than the data goes.

Why every AI writing tell has a shelf life

A tell survives only while nobody is optimising against it. Three separate pressures are grinding away at these markers, and all three show up in the data above.

Vendors tune it out. Model providers read the same threads everyone else does. Once a word becomes a punchline, it becomes a training objective. “Delve” losing 94% of its frequency in two years is not writers acting alone.

Users edit it out. This is the coevolution effect. People accept the draft, then go through and strip the words they know will get them accused. What comes out the far side is machine-shaped text with its most famous fingerprints wiped off, which is arguably harder to catch than the unedited version was.

Innocent writers change too, and this is the one that should worry defenders. We are a case study. Ransomnews scrubbed em dashes from this site because readers had started treating them as an AI signal, and plenty of publications made the same call. Once human writers begin avoiding a construction in order to prove they are human, that construction stops separating machines from people. It starts measuring something else entirely: who is anxious about being mistaken for a machine.

Which may be part of why the arXiv numbers move the way they do. Researchers submitting to cs.CR are not writing for an audience that polices punctuation. They are largely insulated from that third pressure, in a way that a journalist or a marketing team simply is not.

What defenders should do about it

Treat every text-based AI heuristic as a perishable signature rather than a rule. The antivirus industry learned this lesson in the 1990s and the shape of the problem has not changed.

  • Date-stamp the word list. If your phishing triage, CV screening or submission review leans on AI-associated phrases, record when that list was built. Our data puts the working half-life at around eighteen months.
  • Never gate a decision on one marker. Flagging an email, a job application or a bug bounty report because it contains an em dash or the word “delve” was weak in 2024 and is misleading in both directions now.
  • Watch where your false positives land. As a tell becomes common knowledge, the people still using it skew toward those least plugged into the discourse. That correlates with non-native English speakers and with anyone outside tech. There is a fairness problem nested inside the detection problem.
  • Favour provenance over style. Header forensics, DKIM alignment, sending infrastructure and account history age far better than vocabulary does. Our workflow for detecting AI-generated phishing covers that stack in detail.
  • Re-baseline on a calendar. Test your heuristic against known-good and known-machine samples every two quarters. A signature nobody re-tests is a signature you are trusting on faith.

The lesson generalises past writing. Any detection scheme keyed to the surface output of a generative system inherits this treadmill, and the thing on the other side is being retrained continuously. We have made the same argument about prompt injection defences that pattern-match on known payload strings, and about where AI genuinely earns its keep in the SOC.

The mistake that nearly cost us the study

Two bugs are worth publishing because anyone repeating this work will hit both.

The first was sampling. Our initial collector took the first 600 submissions of each year in date order. cs.CR volume has grown roughly tenfold since 2013, so that window silently shrank every year: 2013 spanned eleven months, 2019 spanned three, and 2026 covered January alone. We were comparing most of a year against a few weeks. The fix is the fixed-month design described above.

The second was encoding, and it is the sort of thing that quietly invalidates punctuation research. In LaTeX, -- renders an en dash and --- renders an em dash. Our first metric counted the two-hyphen form and missed the three-hyphen form completely. That would not matter if the convention held steady, but it did not: --- accounted for 40% to 70% of all dashes in the 2013 to 2020 abstracts, then disappeared from the corpus entirely after 2021. Counting one encoding erased most of the early baseline and manufactured part of the rise.

Both fixes are in the numbers above. For what it is worth, the corrected effect is larger than the broken one, not smaller. We would have reported it either way.

What this study cannot tell you

We measure text, not authorship. Nothing here identifies any individual abstract as machine-written, and no author or paper is accused of anything. This is a population-level count and should be read as one.

The base rate is genuinely small. Academic abstracts barely use dashes at all. Even at the 2026 peak, cs.CR runs at one dash per 1,273 words, where general news prose runs four to nine times heavier. Big percentages on a small base deserve suspicion, which is why the raw counts sit next to every rate in this piece.

2026 is incomplete. Our figures cover January, April and July, which matches every other year by design, but the year is not finished and the trend could reverse.

And the marker list belongs to an earlier model generation. If it has stopped working as a detector, that is the finding. It also means our marker curve may be tracking the death of one specific vocabulary rather than the overall level of LLM use, which are not the same thing.

Frequently asked questions

Is the em dash a reliable sign of AI writing?

No. Dash usage varies enormously by house style, by publication and by the content management system rendering the text, and human writers have started avoiding dashes specifically so they do not look machine-generated. One punctuation mark cannot separate the two populations, and treating it as evidence produces false accusations in both directions.

How can you tell if text was written by AI in 2026?

Not reliably from style alone. Every vocabulary marker we tracked lost 48% to 94% of its peak frequency within two years of becoming public knowledge. Provenance signals such as sending infrastructure, DKIM alignment, account history and document metadata hold their value far longer than word choice does.

Why did “delve” become known as a ChatGPT word?

It appeared far more often in post-2022 academic writing than any prior period predicted, which made it a measurable statistical outlier rather than a matter of opinion. Kobak and colleagues used that class of excess vocabulary to estimate LLM involvement across 15 million biomedical abstracts, and “delve” was the clearest single example.

How long does an AI writing tell stay useful?

Roughly eighteen months on this evidence. The ten markers we tracked peaked in 2024 or 2025 and had lost between 48% and 94% of their peak frequency by 2026. “Delve” fell fastest, at 94%.

Can AI detection tools be trusted for security screening?

Not as a sole control. Any tool keyed to surface style features inherits the same decay curve as a hand-built word list, and the cost of a false positive on a job applicant, a bug report or an inbound email is high. Use provenance as the primary signal and style as weak corroboration at most.

Does this research prove academics are using ChatGPT to write papers?

No. It shows that writing in AI-adjacent arXiv categories changed in a way a theory-category control did not. The mechanism is not established, and our document-level test found no association between the two markers that the theory says should appear together.

What is human-LLM coevolution?

It is the term Geng and Trotta use for the feedback loop where writers edit machine output to remove known giveaways, so the observable signature of AI assistance shifts in response to public awareness of it. The practical consequence is that any detection method built on those signatures degrades over time without anyone touching the detector.

What data did Ransomnews analyse for this study?

21,442 abstracts totalling 3,928,378 words, pulled from the public arXiv API across cs.CR, cs.LG and cs.DS. The sample covers January, April and July of every year from 2013 to 2026, drawn at random from the full pool of submissions in each of those months.

Sources and further reading

  • Kobak D., González-Márquez R., Horvát E.-Á., Lause J., Delving into LLM-assisted writing in biomedical publications through excess vocabulary, Science Advances, 2 July 2025
  • Geng M., Trotta R., Human-LLM Coevolution: Evidence from Academic Writing
  • arXiv API, the source of every abstract in this analysis
  • Ransomnews, Detecting AI-generated phishing in 2026
  • Ransomnews, Vibe coding is shipping vulnerabilities at scale in 2026
  • Ransomnews, Our editorial team and research methods

The Ransomnews Monthly

One email a month. Original leak-site data, victim census updates, and the findings that did not make the articles. No spam, unsubscribe any time.

Double opt-in. We store your email, signup time, and IP for consent records (GDPR Art. 7). See our privacy policy.

Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
Previous ArticleRTM Locker interview: a ransomware actor on the RaaS market
Martynas Vareikis

Martynas Vareikis is the AI Editor at Ransomnews. He covers the intersection of artificial intelligence and information security — from machine-learning models in defensive tooling to the adversarial use of LLMs by ransomware operators, deepfake-driven social engineering, and the rise of agentic threats. His reporting focuses on translating fast-moving AI research into practical guidance for defenders, journalists, and the broader security community. Reach Martynas via [email protected].

Related Posts

Vibe coding is shipping vulnerabilities at scale in 2026

July 16, 2026

Prompt injection left the lab in 2026. It is in the wild now

July 16, 2026

Shadow AI is the new stealer-log jackpot in 2026

July 16, 2026

Comments are closed.

// The Ransomnews Monthly

What leaked, what held up

One email a month: the datasets we verified, and the ones that fell apart under scrutiny.

Double opt-in. We store your email, signup time, and IP for consent records (GDPR Art. 7). See our privacy policy.

// Free tool

Were you in a leak?

Check whether an email address has surfaced in infostealer logs. No signup, no data stored.

Run StealerCheck

// Live data

Ransomtracker

Victims as they are posted to ransomware leak sites, tracked continuously and checked against the claims.

Open the tracker

9,459 confirmed attacks tracked

Facebook X (Twitter) LinkedIn
© 2026 Ransomnews.com

Type above and press Enter to search. Press Esc to cancel.

Cookies on Ransomnews

We use strictly-necessary cookies to run the site and may use first-party analytics to understand which articles are read. Some pages contain affiliate links — when you click one, the affiliate network sets cookies on the merchant's domain to attribute the referral. See the Cookie Policy and Affiliate Disclosure for detail.

RANSOMNEWS.COM

Tracking the criminal infrastructure of the internet.

Independent coverage of ransomware, breach economics, threat actors, privacy, AI security, and the open-source investigation toolkit.

// Topics

  • News
  • Security
  • Privacy
  • Cybercrime
  • AI
  • OSINT
  • Threat Groups
  • Stealer Logs
  • Ransomtracker
  • Stealercheck
  • FortiBleed Checker

// Site

  • About Us
  • Editorial Team
  • Contact
  • Tip Line
  • Editorial

// Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Funding & Independence
  • RSS Feed
© 2026 Ransomnews.com · Tracking the criminal infrastructure of the internet.