NewsMacroAI Can Now Doxx Your Anonymous Accounts: What the ETH Zurich and Anthropic Study Found

AI Can Now Doxx Your Anonymous Accounts: What the ETH Zurich and Anthropic Study Found

Author: Decrypt·

Key Takeaways

  • •Researchers from ETH Zurich, MATS, and Anthropic built an AI pipeline that matched anonymized Hacker News accounts to LinkedIn profiles with 90% precision using only web search, embeddings, and reasoning models.
  • •Each deanonymization run cost between $1 and $4 and required no data breach, hack, or insider access, which the researchers say makes the attack hard to block.
  • •In a test involving 338 Hacker News users who had linked their LinkedIn profiles in their bios, the AI correctly identified 226 of them, about 67% of the group.
  • •The study used a controlled best-case setup, and against a pool of 89,000 candidates the strongest method caught only about half of correct matches at 90% precision.
  • •The authors withheld their code, prompts, and all real identities uncovered, and the study passed ETH Zurich's ethics review board before publication.
AI Can Now Doxx Your Anonymous Accounts: What the ETH Zurich and Anthropic Study Found

Researchers at ETH Zurich, the AI safety group MATS, and Anthropic have built an AI pipeline that matched anonymized Hacker News accounts to real LinkedIn profiles with 90% precision — and the attack involved no hacking and no data breach. It ran on nothing but web search, embeddings, and reasoning models such as GPT-5.2, at a cost of roughly $1 to $4 per target, according to the study.

The paper's authors withheld their code, prompts, and every real identity they uncovered, and say the study passed through ETH Zurich's ethics review board before publication.

The research, first published in February, made the rounds on X and Reddit again this week, and the reaction was immediate: scores of alarmed social media users declared it the end of internet anonymity as they know it.

The paper, titled "Large-scale online deanonymization with LLMs," is available on arXiv and comes from researchers at ETH Zurich and the AI safety group MATS, working alongside Nicholas Carlini, a researcher at Anthropic, the company behind Claude. The researchers claim that large language models — the AI systems behind Claude, ChatGPT, and Gemini — can read a person's anonymous posts and figure out who they are in real life. No system gets breached; the models simply read, search, and reason, the same tasks these chatbots already perform millions of times a day.

The alarm online was captured in a widely shared X post:

Researchers built an AI that doxes any "anonymous" reddit account in under a minutes for $2. eth zurich and anthropic published a terrifying paper proving that "practical anonymity" on the internet is officially dead. they built a fully autonomous ai pipeline that takes your… pic.twitter.com/XMeXifhDf3

— Superman (@thesupermannx) September 23, 2026

"Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered," the paper reads.

How the AI actually pulls this off

The researchers broke the attack into four steps they call Extract, Search, Reason, and Calibrate.

"Extract" means feeding an AI a person's posts and having it summarize who they are: where they might live, their job, their hobbies, even a quirk like British spelling or a dog's name mentioned in passing.

"Search" turns that summary into an embedding — a string of numbers a computer can use to instantly find similar-sounding profiles among thousands of candidates.

"Reason" is where a more powerful model examines the best few matches and works out which one actually fits, cross-checking details the way a human investigator would, only much faster.

"Calibrate" has the AI rate its own confidence, so it only guesses when it is fairly sure and stays quiet otherwise.

In one test, the team gathered 338 users of the mostly pseudonymous Hacker News message board who had linked their LinkedIn profiles in their bios, stripped out anything identifying, and then asked an AI agent to figure out who they were using nothing but web search. It correctly named 226 of them — 67% — and was wrong on roughly one in 10 of the guesses it actually made. Those two numbers capture the study's headline logic in miniature: whenever the system did venture a guess, it was right about nine times out of ten, and across the full group it named two of every three targets.

In a separate test on transcripts from Anthropic's own interview research with 125 scientists, the AI correctly identified at least 9 people just from how they described their work.

The real finding is the cost

Running one of these searches costs between $1 and $4 in AI subscription fees, according to the paper. The attack requires no data breach, no hack, and no insider access. It uses the same web search and summarization abilities built into ordinary chatbots.

That is exactly why the researchers say it is hard to block: there is no single "deanonymize this person" switch to disable, just a chain of individually harmless-looking tasks. For the companies building these models, that complicates mitigation: the risk is not a flaw in any one product but an unintended use of the same search-and-summarize abilities that make chatbots useful in the first place.

Nor is this the first time a handful of details has unmasked someone. Back in 2008, researchers cracked Netflix's supposedly anonymous movie-rating dataset by matching it against public IMDb reviews. The difference now is that AI does the matching on messy, unstructured text — jokes, comments, offhand mentions — instead of neat spreadsheets, and it does the work on its own.

Now, the "everyone calm down" part

The scariest-sounding numbers in the paper come with a catch. To measure success, researchers needed subjects whose real identities they already knew, so they picked accounts that had already linked to LinkedIn, or split one person's post history into two halves and hid the connection. That is a controlled, best-case setup — not proof that any random pseudonymous account can be cracked today.

Scale also cuts against the study's most alarming figures. The bigger the haystack of possible candidates, the harder the needle is to find. Against a pool of 89,000 candidates, the strongest AI method caught only about half of correct matches at 90% precision. Precision refers to how often the system's guesses were right; recall refers to how many real matches it actually caught.

The researchers also did not publish their tools, prompts, or any of the real names they uncovered, and the study went through an ethics review board before release. They are not handing anyone a doxxing kit. They are documenting a capability the authors say already exists in current AI models, paper or no paper.

Why this matters even if you have never posted on Reddit

Crypto users already know this fear intimately. A wave of doxxings and kidnapping attempts following the 2025 Coinbase data breach, among many others, showed how fast a leaked identity can turn into a real address at someone's door. AI-driven deanonymization is the same threat with the breach removed: it works off what you have already posted in public, no leak required.

The study also lands a few months after Anthropic disclosed that state-backed hackers used Claude to run most of a cyberespionage campaign on their own, a reminder that AI misuse research keeps outrunning the guardrails meant to contain it.

For anyone who posts under a pseudonym because of activism, abuse, sexuality, immigration status, or a job that frowns on public opinions, the finding means that hometowns, employers, and even a pet's name — scattered across years of old comments — add up into a fingerprint. That was always possible; AI just makes it faster and cheaper to do. Part of the difficulty is an asymmetry: changing how you post tomorrow does nothing about the years of comments already sitting in public, and the method runs on precisely that archive.

For those concerned about this kind of doxxing, the answer is not panic but habit: attach fewer specific, identifying details to any one pseudonym, and for anything genuinely sensitive, use tools built not to retain your data at all.