A study found AI can identify the people behind pseudonymous accounts using only public posts and web searches. [Photo: Shutterstock]

A study has found that artificial intelligence can analyse posts from online accounts run anonymously or under pseudonyms and combine web searches with inference to identify the real person.

On Sept. 24 (local time), blockchain media outlet Decrypt reported that the ability to connect information scattered across public posts without hacking or data leaks raises a new privacy threat to online anonymity.

Researchers in February published a paper titled "Large-Scale Online De-anonymization with Large Language Models (LLMs)." The research involved teams from ETH Zurich, AI safety research organisation MATS and Anthropic.

The researchers built an automated pipeline in which AI extracts identity-related clues from online posts, searches for candidates and then infers from candidate information to confirm whether it matches a real person. They divided the process into four stages: extraction, search, inference and calibration.

In the extraction stage, the system compiles identity-related information such as a user’s area of residence, job, interests and career based on anonymous posts. In the search stage, it converts that information into embeddings to find candidates with similar meaning. In the inference stage, a stronger AI model compares information for each candidate and verifies whether clues extracted from the posts match real profiles. In the final calibration stage, the AI assesses its confidence to reduce the risk of mistaken identification.

To test performance, the researchers ran experiments on 338 users of the tech community Hacker News. These were users who had linked their real LinkedIn profiles in their Hacker News bios. The researchers created profiles that removed direct identifiers such as names, account names and personal URLs, then had an AI agent perform web searches.

The AI correctly identified the real identities of 226 of the 338 people. Recall was 67 percent and precision for identified results was about 90 percent. That is, the AI did not identify everyone, but about 9 out of 10 cases it presented as identified matched the actual person.

In a separate experiment, the researchers also used scientist interview materials released by Anthropic. The materials include interview records of 125 scientists covering AI use and research experience. Based on some of those interviews, the researchers estimated they could identify the real identities of at least 9 people.

This type of identification did not require hacking techniques or intrusion into personal data databases. The researchers conducted the experiment by searching public webpages and summarising and comparing post contents. They said the cost of running an AI agent on a single profile was about $1 to $4, and total experimental costs were under $2,000.

The researchers said the findings should not be interpreted to mean AI can immediately reveal the identities behind all anonymous accounts. To measure performance, they built an evaluation dataset based on data where real identities could be verified. In the Hacker News test, they removed direct identifiers from accounts originally linked to LinkedIn and then attempted to re-identify them. The paper said this evaluation approach could create conditions that are easier to identify than typical pseudonymous accounts.

Performance also fell as the pool of candidates grew. When the researchers expanded the candidate set to as many as 89,000 people, recall for the strongest inference-based method was about 55 percent at a 90 percent precision threshold. They said performance could drop further if the candidate pool grows, and they presented separate estimates for the internet as a whole.

The study has drawn attention because AI automated existing de-anonymisation methods. In the past, collecting clues across multiple posts and checking them against public sources required significant time and manpower. The researchers said LLMs can automate such work and sharply reduce the cost and time needed to identify pseudonymous accounts.

The researchers said in the paper there is a need to re-examine privacy threat models built around "practical obscurity" that has protected online pseudonymous users. Practical obscurity does not mean personal information can never be identified. It assumes that because public information is scattered, finding a specific person requires substantial cost and effort in practice. The researchers said AI can lower that cost barrier by automating search and information cross-checking.

Keyword

#Hacker News #LinkedIn #ETH Zurich #MATS #Anthropic
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.