Researchers analyzed 7,534 citations from Perplexity AI's sonar and sonar-pro models when prompted for software recommendations across 380 categories. The study focused on the quality and origin of the cited domains to understand the models' grounding behavior.
The analysis revealed that 59.8% of the citations pointed to domains ranked worse than #100,000 in the Tranco top-1M list, and 23.4% were to domains not present in the top million at all. The median Tranco rank for cited domains was 71,611, indicating a reliance on less authoritative websites.
A significant finding was the identification of three websites, all created in December 2023, that collectively published 215,128 machine-generated "best <category>" pages. These sites were frequently cited by the Perplexity models, suggesting that the AI is being grounded on newly created, potentially low-quality content farms.
The reliance on such sources raises questions about the accuracy and trustworthiness of the information provided by AI models like Perplexity. The study highlights a potential vulnerability in web-grounded AI systems, where newly generated, unvetted content can disproportionately influence outputs.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
An analysis of Perplexity AI's web-grounded models found that a significant portion of their software recommendations cite low-ranked or unranked domains, including 215,128 machine-generated "best software" pages from three sites created in December 2023. This indicates that Perplexity's grounding process frequently relies on newly created, low-quality content farms rather than established, authoritative sources, raising concerns about the reliability of its recommendations.