Scribd, Inc. used Gemini Enterprise's batch prediction feature to perform trust and safety classification on its entire user-generated content corpus. This initiative involved over 400 million documents and 12 billion pages from Scribd and Slideshare platforms.
The classification process was completed within months, with Google Cloud scaling batch throughput to meet the project timeline. This allowed Scribd to efficiently review a large volume of content for compliance with community guidelines.
Gemini's native PDF understanding was a key factor, enabling more than 99% of the corpus to be processed directly without requiring optical character recognition (OCR), rendering, or screenshotting pipelines. This eliminated the need for Scribd to build additional processing infrastructure.
The use of Gemini Enterprise's batch prediction also offered a 50% discount compared to interactive pricing, making large-scale LLM classification economically viable for Scribd's extensive content library.
Scribd, which includes products like Scribd, Slideshare, Everand, and Fable, manages hundreds of millions of user-uploaded documents. Maintaining trust and safety across this vast collection requires balancing content access with community protection.
Traditional methods for content moderation often involve specialized detection models for each policy area, which can be resource-intensive. Scribd evaluated various off-the-shelf tools and open models but found they did not meet the required quality or scale for their 400-million-document backfill.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Scribd, Inc. utilized Gemini Enterprise's batch prediction capabilities to classify more than 400 million user-uploaded documents across Scribd and Slideshare for trust and safety. This process, completed in months, allowed Scribd to analyze 12 billion pages of content using Gemini's native PDF understanding, avoiding the need for OCR or rendering pipelines.