← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Scribd uses Gemini Enterprise batch inference to classify over 400 million documents

🔄 Updated 2h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Scribd classified 400M+ documents (12B+ pages) with Gemini Enterprise.
  • Corpus-wide backfill completed in months.
  • Native PDF input processed over 99% of content as-is.
  • Gemini Enterprise batch prediction offered a 50% discount.

Scribd Classifies Content with Gemini Enterprise

Scribd, Inc. used Gemini Enterprise's batch prediction feature to perform trust and safety classification on its entire user-generated content corpus. This initiative involved over 400 million documents and 12 billion pages from Scribd and Slideshare platforms.

The classification process was completed within months, with Google Cloud scaling batch throughput to meet the project timeline. This allowed Scribd to efficiently review a large volume of content for compliance with community guidelines.

Technical Implementation and Benefits

Gemini's native PDF understanding was a key factor, enabling more than 99% of the corpus to be processed directly without requiring optical character recognition (OCR), rendering, or screenshotting pipelines. This eliminated the need for Scribd to build additional processing infrastructure.

The use of Gemini Enterprise's batch prediction also offered a 50% discount compared to interactive pricing, making large-scale LLM classification economically viable for Scribd's extensive content library.

Addressing Trust and Safety at Scale

Scribd, which includes products like Scribd, Slideshare, Everand, and Fable, manages hundreds of millions of user-uploaded documents. Maintaining trust and safety across this vast collection requires balancing content access with community protection.

Traditional methods for content moderation often involve specialized detection models for each policy area, which can be resource-intensive. Scribd evaluated various off-the-shelf tools and open models but found they did not meet the required quality or scale for their 400-million-document backfill.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~24 min · 20 stories · Sep 24

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Scribd, Inc. utilized Gemini Enterprise's batch prediction capabilities to classify more than 400 million user-uploaded documents across Scribd and Slideshare for trust and safety. This process, completed in months, allowed Scribd to analyze 12 billion pages of content using Gemini's native PDF understanding, avoiding the need for OCR or rendering pipelines.