turbopuffer is implementing a new storage architecture in its upcoming version 3. This redesign changes how documents and indexes are laid out, written, compacted, and queried within the system. The primary goal is to enhance search performance across the board, including vector search capabilities.
Beyond improving search, the new architecture also establishes a foundation for executing a wider range of SQL queries more quickly within turbopuffer. This expansion of functionality suggests a move towards supporting more diverse data operations than previously possible.
Initially launched as a serverless vector database (v1), turbopuffer specialized in cost-effective and fast vector searches, utilizing object storage and tiered caching. Version 2 added strong text and regex search. However, the original storage architecture, where the ANN vector index was the primary index, constrained certain query plans like GROUP BY and aggregations. The v3 update addresses these limitations by making the ANN index a secondary index.
In turbopuffer v1, documents consisted of an ID and a vector. The system used a hierarchical clustering index, starting with SPANN and later migrating to SPFresh for incremental indexing. This method clustered vectors into groups, forming a tree structure with centroids.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
turbopuffer is changing its storage architecture in version 3, moving away from a vector-primary index design. This update aims to improve search speed, including vector search, and enable faster execution of more SQL queries within turbopuffer.