Whistle is a newly released speech recognition model specifically engineered for on-device applications. It targets a range of devices including mobiles, wearables, robots, smart home systems, automotive platforms, and microcontrollers. The model is distributed as a single 16.9 MB file.
The model performs three primary functions directly on the device: transcription, word timestamps, and speech embedding. Transcription handles 16 kHz mono audio clips up to 30 seconds long in one pass, supporting English, German, French, Spanish, Italian, Dutch, and Polish, with automatic language detection. Word timestamps provide start, end, and probability for each word. Speech embedding outputs encoder data per 80 ms frame without generating a transcript.
Whistle runs on the CPU without external dependencies and utilizes the same C++ engine and quantization as the Needle model. Its front end processes 16 kHz mono audio into 80 log-mel bins. The encoder consists of eight Simple Attention blocks, similar to Needle, using mHC residual lanes and a Monarch Hadamard MLP. The decoder features eight Laddered Simple Attention blocks with a different layer count compared to Needle's block list. Each decoder layer incorporates a gated cross attention to read from the encoder.
Decoding employs five beams scored by length-normalized log probability. Keyword biasing is implemented using an Aho-Corasick automaton to boost the log probability of specified phrases. The transcript is limited to 320 tokens, and the vocabulary includes 8,192 text pieces plus seven language-specific tokens, allowing detected language to be emitted as a token.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A new speech recognition model, Whistle, has been released, designed for on-device transcription, word timestamps, and speech embedding. It operates as a 16.9 MB file on the CPU without external dependencies, supporting seven languages for mobile, wearable, and embedded systems.