← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Whistle Speech Recognition Model Released for On-Device Applications

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Whistle is a 16.9 MB, CPU-only speech recognition model.
  • Performs transcription, word timestamps, and speech embedding on-device.
  • Supports English, German, French, Spanish, Italian, Dutch, and Polish.
  • Uses a similar C++ engine and quantization as the Needle model.
  • Processes up to 30 seconds of 16 kHz mono audio per pass.

Introduction of Whistle Model

Whistle is a newly released speech recognition model specifically engineered for on-device applications. It targets a range of devices including mobiles, wearables, robots, smart home systems, automotive platforms, and microcontrollers. The model is distributed as a single 16.9 MB file.

On-Device Functionality

The model performs three primary functions directly on the device: transcription, word timestamps, and speech embedding. Transcription handles 16 kHz mono audio clips up to 30 seconds long in one pass, supporting English, German, French, Spanish, Italian, Dutch, and Polish, with automatic language detection. Word timestamps provide start, end, and probability for each word. Speech embedding outputs encoder data per 80 ms frame without generating a transcript.

Technical Architecture

Whistle runs on the CPU without external dependencies and utilizes the same C++ engine and quantization as the Needle model. Its front end processes 16 kHz mono audio into 80 log-mel bins. The encoder consists of eight Simple Attention blocks, similar to Needle, using mHC residual lanes and a Monarch Hadamard MLP. The decoder features eight Laddered Simple Attention blocks with a different layer count compared to Needle's block list. Each decoder layer incorporates a gated cross attention to read from the encoder.

Decoding Process

Decoding employs five beams scored by length-normalized log probability. Keyword biasing is implemented using an Aho-Corasick automaton to boost the log probability of specified phrases. The transcript is limited to 320 tokens, and the vocabulary includes 8,192 text pieces plus seven language-specific tokens, allowing detected language to be emitted as a token.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~4 min · 3 stories · Oct 08

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

A new speech recognition model, Whistle, has been released, designed for on-device transcription, word timestamps, and speech embedding. It operates as a 16.9 MB file on the CPU without external dependencies, supporting seven languages for mobile, wearable, and embedded systems.