Cactus has developed a new technique for on-device AI models to determine the reliability of their own answers. Small, local models offer speed and privacy but can sometimes produce incorrect information. The new method addresses this by integrating 'probes' within the model's checkpoint.
These probes score every answer with a confidence level between 0 and 1, returned as structured data. If a model's confidence in an answer is below a set threshold (e.g., 0.85), the query can be automatically rerouted to a larger, more capable model. This allows for efficient local processing of high-confidence queries while leveraging more powerful external models for uncertain cases.
The first implementation of this technology is the Gemma 4 E2B Hybrid model, available in the Cactus Hybrid collection on Hugging Face. This model, the smallest Gemma variant, demonstrates performance comparable to Gemini 3.1 Flash-Lite on most benchmarks. It achieves this by processing the majority of queries itself and routing only 15–35% of them to the Gemini 3.1 Flash-Lite for validation or more complex responses.
Developers can access and integrate the Gemma 4 E2B Hybrid model using the `cactus-compute` library or through `mlx-lm` and `transformers` with provided code examples. Cactus encourages independent benchmarking for various quantization methods like Unsloth, GGUF, and MLX to assess performance.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Cactus developed a method for small, on-device AI models to self-assess the confidence of their answers, outputting a score between 0 and 1. This allows low-confidence queries to be rerouted to larger models, combining the speed and privacy of on-device processing with the accuracy of more powerful models. The Gemma 4 E2B Hybrid model, using this technique, matches Gemini 3.1 Flash-Lite performance by routing only 15–35% of queries.