InstinctFlash has announced the full-source release of its platform, which includes eight robotics model families, acceleration kernels, and Python/WebSocket serving capabilities through a single Runtime. This release aims to facilitate the deployment of large world-action models in real-time robotics applications.
Benchmarks conducted on NVIDIA Jetson Thor show significant performance improvements. InstinctFlash achieved up to a 33.78x speedup with the LingBot-VA model, utilizing FP8 quantization and fewer sampling steps. This optimization was observed without a loss in task performance during real-robot tests, indicating efficient inference for complex models.
In addition to Jetson Thor, InstinctFlash now supports NVIDIA RTX 4090 and RTX 5090 GPUs. This expansion allows developers to deploy and run these 5B world-action models on high-end workstations using the same Runtime API, providing flexibility for development and deployment environments.
The platform offers a simplified deployment process. Users can point the 'serve' command at a fine-tuned checkpoint, and InstinctFlash automatically detects the model family, generates a declaration, and starts serving. For stock releases, models can be loaded using their Hub ID after installing the family's environment, streamlining the inference workflow.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
InstinctFlash released full-source code for running 5B world-action models in real-time on NVIDIA Jetson Thor, achieving up to 33.78x speedup. The release includes support for RTX 4090 and 5090 GPUs, offering a unified Runtime API for deployment across platforms. This allows for more efficient deployment of large robotics models in real-world applications.