← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

InstinctFlash Enables Real-Time 5B World-Action Models on Jetson Thor and RTX GPUs

🔄 Updated 1d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • InstinctFlash released full-source code for robotics models.
  • Achieved up to 33.78x speedup on Jetson Thor with FP8.
  • Supports RTX 4090 and 5090 GPUs with a unified API.
  • Enables real-time inference for 5B world-action models.

Full-Source Release for Robotics Models

InstinctFlash has announced the full-source release of its platform, which includes eight robotics model families, acceleration kernels, and Python/WebSocket serving capabilities through a single Runtime. This release aims to facilitate the deployment of large world-action models in real-time robotics applications.

Performance Benchmarks on Jetson Thor

Benchmarks conducted on NVIDIA Jetson Thor show significant performance improvements. InstinctFlash achieved up to a 33.78x speedup with the LingBot-VA model, utilizing FP8 quantization and fewer sampling steps. This optimization was observed without a loss in task performance during real-robot tests, indicating efficient inference for complex models.

Expanded GPU Support

In addition to Jetson Thor, InstinctFlash now supports NVIDIA RTX 4090 and RTX 5090 GPUs. This expansion allows developers to deploy and run these 5B world-action models on high-end workstations using the same Runtime API, providing flexibility for development and deployment environments.

Simplified Deployment and Inference

The platform offers a simplified deployment process. Users can point the 'serve' command at a fine-tuned checkpoint, and InstinctFlash automatically detects the model family, generates a declaration, and starts serving. For stock releases, models can be loaded using their Hub ID after installing the family's environment, streamlining the inference workflow.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~26 min · 21 stories · Sep 23

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

InstinctFlash released full-source code for running 5B world-action models in real-time on NVIDIA Jetson Thor, achieving up to 33.78x speedup. The release includes support for RTX 4090 and 5090 GPUs, offering a unified Runtime API for deployment across platforms. This allows for more efficient deployment of large robotics models in real-world applications.