From Hugging Face Blog · 35 stories
Hybrid models outperform transformers in predicting meaning-rich tokens
Experiments revealed that hybrid models, like Olmo Hybrid, predict meaning-rich tokens better than transformers. However, on simple repetitive tokens, transformers maintain an edge, indicating differing strengths in architectural approaches.
NVIDIA NeMo AutoModel Enhances Fine-Tuning for Generative AI Models
NVIDIA launched NeMo AutoModel, enhancing fine-tuning for generative AI models by enabling higher training performance. This tool achieves up to 3.7x faster training and reduces GPU memory use by up to 32%, making it easier for developers to implement advanced models without extensive code changes.
Launch of FFASR Leaderboard to Benchmark ASR in Real-World Conditions
Treble Technologies and Hugging Face introduced the FFASR Leaderboard to evaluate Automatic Speech Recognition (ASR) models under far-field conditions. This community-driven benchmark aims to address the significant gap in performance between traditional clean-speech evaluations and real-world usage scenarios involving background noise and reverberation.
IBM's CUGA Offers Lightweight Framework for Building Agentic Apps
IBM has released CUGA, a Configurable Generalist Agent harness that simplifies the development of agentic applications by automating the orchestration and state management. With two dozen example applications provided, developers can create functional agents quickly without extensive groundwork, increasing efficiency in building machine learning applications.
Transformers.js improves browser-based AI model management with Cross-Origin Storage API
Transformers.js now integrates the proposed Cross-Origin Storage API to manage AI model resources more efficiently. This change reduces the redundant downloads of commonly used models across different web applications, addressing issues related to cache storage and data usage.
PP-OCRv6: New OCR Model on Hugging Face with 50-Language Support
PaddleOCR has launched PP-OCRv6, a new OCR model with capabilities in 50 languages and scalability from 1.5M to 34.5M parameters. The model improves text detection and recognition accuracy compared to its predecessor, PP-OCRv5, making it suitable for a variety of real-world OCR applications.
Local Models Triaged Issues in OpenClaw Repository for Free
In June 2026, local AI models were used to efficiently triage issues in the OpenClaw repository. This method allows for real-time notifications and reduces costs associated with cloud-based models, highlighting the growing importance of local AI implementation.
MosaicLeaks addresses privacy risks in deep research agents with new training method
MosaicLeaks reveals privacy vulnerabilities in deep research agents that combine private documents and web searches, leading to potential leakage of sensitive information. The proposed Privacy-Aware Deep Research (PA-DR) method improves task accuracy and decreases information leakage significantly, from 34.0% to 9.9% for full-information leakage.
New ARD Specification Enables Dynamic Agent Searches Across Tools
The Agentic Resource Discovery (ARD) specification has been developed collaboratively by major tech companies to allow agents to discover tools at runtime. This move shifts from a static model requiring pre-installed capabilities to dynamic, intent-based searches, enhancing the ability of agents to access and utilize a broader range of tools effectively.
Migrating CI from GitHub to Hugging Face Jobs for Enhanced Performance
Trackio has migrated its CI from GitHub Actions to Hugging Face Jobs, achieving a 30% reduction in CPU CI time and enabling GPU testing. This step is significant for improving efficiency and expanding testing capabilities in machine learning projects.
OpenEnv Gains Support from Major AI Organizations for Open Source Development
OpenEnv has transitioned to an open-source model coordinated by leading AI organizations such as Meta-PyTorch and Microsoft. This move aims to improve agent training efficiency across various AI harnesses and environments, fostering collaboration within the AI community.
Nemotron 3.5 Enhances Multimodal Content Safety with Custom Policies
Nemotron 3.5 introduces customizable multimodal safety integration, considering user prompts, images, and responses simultaneously. This update captures policy violations emerging from interaction, enhancing deployments across various global languages and industries.
Direct Preference Optimization Reduces Text Degeneration in OCR Models
DharmaOCR introduces Direct Preference Optimization (DPO) to combat text degeneration in OCR models. The second training stage reduced degeneration rates by an average of 59.4%, addressing a significant limitation of supervised fine-tuning.
Holo3.1 Released with Local Execution and Enhanced Performance
Holo3.1 has been released, featuring enhanced robustness for local and mobile environments, quantized checkpoints for local inference, and improved performance across various deployment frameworks. This release addresses the challenges of deployment flexibility and performance consistency in diverse operational settings.
JetBrains Launches Mellum2: 12B Mixture-of-Experts AI Model
JetBrains has released Mellum2, a 12 billion-parameter Mixture-of-Experts model optimized for natural language and coding tasks. With efficient parameter activation and over 2x faster inference compared to similar models, Mellum2 is positioned for high-throughput AI applications.
New Terminal Tools Manage AI Coding Agents and Sessions
Two new terminal-based tools, Agent-Manager and Wallfacer, have been released to help developers manage AI coding agents. Agent-Manager provides a TUI for running and monitoring multiple agents like Claude Code, Codex, OpenCode, and Grok in individual tmux sessions, while Wallfacer acts as a session manager for organizing, searching, and resuming past interactions from agents such as Claude Code, Cursor CLI, Kiro CLI, and Codex.
New Research Addresses LLM Safety Refusal for Specific Harmful Subsets of Topics
New research explores how large language models (LLMs) can be trained to refuse only specific harmful subsets of a topic, rather than the entire topic. This approach allows LLMs to maintain helpfulness while still adhering to safety policies, addressing a limitation where current topic-level safety guards are too broad.
Constraint-Aware GPU Allocator Improves Utilization by 33 Points Over FIFO Scheduling
A new constraint-aware GPU allocator demonstrated up to a 33 percentage point increase in GPU utilization and up to 105% increase in priority-weighted output compared to a FIFO scheduler on identical hardware and workloads. This improvement highlights the significant impact of allocation order on resource efficiency in GPU clusters, particularly when managing diverse workload types like training and real-time inference.
Open Model Landscape in Summer 2026 Shows Rapid Growth and Shifting Development Strategies
The open model ecosystem on Hugging Face grew significantly in 2026, with models increasing from 2.43 to 2.96 million and datasets from 711,000 to 1 million. Chinese labs are increasingly releasing large, high-performance models, often exceeding those from American labs, while hardware vendors like AMD and NVIDIA are becoming major contributors of open models to drive chip sales.
Challenges and Advances in Simulation for Physical AI Systems
The article discusses the challenges of data availability in training physical AI systems and highlights the role of simulation in overcoming these issues. Simulation enables the generation of photorealistic data at lower costs, allowing developers to enhance robot learning and performance in complex physical interactions.
Routing Systems in AI: Complexity Beyond Model Selection
Routing systems for AI agents face complexity beyond simple model selection, involving cost, performance, and compliance challenges. Caching effects and task difficulty assessments must also be factored into routing decisions for optimal efficiency.
Analysis of AI Specialization and Its Emergence as a Key Principle
A recent analysis highlights the inevitability of specialization in effective AI systems, drawing on various domains. It argues that focused AI systems outperform general models, correlating with findings in optimization theory and evolutionary biology.
Exploring Alternatives to LoRA in Parameter-Efficient Fine-Tuning
The article investigates alternatives to LoRA, the predominant technique in parameter-efficient fine-tuning (PEFT). It highlights the potential of PEFT techniques to reduce memory requirements for model fine-tuning and mentions the development of the PEFT library by Hugging Face, which supports various methods and improves accessibility.
Hugging Face Models Now One-Click Deployable to Amazon SageMaker Studio
AWS and Hugging Face have integrated deep-linking, allowing developers to move a model from Hugging Face directly into Amazon SageMaker Studio in a single click. This eliminates prior multi-step processes, enabling quicker model experimentation and deployment. The update is significant for faster AI development and deployment in enterprise environments.
OlmoEarth Studio now offers custom embedding exports for Earth observation data
OlmoEarth Studio has introduced the ability to compute and export embedding vectors from its open-source OlmoEarth foundation models. These embeddings provide compact numerical representations of Earth-observation data, enabling various downstream analytical tasks such as similarity search and segmentation.
Hugging Face Expands PyTorch Profiling Guide with MLP and Attention Techniques
Hugging Face continues its 'Profiling in PyTorch' series, detailing the integration and profiling of nn.Linear and Multilayer Perceptron (MLP) blocks, and expanding to attention mechanisms in transformer models. These insights assist developers in optimizing deep learning models using the PyTorch profiler, showcasing GPU capabilities effectively.
DiScoFormer model estimates density and score for data distributions
The DiScoFormer model estimates both the density and score of data distributions in a single forward pass. This model improves upon existing methods by allowing for high-dimensional data analysis without the need for retraining, addressing challenges in density estimation and score matching.
Hugging Face simplifies vLLM server setup with single command
Hugging Face introduced a command to run a vLLM server easily, facilitating model testing and evaluation. This command allows users to quickly deploy models and interact with them via the OpenAI API using Hugging Face infrastructure.
Hugging Face Enhances CLI and Adopts Weekly Releases for Improved Efficiency
Hugging Face has updated their command-line interface (CLI) to cater to both human and artificial intelligence (AI) agents, optimizing token usage. Additionally, they have shifted to a weekly release schedule for the huggingface_hub Python client to accelerate the implementation of fixes and features. These changes enhance CLI efficiency and streamline the release process.
Strands Robots SDK integrates LeRobot for seamless robot task management
The Strands Robots SDK now integrates LeRobot hardware and simulations, streamlining task management for robots. Users can record, test, and deploy robot tasks with fewer tools, enhancing workflow efficiency across multiple robots.
Agent Creates 3D Paris Gallery Using Hugging Face Spaces
A coding agent utilized Hugging Face Spaces to create a web gallery featuring 3D Gaussian models of Paris monuments without manually engaging with image or 3D tools. This illustrates a shift towards modular software construction where AI integrates existing components easily.
Introduction of MCP Tools for Reachy Mini Enhances Remote Functionality
The Reachy Mini now supports remote tools through MCP canary Space, allowing the addition of external functionalities like weather queries. This update enhances the robot's interactivity and potential use cases without modifying the core app directly.
GPU Utilization Becomes Key Constraint for Enterprise AI, Similar to Airline Aircraft Downtime
The efficiency of GPU utilization is emerging as a critical factor for enterprise AI success, mirroring how aircraft ground time impacts airline profitability. As AI scales, the focus shifts from model quality and raw compute power to maximizing the active use of specialized hardware to control costs and drive output.
Open-source reproduction of AI model for generating watercolor paintings with TRL and OpenEnv
A developer has openly reproduced the training pipeline for an AI model that generates watercolor paintings using JavaScript and the p5.brush library. This reproduction utilizes Hugging Face's TRL and OpenEnv, making the training scripts, RL environment, and models publicly available.
Finetuning Multi-Vector Embedding Models with Sentence Transformers
A guide details how to finetune multi-vector embedding models using Sentence Transformers, including a new `MultiVectorEncoder` for ColBERT-style late interaction retrieval. This method allows users to train models that can outperform general-purpose retrievers on specific datasets.