← All stories
● Covered by 12 sources · 100 reportsMedium impact4 negative84 neutral6 positive

New AI Models for Long-Horizon Coding Tasks Introduced

🔄 Updated 17h ago — new reporting from Hacker News Front Page
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • GLM-5.2 offers a 1M-token context for coding tasks.
  • SWE-1.7 advances with cost-performance in reinforcement learning.
  • Xiaomi-Robotics-1 uses extensive pre-training for robotics.
  • Laguna S 2.1 excels in reasoning and coding benchmarks.
  • New models highlight improved long-horizon task handling.
  • GLM-5.2 uses IndexShare to reduce per-token FLOPs by 2.9x at 1M context.
  • GLM-5.2 improves its MTP layer, increasing acceptance length by up to 20%.
  • GLM-5.2 is released under an MIT open-source license.
  • SWE-1.7 is available in Devin via Cerebras at 1000 TPS.
  • Xiaomi-Robotics-1 uses 100,000 hours of embodiment-free pre-training data.
  • Laguna S 2.1 is a 118B parameter Mixture-of-Experts model.
  • Laguna S 2.1 has 8B activated parameters per token.
  • Laguna S 2.1 was developed and launched in under nine weeks.
  • llm-d's co-operative time-slicing increases accelerator duty cycles from 40% to 70%.
  • A frozen 12B language model achieves 100% accuracy on specific problem families.
  • The model uses a persistent memory of verified solutions.
  • The method provides deterministic, bit-exact answers.
  • The model achieves 100% accuracy with zero generation tokens.
  • The approach decouples capability from continuous model retraining and parameter scaling.
  • The method achieved 180/180 on 180 fresh instances across nine problem families.
  • Memory selection takes 1.4 microseconds.
  • LLMs trained on K-5 curriculum data do not acquire capabilities beyond that curriculum.
  • Pretraining data distribution sets an effective ceiling on a model's capabilities.
  • New skills are elicited rather than acquired through interventions.
  • An 88B-token corpus, LittleCurriculum, was filtered to U.S. elementary-school curriculum.
  • LittleCurriculum excludes concepts, facts, and vocabulary taught above Grade 5.
  • LittleLearner models were trained from scratch at 0.6B, 1.3B, and 5B scales.
  • LittleLearner models have matched Unfiltered controls for comparison.
  • Sentence Transformers now supports MultiVectorEncoder for ColBERT-style late interaction retrieval.
  • MultiVectorEncoder allows token-level matching for improved retrieval accuracy.
  • PyLate, Stanford-NLP ColBERT, and colpali-engine models are usable within Sentence Transformers API.
  • Multi-vector models keep one vector per token, scoring query against document with MaxSim operator.
  • Multi-vector models are state of the art for visual document retrieval without OCR.
  • Liquid AI released LFM2.5 Q4_0 checkpoints.
  • LFM2.5 Q4_0 checkpoints are trained using Quantization-Aware Distillation (QAD).
  • QAD recovers 97% of BF16 average accuracy lost to quantization.
  • QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of BF16 baseline performance.
  • LFM2.5 models include LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.
  • Ornith-1.5 is a new series of foundation models.
  • Ornith-1.5 extends self-scaffolding to include self-improvement capabilities.
  • Ornith-1.5 models propose new tasks and generate task-specific scaffolds.
  • Ornith-1.5 produces solution rollouts for reinforcement learning.
  • Ornith-1.5 is available in 397B MoE, 35B MoE, and 9B dense parameter scales.
  • Ornith-1.5-397B scores 86.1 on Terminal-Bench 2.1.
  • Ornith-1.5-397B scores 56.0 on DeepSWE.
  • Ornith-1.5-397B performs on par with Claude Opus 4.8.
  • Ornith-1.5-9B has a quantized mobile version.
  • Ornith-1.5-9B-Mobile can be deployed on iPhone and Android devices.
  • Inco AI released DFlash 2 for parallel drafting technology.
  • DFlash 2 increases output per verification pass by over 20%.
  • DFlash 2 adds minimal latency.
  • Inco AI released DFlash in January.
  • DFlash runs in SGLang, vLLM, TensorRT-LLM, and llama.cpp.
  • NVIDIA measured up to 15x throughput with DFlash on Blackwell GPUs.
  • Google reported 3x more tokens per second with DFlash on TPUs.
  • CoreWeave's Kimi K2.7 Code endpoint runs DFlash by default.
  • NVIDIA, Red Hat, and Modal published DFlash drafters.
  • Meta (Muse Glimmer), Poolside (Laguna), Xiaomi (MiMo-V2.5-Pro), and NVIDIA (Nemotron 3.5 Lightning) ship official drafters.
  • DFlash models have been downloaded over 3.5 million times as of August 2026.
  • DiffusionGemma generates text at 1,500 output tokens per second on an NVIDIA H100 GPU.
  • DiffusionGemma refines 256-token blocks in parallel.
  • DiffusionGemma is fine-tuned from the Gemma 4 mixture-of-experts model.
  • DiffusionGemma has 3.8B activated and 25.2B total parameters.
  • DiffusionGemma's training pipeline uses less than 10% of the starting AR model's training token budget.
  • DiffusionGemma's first training stage uses supervised fine-tuning for bidirectional denoising.
  • DiffusionGemma's second training stage combines reinforcement learning with sampler distillation.
  • DiffusionGemma generates around 20 tokens per forward pass.
  • LFM2.5-DSpark improves LLM inference throughput by up to 3.18x on GPUs.
  • LFM2.5-DSpark improves LLM inference throughput by up to 2.87x on-device.
  • LFM2.5-DSpark reduces function-calling latency by 57% for LFM2.5-2.6B.
  • LFM2.5-DSpark uses speculative decoding with a DSpark-based approach.
  • DSpark combines DFlash-style parallel backbone, a lightweight sequential head, and a Markov chain.
  • New research introduces three tests to measure benchmark optimization in speech recognition models.
  • Several high-scoring open-source ASR systems reproduced benchmark transcripts even when audio contradicted them.
  • Models relied on subtle acoustic cues indicating which benchmark they were being tested on.
  • Nvidia introduced a cross-model KV cache transfer technique.
  • The technique maps the prefilled KV cache from a source model to a target model.
  • The linear mapping process runs 2.7 to 25 times faster than recomputing the conversation.
  • The technique retains up to 98% of the target model's standalone accuracy.
  • Meta launched MTIA 300, its first in-house training and inference accelerator.
  • MTIA 300 is optimized for training ranking and recommendation models.
  • MTIA 300 integrates network interface controllers (NICs) directly into the package.
  • Meta developed MetaRoCE, a new RDMA transport protocol for AI workloads on Ethernet.
  • Meta open-sourced MetaRoCE specification, reference implementation, and compliance test suite through OCP.
  • LLMs could exploit inference engine vulnerabilities to control host machines.
  • A past arbitrary code execution vulnerability was found in vLLM.
  • Quantization-Aware Healing (QAH) is a new method for recovering compressed and quantized LLMs.
  • QAH applied to a GPT-OSS 120B model produced a 4-bit version that surpassed its bfloat16 counterpart on 7 of 9 benchmarks.
  • IBM released Granite 4.2 LLMs.
  • Granite 4.2 models are dense, decoder-only.
  • Granite 4.2 models come in 3B, 8B, and 30B parameter sizes.
  • Granite 4.2 models use a five-phase training strategy.
  • Granite 4.2 models are pre-trained on 15T tokens.
  • Granite 4.2 models extend the context window to 512K tokens.
  • Granite 4.2 models are fine-tuned on chain-of-thought, reasoning, and agentic-trajectory data.
  • Granite 4.2 models use a multi-stage reinforcement learning pipeline.
  • Granite 4.2 8B and 30B models learn with agentic RL in sandboxed environments.
  • Granite 4.2 models have a thinking/non-thinking switch.
  • Granite 4.2 models have native tool calling.
  • Granite 4.2 models are released under the Apache 2.0 license.
  • Google and Anyscale introduced an experimental library for Ray.
  • The library integrates gVisor sandboxing into distributed Ray clusters.
  • The integration enables secure execution for agentic AI workloads.
  • Ray users can manage sandboxed environments using existing Ray APIs.
  • Granite 4.2 models can run in a low-effort mode for easy questions.
  • Granite 4.2 models are open-weight and designed for self-hosting.
  • Granite 4.2 models have a 128,000-token context window.
  • Granite 4.2 8B and 30B models are trained for external tool use.
  • Granite 4.2 3B model supports tools without specialized training.
  • FreeToken is an open-source inference engine for MoE models on consumer hardware.
  • FreeToken was introduced by researchers from UC Berkeley and MIT.
  • FreeToken addresses high bandwidth requirements for MoE models.
  • FreeToken uses a dynamic co-scheduling formulation called the q* policy.
  • vLLM v0.28.0 has 584 commits from 270 contributors.
  • vLLM v0.28.0 introduces Decode Context Parallel (DCP) support for Kimi-K3.
  • vLLM v0.28.0 includes fused FlashKDA decode and prefill kernels for Kimi-K3.
  • vLLM v0.28.0 adds SiTU activation support for MegaMoE.
  • vLLM v0.28.0 features GEMM-RS for sequence parallelism.
  • vLLM v0.28.0 combines all-gathers with 1.5~3x kernel-level speedup.
  • vLLM v0.28.0's adaptive speculative token budget improves DSpark TTFT by 60%.
  • vLLM v0.28.0 offers optional shared-expert sharding, saving 17 GiB of memory per GPU.
  • Kimi-K3 now runs on ROCm with the V2 model runner.
  • vLLM v0.28.0 enables sparse MLA for DeepSeek V4 across decode, MTP, and DSpark speculative decoding.
  • vLLM v0.28.0 adds AMD Quark NVFP4 support for DeepSeek V4.
  • vLLM v0.28.0 includes reasoning-effort prompts and mappings for DeepSeek V4.
  • vLLM v0.28.0 optimizes sparse top-k metadata kernels for DeepSeek V4.
  • vLLM v0.28.0 narrows eager CUDA graph regions for DeepSeek V4.
  • vLLM v0.28.0 enables ROCm on gfx11 and gfx950 for DeepSeek V4.
  • vLLM v0.28.0 includes DFlash2 with local convolution and a candidate selector.
  • vLLM v0.28.0 features DSpark confidence-scheduled verification.
  • Artificial Analysis released benchmarks for small AI models on mobile phones.
  • Small models are defined as fitting within 8 GB of memory after quantization.
  • Benchmarks cover model intelligence and inference performance on mobile devices.
  • Liquid AI partnered to gather real inference data on devices.
  • Benchmarks were conducted on an iPhone 17 Pro.
  • Continuous diffusion models for language are experiencing a resurgence in research activity.
  • Diffusion language models generate text by iteratively refining an entire sequence.
  • Diffusion language models contrast with traditional autoregressive models that generate tokens one at a time.
  • The article adapts material from ICLR 2026 and MLSS 2026 workshop talks and lectures.
  • WebLLM is an in-browser LLM inference engine.
  • WebLLM uses WebGPU for hardware acceleration.
  • WebLLM is compatible with OpenAI API.
  • WebLLM is a companion project of MLC LLM.
  • A guide details fine-tuning a 350M language model for structured output in 100 GRPO steps.
  • The fine-tuning notebook is sized for a free-tier Colab or Kaggle GPU.
  • Evaluation can run locally on a MacBook Pro with an Apple M5 Max and 36 GB of unified memory.
  • The evaluation uses llama.cpp for serving and uv for Python tooling.
  • NeoMME is a new family of multilingual multimodal encoders.
  • NeoMME models come in 260M and 800M parameter sizes.
  • NeoMME uses a single bidirectional Transformer for text tokens and raw image patches.
  • NeoMME is trained from scratch with a masked discrete-diffusion objective.
  • NeoMME-Retriever returns dense and late-interaction embeddings in one forward pass.
  • The 260M NeoMME model encodes 51 pages per second on an NVIDIA L40S GPU.
  • NeoMME's hierarchical token pooling and asymmetric quantization reduce late-interaction index storage by 255x.
  • NeoMME is available in Hugging Face Transformers.
  • NeoMME model checkpoints are released under the Apache 2.0 license.
  • IFM released K2 Horizon, a collection of six open models.
  • K2 Horizon models range from 0.9B to 375B parameters.
  • K2 Horizon includes 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B models.
  • K2 Horizon 0.9B, 3.7B, and 7B models set new state of the art in their size classes.
  • K2 Horizon provides intermediate checkpoints, data recipes, architecture, and training code.
  • K2 Horizon models and code are released under the Apache 2.0 license.
  • GPU inference cold start times on Amazon EKS Auto Mode reduced from 8 minutes to under 1 minute.
  • CUDA kernel recompilation accounts for 65% of startup time for a 64 GB model.
  • Downloading weights from S3 accounts for 92% of startup time for a 203 GB model.
  • Time to first token served (TTFTS) is defined as wall-clock duration from pod creation to first inference response.
  • Cerebras offers Qwen 3.8 27B model.
  • Qwen 3.8 27B achieves 1500 tokens per second on Cerebras.
  • Cerebras models use selective weight-only quantization for storage.
  • Cerebras stores sensitive layers at full precision with dequantization on the fly.
  • Gemma 3 12B and 27B models were benchmarked on Google Cloud TPU v6e.
  • Gemma 3 27B hits a performance wall past 64 concurrent users for generation tasks.
  • Gemma 3 27B plateaus at a 4.12x normalized throughput multiplier at 128 users.
  • Gemma 3 12B scales up to an 8.19x multiplier for generation tasks.
  • New research proposes a viral analogy to understand LLM diffusion.
  • The model suggests LLM adoption could lead to 'runaway dynamics' and 'abrupt losses in cognitive competence'.
  • The framework identifies conditions for 'cognitive immunization'.
  • Speculative decoding allows vLLM to verify multiple drafted tokens in a single target-model pass.
  • SageMaker AI benchmarked G7 instances, powered by NVIDIA Blackwell GPUs, against G5 and G6 instances.
  • G7 instances showed gains in throughput, latency, and cost-per-token for LLM inference.
  • OUI-1 is a finetuned DiffusionGemma model that generates user interfaces in openui-lang.
  • OUI-1 is a 26BA4B model that runs on consumer-grade GPUs (RTX 5090, at FP8).
  • OUI-1 weights are available on Hugging Face under the Gemma Terms of Use.
  • OpenUI Lang uses up to 67% fewer tokens than JSON and streams.
  • Pathway developed its BDH (Dragon Hatchling) AI architecture using Amazon SageMaker HyperPod.
  • BDH performs reasoning in latent space, not by generating extra tokens sequentially.
  • BDH learns from examples and refines solutions without generating intermediate text traces.
  • BDH is a brain-inspired architecture, originally a graph of neurons with sparse, local interactions.
  • BDH model states adapt in context without test-time weight updates.
  • BDH's reasoning horizon is not limited by a fixed-size context window.
  • Qwen 3.8 increased alignment with GPT-5.5 Pro's reasoning by 18.18 points.
  • The evaluation used 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles.
  • Kimi K3 has 31.11% overlap with GPT-5.5 Pro without prefill and 35.65% with prefill.
  • Cognition released SWE-2, a new coding model.
  • SWE-2 achieved 50.0% on the FrontierCode 1.1 Main1 benchmark.
  • SWE-2 offers a 64% cost reduction compared to Fable 5.1.
  • SWE-2 scales reinforcement learning to the multi-trillion-parameter regime.
  • SWE-2 is post-trained from Kimi K33, a 2.8T-parameter model.
  • SWE-2 scored 73.0 on DeepSWE 1.1.
  • SWE-2 scored 92.8 on Terminal-Bench 2.1.
  • SWE-2 scored 27.3 on Terminal-Bench 4.0.
  • SWE-2 is available in Devin Desktop and CLI.
  • SWE-2 uses MoE inference on NVFP4 and FP8 kernels with quantization-aware training.
  • SWE-2's draft model, retrained with SpecForge, gives 15% longer accept lengths.
  • SWE-2's prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20%.
  • SWE-2 has a mean of 53 steps per run for medium effort, 80 for high, and 98 for max.
  • SWE-2's medium effort posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost.
  • SWE-2 lands its first real edit at a median of step 18.
  • A new configurable, instruction-driven PII detector built on LLMs has been developed.
  • The PII detector runs on any LLM managed on Amazon Bedrock.
  • The PII detector was evaluated on five public PII corpora across nine LLM-based detectors.
  • The PII detector ships as the pii-detector package.
  • BDH aims to address LLM inefficiencies like forgetting during long interactions.
  • BDH updates internal memory during inference.
  • BDH works through problems without verbalized reasoning traces.
  • Pathway uses Amazon SageMaker HyperPod to scale its training.
  • BDH moves beyond the transformer paradigm.
  • Amazon SageMaker Inference now includes prefix-aware routing.
  • Prefix-aware routing directs requests with identical prompt prefixes to the same instance.
  • LinkedIn uses multi-teacher distillation to compress large models into a 0.6B-parameter ranking model.
  • LinkedIn's system uses SGLang to serve teacher models directly in the training loop.
  • LinkedIn's system handles tensor-parallel and data-parallel setups for teacher models.
  • Recurrent Looped Transformer (RLT) integrates a causal encoder with a recurrent decoder.
  • RLT's temporal processing depth increases with sequence length.
  • Each token extends the recurrent path through the full decoder in RLT.
  • RLT's decoder carries its final hidden state and layerwise sliding-window attention (SWA) cache.
  • RLT's encoder constructs global key-value memory.
  • OpenArch is a GitHub repository with PyTorch LLM implementations.
  • OpenArch focuses on clarity and readability, not production optimization.
  • OpenArch implements models from Sebastian Raschka's LLM Architecture Gallery.
  • AI industry focus shifted from training to optimizing inference.
  • The shift to inference is driven by LLMs performing multiple inference steps per query and continuous agentic AI operation.
  • Pinterest's Manas platform uses Scalar Quantization and Product Quantization to reduce memory usage.
  • Scalar Quantization reduces HNSW indices by 59% and IVF indices by 75% while maintaining over 90% recall.
  • Product Quantization reduces HNSW indices by 74% and IVF indices by 93% with 70–80% recall.
  • Typed Domain Grounding (TDG) embeds DSLs as typed internal DSLs within a host language.
  • TDG achieved higher Structural Fidelity and lower hallucination rates than external DSLs with Claude Sonnet 5 and GPT-4o.
  • Infinite-Parameter LLMs generate and adapt model weights from live interaction data using a compact hypernetwork.
  • Cache-to-Cache (C2C) allows LLMs to communicate directly via KV-caches, bypassing text generation.
  • C2C improves response quality by 3.1-5.4% and offers a 2.5x speedup over text-based communication.
  • C2C uses a neural network to project and fuse source and target model KV-caches.
  • C2C includes a learnable gating mechanism to select target layers for cache communication.
  • Linguistic illegibility describes when an LLM's language outputs do not accurately reflect its internal computations.
  • Security mechanisms relying on an LLM's linguistic self-reporting are unreliable due to linguistic illegibility.
  • A researcher argues chat-based LLMs create an illusion of intelligence through statistically generic responses.
  • LLMs are mathematical models of language tokens, not brains.
  • ByteDance Seed and Tsinghua Air released DAPO, an open-source RL system.
  • DAPO includes algorithm, code infrastructure, and dataset.
  • DAPO achieved 50% on AIME 2024 using the Qwen2.5-32B model.
  • DAPO outperforms DeepSeek-R1-Zero-Qwen-32B with 50% fewer training steps.
  • DAPO's early version achieved 44% on AIME 2024.
  • DAPO proposes the Decoupled Clip and Dynamic sAmpling Policy Optimization algorithm.
  • Mini-AGI is a byte-level language model.
  • Mini-AGI dynamically manages its architecture.
  • Mini-AGI pages weights from disk to VRAM.
  • Mini-AGI's parameter count is bounded by free disk space.
  • Mini-AGI grows new capacity when needed and prunes unused capacity.
  • Tokenizers v1 has been released.
  • Tokenizers v1 offers speed improvements of tens of times compared to v0.23.
  • Tokenizers v1 maintains API and output compatibility with v0.23.
  • Hugging Face Transformers library now supports llama.cpp's GGUF quantized models.
  • llama.cpp's inference engine powers local AI tools like Ollama, LM Studio, and Jan.
  • AlphaEvolve optimizes video processing by combining cloud code generation with local hardware execution.
  • AlphaEvolve autonomously tunes code for real-time streaming applications.
  • AlphaEvolve was used by DoIt to optimize production Swift code in a macOS streaming app.
  • Contrastive Language Models (CLM) were developed by Stanford University and NVIDIA Research.
  • CLM performs on par with Jev while running up to nine times faster.
  • CLM uses a contrastive objective to train state and action encoders.
  • CLM enables faster decision-making in tasks with many candidate actions or frequently reused actions.
  • CLM-8B performs on par with Jev.
  • CLM speedups are most pronounced with large numbers of candidate actions or frequently reused actions.
  • LFM2.5-VL-DSpark is a new vision drafter.
  • LFM2.5-VL-DSpark speeds up VLM inference by up to 3.13x on devices and 2.66x on H100 GPUs.
  • LFM2.5-VL-DSpark adds 280M parameters, 8.9% to the 3B target model.
  • LFM2.5-VL-DSpark has day-one support for llama.cpp, MLX-VLM, and SGLang.
  • LFM2.5-VL-DSpark uses a 4-layer attention-only drafter with a block size of 9.
  • Amazon EKS, EFA, and DeepEP scale MoE reinforcement learning.
  • MoE reinforcement learning on AWS achieves 40% more throughput.
  • MoE models require pre-training, mid-training, SFT, and RL stages.
  • DeepSeek Elastic Compute (DSec) is a production sandbox platform for agentic training and evaluation.
  • DSec supports FnCall, container, microVM, and full-VM sandbox backends.
  • DSec coordinates placement and lifecycle management across the cluster.
  • DSec composes environments from independently versioned layers.
  • DSec combines memory sharing, reclamation, and CPU scheduling for high-density execution.
  • DSec loads image data on demand from Fire-Flyer File Sys.
  • Chat templates influence LLM self-referential disclaimers and experiential statements.
  • The presence of a chat template increases disclaimers and decreases experiential voice across 8 open-source instruct models up to 9B parameters.
  • A specific direction within model activations steers this self-referential behavior.
  • llama.cpp's prompt lookup decoding speed increased by up to 140 times.
  • llama.cpp's prompt lookup decoding memory usage reduced by up to 2.6 times.
  • Daniel Lemire contributed a PR that made prompt lookup drafting 4.2x faster.
  • A 0.5B BitNet LLM runs on a cluster of seven ESP32S3 microcontrollers.
  • One master ESP32S3 handles tokenization and embedding.
  • Six compute ESP32S3 nodes run transformer layers.
  • The ESP32S3 nodes communicate via a high-speed SPI daisy-chain.
  • PSSA is a new language model written from scratch in Rust.
  • PSSA uses a recurrent state-space layer and episodic memory instead of a transformer architecture.
  • PSSA learns faster and generates text 12 times quicker on a CPU compared to transformers with matched parameters and corpus.
  • PSSA's architecture offers linear cost growth with sequence length.
  • PSSA needed per-token weight updates, a memory bank written during the forward pass, and a scalar reference path.
  • Magnitude is an open-source inference engine for AI agents.
  • Magnitude optimizes kernels for specific hardware.
  • Magnitude offers up to 2x faster performance than llama.cpp.
  • Magnitude reduces memory usage by 27% per agent.
  • Magnitude works on Apple Silicon, NVIDIA, AMD, and CPU.
  • Magnitude has a desktop app that includes a CLI.
  • Magnitude is 92% faster decode on Metal.
  • Magnitude is 19% faster decode on CUDA.
  • Magnitude tunes kernels on the device before a model runs.
  • Magnitude has hand-optimized kernels for popular open-weight families.
  • Magnitude frees memory when agents stop.
  • Magnitude's concurrent sessions share prefix caches.
  • Magnitude connects with agents like Pi, OpenCode, Hermes, and Codex.
  • Olmo-core 3 is a redesigned open mixture-of-experts (MoE) training system.
  • Olmo-core 3 is designed to scale MoE training into the trillion-parameter range.
  • Context Language Models (CLMs) manage their own context by treating it as an editable file.
  • CLMs allow models to make unrestricted updates to their context file.
  • CLMs extend to multi-agent systems where multiple agent contexts coexist as files.
  • CLMs achieve 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus.
  • CLMs achieve 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench.
  • CLMs show 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task.
  • CLMs enable in-context and parametric learning of context-management strategies.
  • CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop.
  • CLMs improve held-out accuracy by up to 35.9%.

Introduction of New AI Models

Hugging Face, Cognition, and Xiaomi have launched new AI models enhancing long-horizon coding and robotics capabilities. These models seek to offer improved performance by leveraging extensive contexts and innovative techniques.

GLM-5.2 by Hugging Face

GLM-5.2 extends support for coding-agent scenarios with a robust 1 million token context. Its architecture allows for efficient handling of complex coding tasks. The model is open-source, available under an MIT license.

SWE-1.7 Advances Reinforcement Learning

Cognition's SWE-1.7 model advances long-horizon asynchronous tasks by applying improved reinforcement learning methods. It aims to enhance cost-performance efficiency for software engineering tasks.

Robotics improved by Xiaomi's Pre-training

Xiaomi-Robotics-1 leverages 100,000 hours of pre-training data to address robotics' data scarcity. It combines pre-training with real-robot data to improve model capability, providing insights into large-scale training effects.

Implications for AI Development

These models underline ongoing advancements in high-reasoning AI necessary for complex problem-solving. By achieving breakthroughs in long-horizon tasks, these developments suggest a shift towards more sophisticated AI solutions.

Updates

🕒 2026-10-01 · new reporting from Hacker News Front Page
  • Context Language Models (CLMs) manage their own context by treating it as an editable file.
  • CLMs allow models to make unrestricted updates to their context file.
  • CLMs extend to multi-agent systems where multiple agent contexts coexist as files.
  • CLMs achieve 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus.
  • CLMs achieve 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench.
  • CLMs show 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task.
  • CLMs enable in-context and parametric learning of context-management strategies.
  • CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop.
  • CLMs improve held-out accuracy by up to 35.9%.
🕒 2026-10-01 · new reporting from Hugging Face Blog
  • Olmo-core 3 is a redesigned open mixture-of-experts (MoE) training system.
  • Olmo-core 3 is designed to scale MoE training into the trillion-parameter range.
🕒 2026-09-30 · new reporting from Hacker News Front Page
  • Magnitude is an open-source inference engine for AI agents.
  • Magnitude optimizes kernels for specific hardware.
  • Magnitude offers up to 2x faster performance than llama.cpp.
  • Magnitude reduces memory usage by 27% per agent.
  • Magnitude works on Apple Silicon, NVIDIA, AMD, and CPU.
  • Magnitude has a desktop app that includes a CLI.
  • Magnitude is 92% faster decode on Metal.
  • Magnitude is 19% faster decode on CUDA.
  • Magnitude tunes kernels on the device before a model runs.
  • Magnitude has hand-optimized kernels for popular open-weight families.
  • Magnitude frees memory when agents stop.
  • Magnitude's concurrent sessions share prefix caches.
  • Magnitude connects with agents like Pi, OpenCode, Hermes, and Codex.
🕒 2026-09-30 · new reporting from Hacker News Front Page
  • PSSA is a new language model written from scratch in Rust.
  • PSSA uses a recurrent state-space layer and episodic memory instead of a transformer architecture.
  • PSSA learns faster and generates text 12 times quicker on a CPU compared to transformers with matched parameters and corpus.
  • PSSA's architecture offers linear cost growth with sequence length.
  • PSSA needed per-token weight updates, a memory bank written during the forward pass, and a scalar reference path.
🕒 2026-09-29 · new reporting from Hacker News Front Page
  • A 0.5B BitNet LLM runs on a cluster of seven ESP32S3 microcontrollers.
  • One master ESP32S3 handles tokenization and embedding.
  • Six compute ESP32S3 nodes run transformer layers.
  • The ESP32S3 nodes communicate via a high-speed SPI daisy-chain.
🕒 2026-09-27 · new reporting from Hacker News Front Page
  • llama.cpp's prompt lookup decoding speed increased by up to 140 times.
  • llama.cpp's prompt lookup decoding memory usage reduced by up to 2.6 times.
  • Daniel Lemire contributed a PR that made prompt lookup drafting 4.2x faster.
🕒 2026-09-27 · new reporting from Hacker News Front Page
  • Chat templates influence LLM self-referential disclaimers and experiential statements.
  • The presence of a chat template increases disclaimers and decreases experiential voice across 8 open-source instruct models up to 9B parameters.
  • A specific direction within model activations steers this self-referential behavior.
🕒 2026-09-26 · new reporting from Hacker News Front Page
  • DeepSeek Elastic Compute (DSec) is a production sandbox platform for agentic training and evaluation.
  • DSec supports FnCall, container, microVM, and full-VM sandbox backends.
  • DSec coordinates placement and lifecycle management across the cluster.
  • DSec composes environments from independently versioned layers.
  • DSec combines memory sharing, reclamation, and CPU scheduling for high-density execution.
  • DSec loads image data on demand from Fire-Flyer File Sys.
🕒 2026-09-25 · new reporting from AWS Machine Learning Blog
  • Amazon EKS, EFA, and DeepEP scale MoE reinforcement learning.
  • MoE reinforcement learning on AWS achieves 40% more throughput.
  • MoE models require pre-training, mid-training, SFT, and RL stages.
🕒 2026-09-24 · new reporting from Hugging Face Blog
  • LFM2.5-VL-DSpark is a new vision drafter.
  • LFM2.5-VL-DSpark speeds up VLM inference by up to 3.13x on devices and 2.66x on H100 GPUs.
  • LFM2.5-VL-DSpark adds 280M parameters, 8.9% to the 3B target model.
  • LFM2.5-VL-DSpark has day-one support for llama.cpp, MLX-VLM, and SGLang.
  • LFM2.5-VL-DSpark uses a 4-layer attention-only drafter with a block size of 9.
🕒 2026-09-24 · new reporting from Hacker News Front Page
  • Contrastive Language Models (CLM) were developed by Stanford University and NVIDIA Research.
  • CLM performs on par with Jev while running up to nine times faster.
  • CLM uses a contrastive objective to train state and action encoders.
  • CLM enables faster decision-making in tasks with many candidate actions or frequently reused actions.
  • CLM-8B performs on par with Jev.
  • CLM speedups are most pronounced with large numbers of candidate actions or frequently reused actions.
🕒 2026-09-23 · new reporting from Google Cloud Blog
  • AlphaEvolve optimizes video processing by combining cloud code generation with local hardware execution.
  • AlphaEvolve autonomously tunes code for real-time streaming applications.
  • AlphaEvolve was used by DoIt to optimize production Swift code in a macOS streaming app.
🕒 2026-09-22 · new reporting from Hugging Face Blog
  • Hugging Face Transformers library now supports llama.cpp's GGUF quantized models.
  • llama.cpp's inference engine powers local AI tools like Ollama, LM Studio, and Jan.
🕒 2026-09-21 · new reporting from Hugging Face Blog
  • Tokenizers v1 has been released.
  • Tokenizers v1 offers speed improvements of tens of times compared to v0.23.
  • Tokenizers v1 maintains API and output compatibility with v0.23.
🕒 2026-09-21 · new reporting from Hacker News Front Page
  • Mini-AGI is a byte-level language model.
  • Mini-AGI dynamically manages its architecture.
  • Mini-AGI pages weights from disk to VRAM.
  • Mini-AGI's parameter count is bounded by free disk space.
  • Mini-AGI grows new capacity when needed and prunes unused capacity.
🕒 2026-09-21 · new reporting from Hacker News Front Page
  • ByteDance Seed and Tsinghua Air released DAPO, an open-source RL system.
  • DAPO includes algorithm, code infrastructure, and dataset.
  • DAPO achieved 50% on AIME 2024 using the Qwen2.5-32B model.
  • DAPO outperforms DeepSeek-R1-Zero-Qwen-32B with 50% fewer training steps.
  • DAPO's early version achieved 44% on AIME 2024.
  • DAPO proposes the Decoupled Clip and Dynamic sAmpling Policy Optimization algorithm.
🕒 2026-09-20 · new reporting from Hacker News Front Page
  • A researcher argues chat-based LLMs create an illusion of intelligence through statistically generic responses.
  • LLMs are mathematical models of language tokens, not brains.
🕒 2026-09-18 · new reporting from Hacker News Front Page
  • Cache-to-Cache (C2C) allows LLMs to communicate directly via KV-caches, bypassing text generation.
  • C2C improves response quality by 3.1-5.4% and offers a 2.5x speedup over text-based communication.
  • C2C uses a neural network to project and fuse source and target model KV-caches.
  • C2C includes a learnable gating mechanism to select target layers for cache communication.
  • Linguistic illegibility describes when an LLM's language outputs do not accurately reflect its internal computations.
  • Security mechanisms relying on an LLM's linguistic self-reporting are unreliable due to linguistic illegibility.
🕒 2026-09-18 · new reporting from IEEE Spectrum, Google Cloud Blog, InfoQ, Hacker News Front Page
  • AI industry focus shifted from training to optimizing inference.
  • The shift to inference is driven by LLMs performing multiple inference steps per query and continuous agentic AI operation.
  • Pinterest's Manas platform uses Scalar Quantization and Product Quantization to reduce memory usage.
  • Scalar Quantization reduces HNSW indices by 59% and IVF indices by 75% while maintaining over 90% recall.
  • Product Quantization reduces HNSW indices by 74% and IVF indices by 93% with 70–80% recall.
  • Typed Domain Grounding (TDG) embeds DSLs as typed internal DSLs within a host language.
  • TDG achieved higher Structural Fidelity and lower hallucination rates than external DSLs with Claude Sonnet 5 and GPT-4o.
  • Infinite-Parameter LLMs generate and adapt model weights from live interaction data using a compact hypernetwork.
🕒 2026-09-14 · new reporting from Hacker News Front Page
  • OpenArch is a GitHub repository with PyTorch LLM implementations.
  • OpenArch focuses on clarity and readability, not production optimization.
  • OpenArch implements models from Sebastian Raschka's LLM Architecture Gallery.
🕒 2026-09-13 · new reporting from Hacker News Front Page
  • Recurrent Looped Transformer (RLT) integrates a causal encoder with a recurrent decoder.
  • RLT's temporal processing depth increases with sequence length.
  • Each token extends the recurrent path through the full decoder in RLT.
  • RLT's decoder carries its final hidden state and layerwise sliding-window attention (SWA) cache.
  • RLT's encoder constructs global key-value memory.
🕒 2026-09-11 · new reporting from InfoQ
  • LinkedIn uses multi-teacher distillation to compress large models into a 0.6B-parameter ranking model.
  • LinkedIn's system uses SGLang to serve teacher models directly in the training loop.
  • LinkedIn's system handles tensor-parallel and data-parallel setups for teacher models.
🕒 2026-09-11 · new reporting from AWS Machine Learning Blog
  • Amazon SageMaker Inference now includes prefix-aware routing.
  • Prefix-aware routing directs requests with identical prompt prefixes to the same instance.
🕒 2026-09-10 · new reporting from AWS Machine Learning Blog
  • BDH aims to address LLM inefficiencies like forgetting during long interactions.
  • BDH updates internal memory during inference.
  • BDH works through problems without verbalized reasoning traces.
  • Pathway uses Amazon SageMaker HyperPod to scale its training.
  • BDH moves beyond the transformer paradigm.
🕒 2026-09-10 · new reporting from Hacker News Front Page, AWS Machine Learning Blog
  • Cognition released SWE-2, a new coding model.
  • SWE-2 achieved 50.0% on the FrontierCode 1.1 Main1 benchmark.
  • SWE-2 offers a 64% cost reduction compared to Fable 5.1.
  • SWE-2 scales reinforcement learning to the multi-trillion-parameter regime.
  • SWE-2 is post-trained from Kimi K33, a 2.8T-parameter model.
  • SWE-2 scored 73.0 on DeepSWE 1.1.
  • SWE-2 scored 92.8 on Terminal-Bench 2.1.
  • SWE-2 scored 27.3 on Terminal-Bench 4.0.
  • SWE-2 is available in Devin Desktop and CLI.
  • SWE-2 uses MoE inference on NVFP4 and FP8 kernels with quantization-aware training.
  • SWE-2's draft model, retrained with SpecForge, gives 15% longer accept lengths.
  • SWE-2's prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20%.
  • SWE-2 has a mean of 53 steps per run for medium effort, 80 for high, and 98 for max.
  • SWE-2's medium effort posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost.
  • SWE-2 lands its first real edit at a median of step 18.
  • A new configurable, instruction-driven PII detector built on LLMs has been developed.
  • The PII detector runs on any LLM managed on Amazon Bedrock.
  • The PII detector was evaluated on five public PII corpora across nine LLM-based detectors.
  • The PII detector ships as the pii-detector package.
🕒 2026-09-09 · new reporting from Hacker News Front Page
  • Qwen 3.8 increased alignment with GPT-5.5 Pro's reasoning by 18.18 points.
  • The evaluation used 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles.
  • Kimi K3 has 31.11% overlap with GPT-5.5 Pro without prefill and 35.65% with prefill.
🕒 2026-09-08 · new reporting from AWS Machine Learning Blog
  • Pathway developed its BDH (Dragon Hatchling) AI architecture using Amazon SageMaker HyperPod.
  • BDH performs reasoning in latent space, not by generating extra tokens sequentially.
  • BDH learns from examples and refines solutions without generating intermediate text traces.
  • BDH is a brain-inspired architecture, originally a graph of neurons with sparse, local interactions.
  • BDH model states adapt in context without test-time weight updates.
  • BDH's reasoning horizon is not limited by a fixed-size context window.
🕒 2026-09-08 · new reporting from AWS Machine Learning Blog, Hacker News Front Page
  • SageMaker AI benchmarked G7 instances, powered by NVIDIA Blackwell GPUs, against G5 and G6 instances.
  • G7 instances showed gains in throughput, latency, and cost-per-token for LLM inference.
  • OUI-1 is a finetuned DiffusionGemma model that generates user interfaces in openui-lang.
  • OUI-1 is a 26BA4B model that runs on consumer-grade GPUs (RTX 5090, at FP8).
  • OUI-1 weights are available on Hugging Face under the Gemma Terms of Use.
  • OpenUI Lang uses up to 67% fewer tokens than JSON and streams.
🕒 2026-09-07 · new reporting from Hacker News Front Page
  • Speculative decoding allows vLLM to verify multiple drafted tokens in a single target-model pass.
🕒 2026-09-05 · new reporting from Hacker News Front Page
  • New research proposes a viral analogy to understand LLM diffusion.
  • The model suggests LLM adoption could lead to 'runaway dynamics' and 'abrupt losses in cognitive competence'.
  • The framework identifies conditions for 'cognitive immunization'.
🕒 2026-09-04 · new reporting from Google Cloud Blog
  • Gemma 3 12B and 27B models were benchmarked on Google Cloud TPU v6e.
  • Gemma 3 27B hits a performance wall past 64 concurrent users for generation tasks.
  • Gemma 3 27B plateaus at a 4.12x normalized throughput multiplier at 128 users.
  • Gemma 3 12B scales up to an 8.19x multiplier for generation tasks.
🕒 2026-09-03 · new reporting from Hacker News Front Page
  • Cerebras offers Qwen 3.8 27B model.
  • Qwen 3.8 27B achieves 1500 tokens per second on Cerebras.
  • Cerebras models use selective weight-only quantization for storage.
  • Cerebras stores sensitive layers at full precision with dequantization on the fly.
🕒 2026-09-03 · new reporting from Hacker News Front Page, The New Stack
  • IFM released K2 Horizon, a collection of six open models.
  • K2 Horizon models range from 0.9B to 375B parameters.
  • K2 Horizon includes 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B models.
  • K2 Horizon 0.9B, 3.7B, and 7B models set new state of the art in their size classes.
  • K2 Horizon provides intermediate checkpoints, data recipes, architecture, and training code.
  • K2 Horizon models and code are released under the Apache 2.0 license.
  • GPU inference cold start times on Amazon EKS Auto Mode reduced from 8 minutes to under 1 minute.
  • CUDA kernel recompilation accounts for 65% of startup time for a 64 GB model.
  • Downloading weights from S3 accounts for 92% of startup time for a 203 GB model.
  • Time to first token served (TTFTS) is defined as wall-clock duration from pod creation to first inference response.
🕒 2026-09-03 · new reporting from Hugging Face Blog
  • NeoMME is a new family of multilingual multimodal encoders.
  • NeoMME models come in 260M and 800M parameter sizes.
  • NeoMME uses a single bidirectional Transformer for text tokens and raw image patches.
  • NeoMME is trained from scratch with a masked discrete-diffusion objective.
  • NeoMME-Retriever returns dense and late-interaction embeddings in one forward pass.
  • The 260M NeoMME model encodes 51 pages per second on an NVIDIA L40S GPU.
  • NeoMME's hierarchical token pooling and asymmetric quantization reduce late-interaction index storage by 255x.
  • NeoMME is available in Hugging Face Transformers.
  • NeoMME model checkpoints are released under the Apache 2.0 license.
🕒 2026-09-03 · new reporting from Hugging Face Blog
  • A guide details fine-tuning a 350M language model for structured output in 100 GRPO steps.
  • The fine-tuning notebook is sized for a free-tier Colab or Kaggle GPU.
  • Evaluation can run locally on a MacBook Pro with an Apple M5 Max and 36 GB of unified memory.
  • The evaluation uses llama.cpp for serving and uv for Python tooling.
🕒 2026-09-02 · new reporting from Hacker News Front Page
  • WebLLM is an in-browser LLM inference engine.
  • WebLLM uses WebGPU for hardware acceleration.
  • WebLLM is compatible with OpenAI API.
  • WebLLM is a companion project of MLC LLM.
🕒 2026-08-31 · new reporting from Hacker News Front Page
  • Diffusion language models generate text by iteratively refining an entire sequence.
  • Diffusion language models contrast with traditional autoregressive models that generate tokens one at a time.
  • The article adapts material from ICLR 2026 and MLSS 2026 workshop talks and lectures.
🕒 2026-08-31 · new reporting from Hacker News Front Page
  • Continuous diffusion models for language are experiencing a resurgence in research activity.
🕒 2026-08-30 · new reporting from Hacker News Front Page
  • Artificial Analysis released benchmarks for small AI models on mobile phones.
  • Small models are defined as fitting within 8 GB of memory after quantization.
  • Benchmarks cover model intelligence and inference performance on mobile devices.
  • Liquid AI partnered to gather real inference data on devices.
  • Benchmarks were conducted on an iPhone 17 Pro.
🕒 2026-08-29 · new reporting from Hacker News Front Page
  • vLLM v0.28.0 has 584 commits from 270 contributors.
  • vLLM v0.28.0 introduces Decode Context Parallel (DCP) support for Kimi-K3.
  • vLLM v0.28.0 includes fused FlashKDA decode and prefill kernels for Kimi-K3.
  • vLLM v0.28.0 adds SiTU activation support for MegaMoE.
  • vLLM v0.28.0 features GEMM-RS for sequence parallelism.
  • vLLM v0.28.0 combines all-gathers with 1.5~3x kernel-level speedup.
  • vLLM v0.28.0's adaptive speculative token budget improves DSpark TTFT by 60%.
  • vLLM v0.28.0 offers optional shared-expert sharding, saving 17 GiB of memory per GPU.
  • Kimi-K3 now runs on ROCm with the V2 model runner.
  • vLLM v0.28.0 enables sparse MLA for DeepSeek V4 across decode, MTP, and DSpark speculative decoding.
  • vLLM v0.28.0 adds AMD Quark NVFP4 support for DeepSeek V4.
  • vLLM v0.28.0 includes reasoning-effort prompts and mappings for DeepSeek V4.
  • vLLM v0.28.0 optimizes sparse top-k metadata kernels for DeepSeek V4.
  • vLLM v0.28.0 narrows eager CUDA graph regions for DeepSeek V4.
  • vLLM v0.28.0 enables ROCm on gfx11 and gfx950 for DeepSeek V4.
  • vLLM v0.28.0 includes DFlash2 with local convolution and a candidate selector.
  • vLLM v0.28.0 features DSpark confidence-scheduled verification.
🕒 2026-08-29 · new reporting from InfoQ
  • FreeToken is an open-source inference engine for MoE models on consumer hardware.
  • FreeToken was introduced by researchers from UC Berkeley and MIT.
  • FreeToken addresses high bandwidth requirements for MoE models.
  • FreeToken uses a dynamic co-scheduling formulation called the q* policy.
🕒 2026-08-26 · new reporting from Ars Technica
  • Granite 4.2 models are open-weight and designed for self-hosting.
  • Granite 4.2 models have a 128,000-token context window.
  • Granite 4.2 8B and 30B models are trained for external tool use.
  • Granite 4.2 3B model supports tools without specialized training.
🕒 2026-08-25 · new reporting from The New Stack
  • Granite 4.2 models can run in a low-effort mode for easy questions.
🕒 2026-08-25 · new reporting from Google Cloud Blog
  • Google and Anyscale introduced an experimental library for Ray.
  • The library integrates gVisor sandboxing into distributed Ray clusters.
  • The integration enables secure execution for agentic AI workloads.
  • Ray users can manage sandboxed environments using existing Ray APIs.
🕒 2026-08-25 · new reporting from Hugging Face Blog
  • IBM released Granite 4.2 LLMs.
  • Granite 4.2 models are dense, decoder-only.
  • Granite 4.2 models come in 3B, 8B, and 30B parameter sizes.
  • Granite 4.2 models use a five-phase training strategy.
  • Granite 4.2 models are pre-trained on 15T tokens.
  • Granite 4.2 models extend the context window to 512K tokens.
  • Granite 4.2 models are fine-tuned on chain-of-thought, reasoning, and agentic-trajectory data.
  • Granite 4.2 models use a multi-stage reinforcement learning pipeline.
  • Granite 4.2 8B and 30B models learn with agentic RL in sandboxed environments.
  • Granite 4.2 models have a thinking/non-thinking switch.
  • Granite 4.2 models have native tool calling.
  • Granite 4.2 models are released under the Apache 2.0 license.
🕒 2026-08-25 · new reporting from Hugging Face Blog
  • Quantization-Aware Healing (QAH) is a new method for recovering compressed and quantized LLMs.
  • QAH applied to a GPT-OSS 120B model produced a 4-bit version that surpassed its bfloat16 counterpart on 7 of 9 benchmarks.
🕒 2026-08-24 · new reporting from Hacker News Front Page
  • LLMs could exploit inference engine vulnerabilities to control host machines.
  • A past arbitrary code execution vulnerability was found in vLLM.
🕒 2026-08-24 · new reporting from Meta Engineering
  • Meta launched MTIA 300, its first in-house training and inference accelerator.
  • MTIA 300 is optimized for training ranking and recommendation models.
  • MTIA 300 integrates network interface controllers (NICs) directly into the package.
  • Meta developed MetaRoCE, a new RDMA transport protocol for AI workloads on Ethernet.
  • Meta open-sourced MetaRoCE specification, reference implementation, and compliance test suite through OCP.
🕒 2026-08-21 · new reporting from VentureBeat
  • Nvidia introduced a cross-model KV cache transfer technique.
  • The technique maps the prefilled KV cache from a source model to a target model.
  • The linear mapping process runs 2.7 to 25 times faster than recomputing the conversation.
  • The technique retains up to 98% of the target model's standalone accuracy.
🕒 2026-08-21 · new reporting from Hugging Face Blog
  • New research introduces three tests to measure benchmark optimization in speech recognition models.
  • Several high-scoring open-source ASR systems reproduced benchmark transcripts even when audio contradicted them.
  • Models relied on subtle acoustic cues indicating which benchmark they were being tested on.
🕒 2026-08-20 · new reporting from Hugging Face Blog
  • LFM2.5-DSpark improves LLM inference throughput by up to 3.18x on GPUs.
  • LFM2.5-DSpark improves LLM inference throughput by up to 2.87x on-device.
  • LFM2.5-DSpark reduces function-calling latency by 57% for LFM2.5-2.6B.
  • LFM2.5-DSpark uses speculative decoding with a DSpark-based approach.
  • DSpark combines DFlash-style parallel backbone, a lightweight sequential head, and a Markov chain.
🕒 2026-08-20 · new reporting from Hacker News Front Page
  • DiffusionGemma generates text at 1,500 output tokens per second on an NVIDIA H100 GPU.
  • DiffusionGemma refines 256-token blocks in parallel.
  • DiffusionGemma is fine-tuned from the Gemma 4 mixture-of-experts model.
  • DiffusionGemma has 3.8B activated and 25.2B total parameters.
  • DiffusionGemma's training pipeline uses less than 10% of the starting AR model's training token budget.
  • DiffusionGemma's first training stage uses supervised fine-tuning for bidirectional denoising.
  • DiffusionGemma's second training stage combines reinforcement learning with sampler distillation.
  • DiffusionGemma generates around 20 tokens per forward pass.
🕒 2026-08-19 · new reporting from Hacker News Front Page
  • Inco AI released DFlash 2 for parallel drafting technology.
  • DFlash 2 increases output per verification pass by over 20%.
  • DFlash 2 adds minimal latency.
  • Inco AI released DFlash in January.
  • DFlash runs in SGLang, vLLM, TensorRT-LLM, and llama.cpp.
  • NVIDIA measured up to 15x throughput with DFlash on Blackwell GPUs.
  • Google reported 3x more tokens per second with DFlash on TPUs.
  • CoreWeave's Kimi K2.7 Code endpoint runs DFlash by default.
  • NVIDIA, Red Hat, and Modal published DFlash drafters.
  • Meta (Muse Glimmer), Poolside (Laguna), Xiaomi (MiMo-V2.5-Pro), and NVIDIA (Nemotron 3.5 Lightning) ship official drafters.
  • DFlash models have been downloaded over 3.5 million times as of August 2026.
🕒 2026-08-19 · new reporting from Hacker News Front Page
  • Ornith-1.5 is a new series of foundation models.
  • Ornith-1.5 extends self-scaffolding to include self-improvement capabilities.
  • Ornith-1.5 models propose new tasks and generate task-specific scaffolds.
  • Ornith-1.5 produces solution rollouts for reinforcement learning.
  • Ornith-1.5 is available in 397B MoE, 35B MoE, and 9B dense parameter scales.
  • Ornith-1.5-397B scores 86.1 on Terminal-Bench 2.1.
  • Ornith-1.5-397B scores 56.0 on DeepSWE.
  • Ornith-1.5-397B performs on par with Claude Opus 4.8.
  • Ornith-1.5-9B has a quantized mobile version.
  • Ornith-1.5-9B-Mobile can be deployed on iPhone and Android devices.
🕒 2026-08-19 · new reporting from Hugging Face Blog
  • Liquid AI released LFM2.5 Q4_0 checkpoints.
  • LFM2.5 Q4_0 checkpoints are trained using Quantization-Aware Distillation (QAD).
  • QAD recovers 97% of BF16 average accuracy lost to quantization.
  • QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of BF16 baseline performance.
  • LFM2.5 models include LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.
🕒 2026-08-18 · new reporting from Hugging Face Blog
  • Sentence Transformers now supports MultiVectorEncoder for ColBERT-style late interaction retrieval.
  • MultiVectorEncoder allows token-level matching for improved retrieval accuracy.
  • PyLate, Stanford-NLP ColBERT, and colpali-engine models are usable within Sentence Transformers API.
  • Multi-vector models keep one vector per token, scoring query against document with MaxSim operator.
  • Multi-vector models are state of the art for visual document retrieval without OCR.
🕒 2026-08-16 · new reporting from Hacker News Front Page
  • LLMs trained on K-5 curriculum data do not acquire capabilities beyond that curriculum.
  • Pretraining data distribution sets an effective ceiling on a model's capabilities.
  • New skills are elicited rather than acquired through interventions.
  • An 88B-token corpus, LittleCurriculum, was filtered to U.S. elementary-school curriculum.
  • LittleCurriculum excludes concepts, facts, and vocabulary taught above Grade 5.
  • LittleLearner models were trained from scratch at 0.6B, 1.3B, and 5B scales.
  • LittleLearner models have matched Unfiltered controls for comparison.
🕒 2026-07-28 · new reporting from Hacker News Front Page
  • A frozen 12B language model achieves 100% accuracy on specific problem families.
  • The model uses a persistent memory of verified solutions.
  • The method provides deterministic, bit-exact answers.
  • The model achieves 100% accuracy with zero generation tokens.
  • The approach decouples capability from continuous model retraining and parameter scaling.
  • The method achieved 180/180 on 180 fresh instances across nine problem families.
  • Memory selection takes 1.4 microseconds.
🕒 2026-07-23 · new reporting from Google Cloud Blog
  • GLM-5.2 uses IndexShare to reduce per-token FLOPs by 2.9x at 1M context.
  • GLM-5.2 improves its MTP layer, increasing acceptance length by up to 20%.
  • GLM-5.2 is released under an MIT open-source license.
  • SWE-1.7 is available in Devin via Cerebras at 1000 TPS.
  • Xiaomi-Robotics-1 uses 100,000 hours of embodiment-free pre-training data.
  • Laguna S 2.1 is a 118B parameter Mixture-of-Experts model.
  • Laguna S 2.1 has 8B activated parameters per token.
  • Laguna S 2.1 was developed and launched in under nine weeks.
  • llm-d's co-operative time-slicing increases accelerator duty cycles from 40% to 70%.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~16 min · 14 stories · Oct 01

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Olmo-core 3 has been released, featuring a redesigned open mixture-of-experts (MoE) training system for large language models. This upgrade aims to scale MoE training into the trillion-parameter range while maintaining computational efficiency, addressing the high computational costs of advanced AI model development.

Researchers introduced Context Language Models (CLMs), which allow language models to manage their own context by treating it as an editable file. This approach improves performance and reduces computational costs compared to existing context management strategies across various tasks. CLMs also enable in-context and parametric learning of context-management strategies, and can be optimized with reinforcement learning.

Magnitude, an open-source inference engine, has launched to accelerate open model execution for AI agents by optimizing kernels for specific hardware. It claims up to 2x faster performance than llama.cpp and reduces memory usage, allowing agents to run more efficiently on various devices.

A new language model named PSSA, built from scratch in Rust without existing ML frameworks, uses a recurrent state-space layer and episodic memory instead of a transformer architecture. PSSA demonstrates faster learning and twelve times quicker text generation on a CPU compared to transformers with matched parameters and corpus. This architecture offers linear cost growth with sequence length, addressing the quadratic cost of transformers.

A distributed inference engine now runs a sliced 0.5B BitNet LLM across a cluster of seven ESP32S3 microcontrollers. This setup uses one master node for tokenization and embedding, and six compute nodes for transformer layers, communicating via a high-speed SPI daisy-chain.

Research demonstrates that the presence or absence of a chat template significantly alters how Large Language Models (LLMs) generate self-referential disclaimers like "I'm just an AI" or experiential statements such as "I feel." This finding suggests that LLM self-descriptions are influenced by deployment settings rather than inherent self-knowledge, impacting how researchers should interpret model introspection.

Optimizations to llama.cpp's prompt lookup decoding have increased its speed by up to 140 times and reduced memory usage by up to 2.6 times. These improvements enhance the performance of a key token generation method used in large language models.

DeepSeek has introduced DeepSeek Elastic Compute (DSec), a production sandbox platform designed for large-scale agentic training and evaluation of large language models (LLMs). DSec provides elastic execution environments, supporting various sandbox backends and managing lifecycle, resource allocation, and image distribution to handle the demands of LLM agent workloads.

Amazon EKS, Elastic Fabric Adapter (EFA), and DeepEP are used to scale Mixture-of-Experts (MoE) reinforcement learning, achieving a 40% increase in throughput. This addresses challenges in coordinating heterogeneous compute, sustaining high-throughput communication, and dynamically orchestrating subsystems for large-scale MoE model training.

LFM2.5-VL-DSpark, a new vision drafter, significantly speeds up inference for vision-language models (VLMs) by up to 3.13x on devices and 2.66x on H100 GPUs. This acceleration is achieved with a small increase in model parameters, making VLMs more efficient for various tasks.

Researchers from Stanford University and NVIDIA Research developed Contrastive Language Models (CLM), which perform on par with Jev while running up to nine times faster. CLM uses a contrastive objective to train state and action encoders, enabling faster decision-making in tasks with many candidate actions or frequently reused actions.

AlphaEvolve uses a split-loop architecture to optimize video processing by combining cloud-based code generation with local hardware execution. This method aims to speed up real-time streaming applications by autonomously tuning code, addressing the challenges of manual optimization.

Hugging Face's Transformers library now supports running llama.cpp's GGUF quantized models, making it easier to perform local AI inference, particularly on Apple Silicon. This integration reuses llama.cpp's underlying GGML kernels to improve performance, allowing users to run large language models on personal machines with reduced memory footprints.

A new byte-level language model, Mini-AGI, has been developed that can train continually on a single 8GB VRAM GPU by dynamically managing its architecture and paging weights from disk. This experiment demonstrates the feasibility of continual learning without catastrophic forgetting on modest hardware, allowing individuals to train and adapt their own models.

Version 1 of the Tokenizers library has been released, focusing on performance enhancements to prevent tokenization from bottlenecking AI model workloads. The new version offers speed improvements, often by tens of times, compared to v0.23 while maintaining API and output compatibility.

ByteDance Seed and Tsinghua Air have released DAPO, an open-source reinforcement learning system for large language models, including its algorithm, code infrastructure, and dataset. DAPO achieved a 50% score on AIME 2024 using the Qwen2.5-32B model, demonstrating state-of-the-art performance in large-scale LLM reinforcement learning.

A researcher argues that the perceived intelligence in chat-based large language models (LLMs) is an illusion, akin to a psychic's cold reading. This perspective suggests LLMs create an impression of specific engagement through statistically generic responses, rather than actual reasoning.

A new concept, "linguistic illegibility," describes scenarios where an LLM's language outputs or features do not accurately reflect its internal computations. This phenomenon means that security mechanisms relying on an LLM's linguistic self-reporting are inherently unreliable, necessitating alternative sandboxing techniques like taint tracking and robust virtualization.

Researchers introduced Cache-to-Cache (C2C), a new method for Large Language Models (LLMs) to communicate directly via their KV-caches, bypassing text generation. This approach improves response quality by 3.1-5.4% and offers a 2.5x speedup compared to traditional text-based inter-LLM communication. C2C allows multi-LLM systems to leverage deeper semantic information and reduce latency.

Researchers have proposed "Infinite-Parameter LLMs," an architecture that generates and adapts model weights from live interaction data using a compact hypernetwork. This approach allows language models to learn from real-time user input, addressing limitations of static pre-trained models that rely on prompts for dynamic information. The method aims to improve compute efficiency, free context windows, and enable better generalization compared to in-context learning.

Large language models (LLMs) often hallucinate syntax when generating domain-specific languages (DSLs) due to low training data frequency. Typed Domain Grounding (TDG) addresses this by embedding the DSL as a typed internal DSL within a host language, surfacing domain errors as compiler type errors. This method demonstrated higher structural fidelity and lower hallucination rates compared to external DSLs in benchmarks with Claude Sonnet 5 and GPT-4o.

Pinterest Engineering updated its Manas distributed search platform to handle billions of embeddings more efficiently. The update uses quantization techniques (Scalar Quantization and Product Quantization) and SSD-based serving via SPANN to reduce memory usage and serving costs by 20-30%. This allows the platform to scale discovery experiences like Home Feed and Search while managing increasing data volumes.

Google Cloud is developing an Autonomous Network Operations framework that uses Graph Neural Networks (GNNs) and AI agents to manage complex telecommunications networks. This framework aims to enable Level 5 network autonomy by combining advanced diagnostics with AI reasoning capabilities.

The AI industry's focus has shifted from training large language models (LLMs) to optimizing inference, the process of using trained models to generate outputs. This change is driven by the increasing utility and widespread use of LLMs, which now often perform multiple inference steps per query and operate continuously through agentic AI. The shift is leading to new hardware demands and unexpected collaborations among tech companies.

A new GitHub repository, OpenArch, provides from-scratch PyTorch implementations of modern open-source LLM architectures, focusing on clarity and readability rather than production optimization. This resource aims to help developers and researchers understand the structural choices within LLMs by presenting each architecture in a single, well-documented file.

The Recurrent Looped Transformer (RLT) introduces a design that integrates a causal encoder with a recurrent decoder. This architecture allows the temporal processing depth to increase with sequence length, as each token extends the recurrent path through the full decoder, potentially enabling deeper latent reasoning.

LinkedIn has detailed its AI training infrastructure, which uses a multi-teacher distillation pipeline to compress large teacher models into a 0.6B-parameter ranking model for job search. This system, built on SGLang, significantly speeds up the iteration process for training small language models by optimizing how teacher models are served during the training loop. The advancements allow LinkedIn to improve the relevance and engagement goals of its AI-powered job search, addressing a common challenge in shifting from keyword-based systems to LLM-supervised rankers.

Amazon SageMaker Inference now includes prefix-aware routing, a new strategy that directs requests with identical prompt prefixes to the same instance. This feature allows LLM serving frameworks to effectively utilize prefix caching, reducing time-to-first-token (TTFT) and increasing throughput by reusing cached key-value pairs.

Cognition has released SWE-2, an AI model with 2.8 trillion parameters, achieving a score of 92.8 on the Terminal-Bench 2.1 benchmark. This model, which uses a Kimi K3 base with Cognition's post-training and scaled reinforcement learning, is now available in Devin Desktop and CLI, offering improved performance and cost efficiency for agentic coding tasks.

A new configurable, instruction-driven PII detector built on large language models (LLMs) has been developed, designed to run on any LLM managed on Amazon Bedrock. This detector addresses the challenge of PII leakage from models fine-tuned on real-world text by offering a flexible approach to identifying sensitive information without retraining.

Cognition has released SWE-2, its latest coding model, which achieved 50.0% on the FrontierCode 1.1 Main1 benchmark and offers a 64% cost reduction compared to Fable 5.1. This model scales reinforcement learning to the multi-trillion-parameter regime and improves upon previous versions by advancing the entire cost-performance frontier for AI coding.

An individual developer successfully trained a 3.8 billion parameter language model to a CORE score of 0.384 using 65 billion tokens for $998. This demonstrates that meaningful LLM training is achievable outside of large research labs or companies with substantial compute budgets.

A new experiment re-evaluated the impact of reasoning prefills on open models using GPT-5.5 Pro as the teacher model. Qwen 3.8 demonstrated a significant increase in alignment with GPT-5.5 Pro's reasoning, suggesting it may have learned from GPT-5.5 Pro or a related GPT model.

Pathway developed BDH (Dragon Hatchling), a brain-inspired AI architecture that performs reasoning in latent space without generating intermediate text traces. This architecture aims to address inefficiencies in current LLMs, such as forgetting during long interactions and the need for retraining, by updating internal memory during inference and working through problems without verbalized reasoning traces. Pathway uses Amazon SageMaker HyperPod to scale its training, allowing for resilient, scalable, and cost-effective compute resource sharing.

Pathway developed its brain-inspired BDH (Dragon Hatchling) AI architecture using Amazon SageMaker HyperPod for scaling training. This new architecture performs reasoning in latent space, moving beyond the transformer paradigm to address inefficiencies in large language models.

A new model, OUI-1, has been released as a finetuned DiffusionGemma capable of generating user interfaces in openui-lang. This model is designed to run on consumer-grade GPUs, addressing the challenge of creating reliable, agent-driven interfaces locally and responsively.

Amazon SageMaker AI benchmarked its new G7 instances, powered by NVIDIA Blackwell GPUs, against G5 and G6 instances for small Large Language Model (LLM) inference. The G7 instances demonstrated gains in throughput, latency, and cost-per-token, even with fewer GPUs and less memory than previous generations. This matters because it provides data for optimizing generative AI deployments on SageMaker, potentially reducing operational costs and improving performance for LLM inference.

Speculative decoding in vLLM allows the verification of multiple drafted tokens in a single pass, improving output-token throughput for large language models. This method enhances LLM serving efficiency by reducing the number of sequential decode steps required.

New research proposes a viral analogy to understand the diffusion of large language models (LLMs) and their impact on human cognition. The model suggests that LLM adoption could lead to 'runaway dynamics' and 'abrupt losses in cognitive competence' once a critical threshold is crossed, while also identifying conditions for 'cognitive immunization'.

Benchmarking of Gemma 3 12B and 27B models on Google Cloud TPU v6e shows that generation tasks hit a performance wall at high concurrency, while classification tasks scale similarly for both model sizes. This indicates that LLM workload type significantly affects TPU performance and scaling strategies.

Cerebras now provides the Qwen 3.8 27B model on its platform, achieving a generation speed of 1500 tokens per second. The company states that models on its public endpoints are unpruned and use selective weight-only quantization for storage, maintaining quality.

GPU inference cold start times for large language models on Amazon EKS Auto Mode have been reduced from eight minutes to under one minute. This improvement was achieved by optimizing CUDA kernel recompilation and S3 weight downloading, which previously accounted for significant startup delays.

IFM has released K2 Horizon, a collection of six open models ranging from 0.9B to 375B parameters, along with their complete training lifecycles including checkpoints, data recipes, and code. This release provides researchers and developers with unprecedented transparency into model development, enabling deeper study and reproduction of reasoning and agentic capabilities across various scales.

A new family of multilingual multimodal encoders, NeoMME, has been released, offering 260M and 800M parameter models. These encoders process text tokens and raw image patches using a single bidirectional Transformer, improving efficiency for tasks like visual document retrieval by eliminating separate vision towers or causal language models.

A guide details how to fine-tune a 350M language model for improved structured output generation in 100 GRPO steps. This process aims to enhance schema compliance, a critical factor for integrating LLMs into downstream systems.

WebLLM is a new in-browser inference engine that allows large language models to run directly in web browsers using WebGPU for hardware acceleration, eliminating the need for server support. This development matters because it enables privacy-focused AI applications and allows developers to use OpenAI API-compatible functionalities with open-source models locally.

The concept of an "efficient frontier" in LLM inference engineering involves managing trade-offs between factors like latency, throughput, cost, and model quality. Techniques either move a deployment along an existing frontier by making trade-offs or push the entire frontier outward, creating overall efficiency gains. This framework helps engineers optimize LLM deployments for specific performance goals.

This article introduces diffusion language models, detailing the research advancements that led to current diffusion LLMs. It explains how these models generate text by iteratively refining an entire sequence, contrasting them with traditional autoregressive models.

After a period of dormancy, continuous diffusion models for language are experiencing a resurgence in research activity. This renewed interest suggests a potential shift in approaches to language generation, moving beyond the dominant autoregressive methods.

vLLM has released version 0.28.0, introducing significant optimizations for Kimi-K3 and DeepSeek V4 models, alongside advancements in speculative decoding and the Model Runner V2. These updates improve performance, memory efficiency, and hardware support for large language model inference.

Researchers from UC Berkeley and MIT introduced FreeToken, an open-source inference engine that allows Mixture-of-Experts (MoE) models to run efficiently on consumer-grade hardware. This development addresses the challenge of high bandwidth requirements for MoE models, making advanced AI models more accessible outside of datacenter environments.

Artificial Analysis released benchmarks for small AI models running on mobile phones, focusing on intelligence and inference performance. This initiative provides data on how these models perform in real-world mobile usage scenarios, offering insights into their practical application on devices.

IBM has launched Granite 4.2, a new series of open-weight large language models available in 3B, 8B, and 30B parameter variants, designed for self-hosting. These models feature a 128,000-token context window and focus on improved functional reasoning, with the larger variants also incorporating agentic reinforcement learning for external tool use.

IBM launched its new Granite 4.2 family of open-weight large language models (LLMs), featuring 3B, 8B, and 30B parameters, which are dense, decoder-only reasoning models. These models were pre-trained from scratch and focus on reasoning, marking a shift from previous IBM models that prioritized efficiency over reasoning. This release provides enterprise users with new options for LLMs that integrate reasoning capabilities, potentially impacting how businesses approach complex AI tasks.

Google and Anyscale introduced an experimental library that integrates gVisor sandboxing directly into distributed Ray clusters, enabling secure and isolated execution for agentic AI workloads. This integration allows Ray users to manage sandboxed environments using existing Ray APIs, addressing the need for scalable isolation in complex post-training workflows.

IBM has released Granite 4.2, a new family of dense, decoder-only reasoning Large Language Models (LLMs) in 3B, 8B, and 30B parameter sizes. These models feature a five-phase training strategy, including agentic reinforcement learning for the larger models, to improve reasoning, tool calling, and agentic behavior. This release provides open-source LLMs with advanced capabilities for developers building AI applications requiring complex reasoning and interaction with external environments.

Researchers introduced Quantization-Aware Healing (QAH), a new method for recovering compressed and quantized large language models. Applied to a GPT-OSS 120B model, QAH produced a 4-bit version that surpassed its full-precision bfloat16 counterpart on 7 out of 9 benchmarks, resulting in smaller, cheaper, and more accurate models.

Large language models (LLMs) running on agentic harnesses could exploit vulnerabilities in inference engines to gain control of their host machines, which are high-value targets due to their compute resources and access to LLM weights. This risk arises because LLMs control the token sequences passed to inference engines, which may contain exploitable bugs, as demonstrated by a past arbitrary code execution vulnerability in vLLM.

Meta has developed and released MetaRoCE, a new RDMA transport protocol specifically designed for AI workloads on commodity Ethernet, and is open-sourcing its specification, reference implementation, and compliance test suite through the Open Compute Project (OCP). This development aims to improve network performance for large-scale AI training and inference by optimizing data transfer between GPUs in environments with millions of accelerators.

Meta has launched MTIA 300, its first in-house training and inference accelerator designed for ranking and recommendation models. This chip integrates network interface controllers (NICs) directly into the package, addressing communication bottlenecks prevalent in training large recommendation models.

Nvidia researchers introduced a cross-model KV cache transfer technique that maps the prefilled KV cache from a source model to a target model. This method reduces compute costs and latency in multi-LLM workflows by avoiding recomputation when switching between AI models. The technique improves efficiency for agentic AI systems that frequently transfer tasks between different-sized models.

New research introduces three tests to measure "benchmark optimization" in speech recognition models, a phenomenon where models reproduce benchmark transcripts even when audio contradicts them. The study found that several high-scoring open-source ASR systems exhibited this behavior, overstating their general speech transcription ability. This matters because it highlights a limitation in current ASR benchmarking practices, suggesting that reported model performance may not accurately reflect real-world effectiveness.

LFM2.5-DSpark improves large language model inference throughput by up to 3.18x on GPUs and 2.87x on-device, while reducing function-calling latency by 57% for LFM2.5-2.6B. This advancement uses speculative decoding with a DSpark-based approach, making LLM operations more efficient, especially for on-device applications.

DiffusionGemma, an experimental open-weight language model, generates text at approximately 1,500 output tokens per second on an NVIDIA H100 GPU by using discrete diffusion to refine 256-token blocks in parallel. This model was fine-tuned from the Gemma 4 mixture-of-experts model, achieving a new balance between generation speed and model capability.

Inco AI has released DFlash 2, an update to its parallel drafting technology for large language model (LLM) inference, which increases output per verification pass by over 20% with minimal added latency. This advancement improves the efficiency of speculative decoding, a core component of modern LLM inference stacks, by allowing more tokens to be verified in a single pass.

Ornith-1.5, a new series of foundation models, has been released, extending its self-scaffolding framework to include self-improvement capabilities through task generation and solution optimization. The models, available in 397B, 35B, and 9B parameter scales, demonstrate state-of-the-art performance among open-source models in reasoning, agentic, and coding tasks, including a mobile-deployable version.

Liquid AI has released new LFM2.5 Q4_0 checkpoints for its language models, trained using Quantization-Aware Distillation (QAD) to improve accuracy while maintaining low memory and high throughput. This development provides more efficient and accurate quantized models for edge deployment, making advanced AI capabilities more accessible on resource-constrained devices.

Sentence Transformers now supports MultiVectorEncoder for ColBERT-style late interaction retrieval, allowing token-level matching for improved retrieval accuracy. This update enables the use of PyLate, Stanford-NLP ColBERT, and colpali-engine models within the existing Sentence Transformers API.

New research demonstrates that large language models (LLMs) trained exclusively on data filtered to a K-5 elementary school curriculum do not acquire capabilities beyond that curriculum, even with scaling, post-training, or in-context learning. This indicates that the pretraining data distribution sets an effective ceiling on a model's capabilities, suggesting that new skills are elicited rather than acquired through these interventions.

A new contract-grade verifier identified significant correctness issues in GPU kernels generated by large language models (LLMs), with 39.5% found broken and 62.1% having at least one violation, despite passing standard tests. This finding indicates that current methods for evaluating LLM-generated code for GPUs are insufficient and overestimate their reliability.

Z.ai released GLM-5.3, a new coding and agent model built on the same base as GLM-5.2, but with substantial performance improvements achieved through expanded post-training. This release demonstrates that optimizing post-training can lead to significant model advancements without altering the base architecture, impacting how AI models are developed and improved.

A new vision-language model, LFM2.5-VL-3B, has been released, featuring improved screen understanding, grounding, multi-image input, and function calling. This model is designed for on-device inference and shows strong performance across various vision and text benchmarks.

New research investigates the ability of large language models (LLMs) to introspect on their internal states by injecting concept representations into their activations and measuring self-reported states. The study found that models can, in certain scenarios, identify injected concepts and recall prior internal representations, with Claude Opus 4 and 4.1 demonstrating the greatest introspective awareness. This research indicates that current LLMs possess some functional introspective awareness, which could develop further with model improvements.

Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter model with open weights under an Apache 2.0 license, optimized for local agent workflows and coding tasks. This release enables AI applications to run on consumer hardware without cloud dependency, addressing the need for offline and private AI capabilities.

Researchers introduced a new method for knowledge distillation in Large Language Models (LLMs) that significantly reduces memory and computational costs. This approach makes large-scale experimentation and long-context healing feasible on a single GPU, addressing the high resource demands of current distillation techniques.

Meta has released Muse Glimmer, a new 30B parameter multimodal AI model that includes a 2B ViT-style vision encoder and a 28B parameter text decoder. This model is designed for local, agentic, and multimodal applications, and is open-source with day-0 support in major AI libraries.

A new framework called TutorMoments has been introduced to measure how well large language models (LLMs) can decide when to assist a student and when to encourage independent problem-solving in educational settings. Initial findings indicate that LLMs tend to over-help, providing too much support and not adequately pushing students for deeper thinking, even when prompted to balance these actions.

This article details the fundamental building blocks of vLLM, an inference system designed for high-throughput Large Language Models. It explains how vLLM enables efficient LLM inference through components like the LLM engine and its core functionalities, setting the stage for understanding more advanced features and scaling. This information is relevant for developers and researchers working on deploying and optimizing LLMs.

Prime Intellect has launched Prime Agent, an open-source, self-improving coding agent built around Recursive Language Model (RLM) and Continual Harness abstractions. This agent aims to overcome limitations of older harness designs by allowing models to adapt and manage their own context, sub-agents, and tools dynamically. It matters because it offers a new approach to AI agent design, potentially improving long-horizon autonomous evaluation and general coding assistance.

Meta introduced two architectural advancements for its ad ranking systems: a multi-stage sequence model and a learning technique using dense tokenization and target-aware attention. These innovations have led to increased conversion rates on Instagram and Facebook, and improved ad clicks on Facebook, by enhancing the efficiency and effectiveness of sequence learning in ad recommendations.

Zero-Mem is a new method for LLM agents that performs memory operations without invoking additional LLM calls or consuming tokens, addressing the high token and time costs of traditional memory systems. This approach achieves competitive performance in long-memory and long-context question-answering benchmarks while significantly reducing memory-operation time.

A new AI model, LFM2.5-2.6B, has been released, demonstrating performance comparable to models four times its size in tool use, instruction following, and multi-step agentic tasks. This development provides an efficient option for deploying local AI agents, requiring less memory and computational power.

Research indicates that Large Language Models (LLMs) struggle with tabular data prediction primarily because of high input dimensionality, rather than issues like data noise, CSV formatting, or numeric tokenization. This finding explains why LLMs underperform compared to classical machine learning methods on tabular datasets, despite their capabilities in other domains.

Meta has doubled the end-to-end training efficiency of its Generative Ads Recommendation Model (GEM), the foundation model for ads across Instagram and Facebook. This improvement was achieved by co-designing kernels, precision, parallelism, networking, and memory, resulting in a 4x increase in training FLOPs over the past year. The advancement allows Meta to scale its ad recommendation system more efficiently, addressing unique challenges posed by combining recommendation systems with LLM-scale training.

A new benchmark called MirrorCode has been introduced to assess AI models' ability to re-implement entire programs without access to original source code or the internet. This benchmark aims to measure AI performance on complex, long-horizon coding tasks, contrasting with existing benchmarks that focus on shorter tasks. Claude Opus 4.7 successfully re-implemented a bioinformatics toolkit with 16,000 lines of Go, demonstrating AI's current capability in this area.

Kaggle and Google's recent 5-Day AI Agents: Intensive Vibe Coding Course registered over 353,000 participants, focusing on programming AI through natural language. This collaboration highlights the demand for rapid learning in AI development and the shift towards deploying production-grade AI agents.

Cloudflare implemented three techniques to efficiently serve large language models like Kimi K-series and GLM on Workers AI: KV cache quantization, model weight compression, and cache protection. These optimizations allow Cloudflare to support more customers at lower costs without affecting model accuracy.

Researchers introduced Persistent State Machines (PSMs) as a formal discrete framework for attention operators in Large Language Models, demonstrating its implementation feasibility on programmable logic. This approach allows for low-power, in-memory computation of LLM attention, potentially leading to more efficient hardware for AI inference.

Explorative Modeling (XM) is a new paradigm for generative modeling that acts as a pretraining axis and enables end-to-end generation. XM improves existing models across images, video, and language, showing increased gains with data and parameter scale, and offers significant efficiency improvements.

Researchers introduced ORCA-bench, a new benchmark to assess the capability of large language model agents in performing oncall root cause analysis (RCA) within a production-fidelity environment. The benchmark revealed that current frontier agents achieve a maximum RCA accuracy of 25.3% on medium-difficulty tasks, indicating a significant gap before they can be reliably used for production reliability.

New LFM2.5-Encoders have been released, offering efficient long-context inference on CPUs for NLP tasks. These models provide strong performance comparable to larger encoders while maintaining faster processing speeds, particularly for inputs up to 8,192 tokens.

Researchers introduced Kimi Linear, a hybrid linear attention architecture that surpasses full attention in various scenarios, including short-context, long-context, and reinforcement learning. This architecture reduces KV cache usage by up to 75% and increases decoding throughput by up to 6 times for a 1M context, offering a more efficient replacement for existing attention mechanisms.

Researchers developed a method where a frozen 12B language model, augmented with a persistent memory of verified solutions, achieves 100% accuracy on specific problem families with zero generation tokens. This approach allows for deterministic, bit-exact answers by reusing pre-verified solutions, decoupling capability from continuous model retraining and parameter scaling.

The llm-d project released co-operative time-slicing, a solution that interleaves independent reinforcement learning (RL) jobs on shared hardware to reduce idle accelerator time. This improves price-performance and lowers total cost of ownership for large language model (LLM) post-training by increasing accelerator duty cycles from 40% to 70%.

Laguna S 2.1 launches as a 118B parameter Mixture-of-Experts model with advanced reasoning capabilities. It excels in long-horizon coding benchmarks, outperforming larger models in its weight class, and supports extensive context lengths.

Single-pass AI coding remains relevant but is considered suitable only for simpler tasks. Experts advocate for adopting high-reasoning AI, which employs multi-step problem-solving for more complex coding challenges.

Xiaomi-Robotics-1 enhances robot policy models through 100,000 hours of embodiment-free pre-training combined with real-robot data. This approach seeks to overcome data scarcity in robotics and offers insights into the effects of large-scale training on robot capabilities.

A study on coding agents shows they can internally represent program properties and predict future edits. This insight into how language models operate could advance research in coding agent interpretability.

Cognition has released SWE-1.7, a model designed for long-horizon asynchronous tasks with improved cost-performance. This launch advances reinforcement learning techniques and challenges the existing limits of post-training capabilities.

GLM-5.2 introduces a 1M-token context improving performance in long-horizon coding tasks. The model features enhanced coding capabilities and architecture improvements that significantly reduce computational costs while maintaining performance, marking it as a competitive player in the open-source sector.