← All stories
● Covered by 7 sources · 20 reportsMedium impact18 neutral2 positive

DeepSeek Launches Open-Source AI Agent Harness and Updates Flagship Model

🔄 Updated 7h ago — new reporting from Hacker News Front Page
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • DeepSeek Harness (dsh) is an open-source AI agent runtime.
  • It features a plugin-based architecture, allowing customization of all components.
  • DeepSeek-V4-Pro model is released, focused on agentic workloads.
  • API pricing for DeepSeek-V4 will change to peak and off-peak rates.
  • The harness is available on GitHub and built on Node.js and Cordis.
  • TrueForge is an open-source AI agent harness released by TrueFoundry.
  • TrueForge is released under the MIT License.
  • TrueForge reduces task completion costs by 30-75% compared to Anthropic's Claude Managed Agents.
  • TrueForge paired with GLM-5.2 LLM cost 75% less than Claude Opus 4.8 on DevRev's Enterprise-Bench ($2.90 vs $11.80).
  • TrueForge with Opus 4.8 saved 30% compared to Claude Managed Agents with Opus 4.8 ($8.50 vs $11.80).
  • DeepSeek Harness is available in developer preview.
  • DeepSeek Harness is released under the MIT license.
  • DeepSeek Harness is built on the Cordis meta-framework.
  • DeepSeek Harness uses a micro-kernel architecture.
  • DeepSeek Harness components include model adapters, tool registries, sandboxing, session state handlers, event dispatchers, and UIs.
  • DeepSeek Harness configuration schemas support YAML or JSON definitions.
  • Nvidia research shows the harness is more important than the model for long-horizon tasks.
  • A custom harness with memory management and a supervisor component enabled Claude Opus 5 to score 100% on ARC-AGI-3.
  • Without the harness, Opus 5 scored 30% on ARC-AGI-3.
  • Harness engineering is a methodology for managing AI-assisted code generation.
  • Harness engineering ensures AI-generated code remains correct and consistent over time.
  • Harness engineering addresses AI assistants' tendency to introduce inconsistencies and drift.
  • Harness engineering provides a framework for integrating AI coding tools without compromising code quality.
  • Harness engineering uses deterministic tooling, agent-based review, and periodic entropy checks.
  • The term "harness engineering" comes from Birgitta Boeckeler's article on martinfowler.com.
  • Aider, Claude Code, and OpenClaw ran an identical model, but token use varied 70-fold.
  • Composio compared eight harnesses using DeepSeek V4 Flash on 30 enterprise workflows.
  • Artificial Analysis continuously tracks harness-model pairings in its coding-agent index.
  • Composio's workflows spanned Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack, and PostHog.
  • Composio's tasks ran under a 900-second ceiling.
  • Composio used a programmatic verifier, not an LLM judge, to grade outcomes.
  • Meta AI and University of Illinois Urbana–Champaign developed EvoHarness-RL.
  • EvoHarness-RL enables AI agents to autonomously manage complex enterprise workflows.
  • EvoHarness-RL helps agents learn when to interact with their environment.
  • EvoHarness-RL addresses limitations of AI agents that rely on rigid, human-written rules.
  • EvoHarness-RL improves agents' ability to handle dynamic situations and recover from errors.
  • The harness manages inputs, checks outputs, and handles failures.
  • Demos often overstate an agent's readiness by operating under ideal conditions.
  • Production environments require systems to manage incomplete data, tool errors, and policy changes.
  • The harness decides what data the agent sees.
  • Harness rebuilt its Git repository.
  • Harness introduced a new AI Code Review product.
  • The new products address the surge in pull requests from AI coding agents.
  • GitHub experienced a nearly eight-hour platform outage on August 17.
  • Microsoft Agent Framework includes an agent harness.
  • A four-part live series, "From Model to Agent: The Agent Framework Harness, Live in C#," demonstrates building a C# AI agent.
  • The series covers adding tools, memory, skills, observability, and deployment to an agent.
  • The series streams live on the .NET YouTube channel and Microsoft Reactor.
  • The series runs for four consecutive Thursdays in September.
  • Zed is rebuilding how people review agent-written code.
  • OpenRouter gives companies more control over where requests are processed.
  • Anthropic is rebuilding its stack to be less confusing.
  • Vercel's AI Gateway reported the average price per token fell 23.2% in August.
  • August was the third straight monthly decline in average token price.
  • AWS open-sourced Strands Harness, a general-purpose AI agent.
  • Strands Harness is built on the existing Strands SDK.
  • Strands Harness is 45% cheaper than Claude Code and Codex.
  • Strands SDK debuted in May 2025 as an open-source Python SDK.
  • AWS later brought Strands to TypeScript.
  • AWS created Strands Labs in February for experimental projects.
  • Marc Brooker is VP and distinguished engineer at AWS.
  • Unreal Labs developed Unreal Agent.
  • Unreal Agent is a harness that manages AI agent tool calls asynchronously.
  • Unreal Agent reduces the underlying model's overhead.
  • Unreal Agent allows user steering during tool calls.
  • Unreal Agent schedules more efficient work.
  • Unreal Agent achieves up to 40% cost savings compared to Codex.
  • Unreal Agent achieves up to 20% cost savings compared to Pi.
  • Strands Harness is Apache 2.0 licensed.
  • Strands Harness is a general-purpose agent, not a coding agent.
  • Strands Harness costs 28% less using the same Claude or GPT models across six benchmarks.
  • Ryan Lopopolo, a Google Cloud software engineer, coined the term "agent harness".
  • Google Antigravity is a harness for Gemini 3.8 Flash.
  • Raven is a new AI harness.
  • Raven generates Directed Acyclic Graphs (DAGs) to orchestrate multiple specialized AI agents.
  • Raven supports recursive self-improvement by proposing, evaluating, and adopting changes to its own planning and action processes.
  • Raven is a Host Agent that brings built-in and third-party agents together.
  • Raven's modular architecture supports iterative improvement of its own harness.
  • Raven is powered by EverOS.
  • Raven carries memory and context across sessions.
  • Raven includes built-in agents: Raven-Research, Raven-Code, Raven-Design, and Raven-Oncall.
  • Raven-Research supports research.
  • Raven-Code supports coding.
  • Raven-Design supports visual design.
  • Raven-Oncall supports unattended workflow automation.
  • Raven is pre-alpha.
  • Raven completed three complete projects on the Multi-Agent Orchestration Benchmark.
  • Raven worked autonomously for about 4 days to complete projects.
  • Raven completed 42 rounds of work on projects.
  • DeepSeek Harness allows users to install or create plugins through chat in "Creator mode".
  • DeepSeek Harness enables organizing documents, analyzing spreadsheets, and writing code.
  • DeepSeek Harness allows previewing results and refining them through conversation.
  • DeepSeek Harness has a Scheduled tasks plugin for recurring work.
  • DeepSeek Harness allows inspecting execution traces and tool call details.
  • DeepSeek Harness can be launched via `npx @deepseek-ai/dsh web`.

DeepSeek Harness Enters Developer Preview

DeepSeek has introduced DeepSeek Harness (dsh), an open-source agent harness now available in developer preview. This Node.js-based runtime is distributed under an MIT license and its source code is accessible on GitHub. The project has garnered significant attention, accumulating over 33,000 GitHub stars within hours of its release.

Plugin-Based Architecture

A core feature of DeepSeek Harness is its modular design, where every capability functions as a plugin. This includes models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the user interface. This architecture, built on Cordis's plugin system, allows developers to select, swap, or extend any component through configuration without altering the core source code. The design ensures that every operation within the agent is traceable, with all model interactions recorded in an append-only session log.

DeepSeek-V4-Pro Model Release

Alongside the harness, DeepSeek launched DeepSeek-V4-Pro, an updated flagship AI model. This model is specifically designed for agentic workloads and is available through DeepSeek’s web interface, mobile app, and API. It includes native support for the OpenAI Responses API and integration with Codex.

API Pricing Changes

DeepSeek is also modifying its API pricing structure for V4 access. The existing flat API pricing will be replaced by peak and off-peak rates, effective from 16:00 UTC on Sunday, August 16. This change will result in higher costs for developers accessing V4 through the API.

Developer Tools Expansion

These releases signify DeepSeek's expansion into developer tools for AI agents, moving beyond just the model layer. DeepSeek Harness offers an alternative to integrated coding-agent environments and aims to provide developers with greater flexibility in building and deploying AI agents.

Updates

🕒 2026-10-02 · new reporting from Hacker News Front Page
  • DeepSeek Harness allows users to install or create plugins through chat in "Creator mode".
  • DeepSeek Harness enables organizing documents, analyzing spreadsheets, and writing code.
  • DeepSeek Harness allows previewing results and refining them through conversation.
  • DeepSeek Harness has a Scheduled tasks plugin for recurring work.
  • DeepSeek Harness allows inspecting execution traces and tool call details.
  • DeepSeek Harness can be launched via `npx @deepseek-ai/dsh web`.
🕒 2026-09-29 · new reporting from Hacker News Front Page
  • Raven is a new AI harness.
  • Raven generates Directed Acyclic Graphs (DAGs) to orchestrate multiple specialized AI agents.
  • Raven supports recursive self-improvement by proposing, evaluating, and adopting changes to its own planning and action processes.
  • Raven is a Host Agent that brings built-in and third-party agents together.
  • Raven's modular architecture supports iterative improvement of its own harness.
  • Raven is powered by EverOS.
  • Raven carries memory and context across sessions.
  • Raven includes built-in agents: Raven-Research, Raven-Code, Raven-Design, and Raven-Oncall.
  • Raven-Research supports research.
  • Raven-Code supports coding.
  • Raven-Design supports visual design.
  • Raven-Oncall supports unattended workflow automation.
  • Raven is pre-alpha.
  • Raven completed three complete projects on the Multi-Agent Orchestration Benchmark.
  • Raven worked autonomously for about 4 days to complete projects.
  • Raven completed 42 rounds of work on projects.
🕒 2026-09-25 · new reporting from Google Cloud Blog
  • Ryan Lopopolo, a Google Cloud software engineer, coined the term "agent harness".
  • Google Antigravity is a harness for Gemini 3.8 Flash.
🕒 2026-09-23 · new reporting from Hacker News Front Page
  • Strands Harness is Apache 2.0 licensed.
  • Strands Harness is a general-purpose agent, not a coding agent.
  • Strands Harness costs 28% less using the same Claude or GPT models across six benchmarks.
🕒 2026-09-22 · new reporting from Hacker News Front Page
  • Unreal Labs developed Unreal Agent.
  • Unreal Agent is a harness that manages AI agent tool calls asynchronously.
  • Unreal Agent reduces the underlying model's overhead.
  • Unreal Agent allows user steering during tool calls.
  • Unreal Agent schedules more efficient work.
  • Unreal Agent achieves up to 40% cost savings compared to Codex.
  • Unreal Agent achieves up to 20% cost savings compared to Pi.
🕒 2026-09-21 · new reporting from The New Stack
  • AWS open-sourced Strands Harness, a general-purpose AI agent.
  • Strands Harness is built on the existing Strands SDK.
  • Strands Harness is 45% cheaper than Claude Code and Codex.
  • Strands SDK debuted in May 2025 as an open-source Python SDK.
  • AWS later brought Strands to TypeScript.
  • AWS created Strands Labs in February for experimental projects.
  • Marc Brooker is VP and distinguished engineer at AWS.
🕒 2026-09-19 · new reporting from The New Stack
  • Zed is rebuilding how people review agent-written code.
  • OpenRouter gives companies more control over where requests are processed.
  • Anthropic is rebuilding its stack to be less confusing.
  • Vercel's AI Gateway reported the average price per token fell 23.2% in August.
  • August was the third straight monthly decline in average token price.
🕒 2026-09-18 · new reporting from .NET Blog
  • Microsoft Agent Framework includes an agent harness.
  • A four-part live series, "From Model to Agent: The Agent Framework Harness, Live in C#," demonstrates building a C# AI agent.
  • The series covers adding tools, memory, skills, observability, and deployment to an agent.
  • The series streams live on the .NET YouTube channel and Microsoft Reactor.
  • The series runs for four consecutive Thursdays in September.
🕒 2026-09-09 · new reporting from The New Stack
  • Harness rebuilt its Git repository.
  • Harness introduced a new AI Code Review product.
  • The new products address the surge in pull requests from AI coding agents.
  • GitHub experienced a nearly eight-hour platform outage on August 17.
🕒 2026-08-30 · new reporting from The New Stack
  • The harness manages inputs, checks outputs, and handles failures.
  • Demos often overstate an agent's readiness by operating under ideal conditions.
  • Production environments require systems to manage incomplete data, tool errors, and policy changes.
  • The harness decides what data the agent sees.
🕒 2026-08-28 · new reporting from VentureBeat
  • Meta AI and University of Illinois Urbana–Champaign developed EvoHarness-RL.
  • EvoHarness-RL enables AI agents to autonomously manage complex enterprise workflows.
  • EvoHarness-RL helps agents learn when to interact with their environment.
  • EvoHarness-RL addresses limitations of AI agents that rely on rigid, human-written rules.
  • EvoHarness-RL improves agents' ability to handle dynamic situations and recover from errors.
🕒 2026-08-27 · new reporting from The New Stack
  • Aider, Claude Code, and OpenClaw ran an identical model, but token use varied 70-fold.
  • Composio compared eight harnesses using DeepSeek V4 Flash on 30 enterprise workflows.
  • Artificial Analysis continuously tracks harness-model pairings in its coding-agent index.
  • Composio's workflows spanned Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack, and PostHog.
  • Composio's tasks ran under a 900-second ceiling.
  • Composio used a programmatic verifier, not an LLM judge, to grade outcomes.
🕒 2026-08-27 · new reporting from Hacker News Front Page
  • Harness engineering is a methodology for managing AI-assisted code generation.
  • Harness engineering ensures AI-generated code remains correct and consistent over time.
  • Harness engineering addresses AI assistants' tendency to introduce inconsistencies and drift.
  • Harness engineering provides a framework for integrating AI coding tools without compromising code quality.
  • Harness engineering uses deterministic tooling, agent-based review, and periodic entropy checks.
  • The term "harness engineering" comes from Birgitta Boeckeler's article on martinfowler.com.
🕒 2026-08-21 · new reporting from TechCrunch
  • Nvidia research shows the harness is more important than the model for long-horizon tasks.
  • A custom harness with memory management and a supervisor component enabled Claude Opus 5 to score 100% on ARC-AGI-3.
  • Without the harness, Opus 5 scored 30% on ARC-AGI-3.
🕒 2026-08-20 · new reporting from InfoQ
  • DeepSeek Harness is available in developer preview.
  • DeepSeek Harness is released under the MIT license.
  • DeepSeek Harness is built on the Cordis meta-framework.
  • DeepSeek Harness uses a micro-kernel architecture.
  • DeepSeek Harness components include model adapters, tool registries, sandboxing, session state handlers, event dispatchers, and UIs.
  • DeepSeek Harness configuration schemas support YAML or JSON definitions.
🕒 2026-08-20 · new reporting from VentureBeat
  • TrueForge is an open-source AI agent harness released by TrueFoundry.
  • TrueForge is released under the MIT License.
  • TrueForge reduces task completion costs by 30-75% compared to Anthropic's Claude Managed Agents.
  • TrueForge paired with GLM-5.2 LLM cost 75% less than Claude Opus 4.8 on DevRev's Enterprise-Bench ($2.90 vs $11.80).
  • TrueForge with Opus 4.8 saved 30% compared to Claude Managed Agents with Opus 4.8 ($8.50 vs $11.80).

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

DeepSeek AI launched DeepSeek Harness, a new open-source platform that allows users to install or create plugins to extend AI tools, skills, and interfaces. This platform enables organizing documents, analyzing spreadsheets, and writing code, with features like scheduled tasks and execution trace inspection.

Raven, a new AI harness, generates Directed Acyclic Graphs (DAGs) to orchestrate multiple specialized AI agents for complex tasks. It features a modular architecture that supports recursive self-improvement by proposing, evaluating, and adopting changes to its own planning and action processes.

The Agent Factory podcast featured Ryan Lopopolo, a Google Cloud software engineer, who discussed the concept of an "agent harness" and its role in enabling autonomous coding workflows. An agent harness is defined as the surrounding components that augment a large language model (LLM) to allow it to interact with external tools and context, moving beyond simple question-answering to perform complex tasks.

Strands harness, an Apache 2.0 licensed agent harness, has been released to simplify AI agent development and deployment in the cloud or locally. It offers improved cost efficiency and comparable accuracy to existing solutions like Claude Code and Codex.

Unreal Labs developed Unreal Agent, a harness that manages AI agent tool calls asynchronously. This approach reduces the underlying model's overhead, allowing for user steering during tool calls and scheduling more efficient work, resulting in up to 40% cost savings compared to Codex and 20% compared to Pi.

AWS has open-sourced Strands Harness, a general-purpose AI agent built on the existing Strands SDK, which provides a preconfigured agent with tools for developers. This release allows developers to run AI agents locally or deploy them to the cloud, offering a foundation for longer-running tasks and potentially reducing costs compared to other AI models.

Recent updates from Zed, Anthropic, and OpenRouter demonstrate a focus on improving the software surrounding AI models to make them more usable for end-users. These developments address challenges in code collaboration, model routing, and user interface complexity, indicating a shift towards better integration and practical application of AI. The trend suggests that the industry is prioritizing the 'harness' around AI inference as the cost of inference decreases.

A new four-part live series, "From Model to Agent: The Agent Framework Harness, Live in C#," demonstrates building a complete C# AI agent using the Microsoft Agent Framework's agent harness. The series covers adding tools, memory, skills, observability, and deployment to an agent, providing a practical guide for developers.

Harness has rebuilt its Git repository and introduced a new AI Code Review product to address the surge in pull requests generated by AI coding agents. This development aims to help engineering teams manage the increased volume of code contributions and prevent bottlenecks in the development pipeline.

The effectiveness of AI agents in real-world applications relies heavily on the surrounding 'harness' or scaffolding, which manages inputs, checks outputs, and handles failures, rather than solely on the underlying language model's reasoning capabilities. This is important because demos often overstate an agent's readiness by operating under ideal conditions, while production environments require systems to manage incomplete data, tool errors, and policy changes.

Researchers at Meta AI and the University of Illinois Urbana–Champaign developed EvoHarness-RL, a framework that enables AI agents to autonomously manage complex enterprise workflows by learning when to interact with their environment. This framework addresses the limitations of current AI agents that rely on rigid, human-written rules for task execution, improving their ability to handle dynamic situations and recover from errors.

Recent benchmarks indicate that the software harness steering AI coding agents significantly impacts token usage and cost, even when using identical underlying models. This suggests that harness efficiency is as crucial as model selection for optimizing AI coding agent performance and expense.

Harness engineering is a methodology for managing AI-assisted code generation to ensure code remains correct and consistent over time. It addresses the tendency of AI assistants to introduce inconsistencies and drift from established conventions in a codebase. This approach provides a framework for integrating AI coding tools without compromising code quality.

Nvidia research indicates that the software 'harness' surrounding an AI model is more critical than the model itself for long-horizon tasks, enabling significant performance improvements. This finding suggests that memory management and supervisory components are key to developing effective AI agents capable of complex, multi-step operations.

DeepSeek has released DeepSeek Harness (dsh), an open-source execution runtime for building autonomous AI agents, available in developer preview under the MIT license. This release introduces a modular, micro-kernel architecture that allows developers to interchange components like model adapters and sandboxing environments, signifying a shift towards unbundled infrastructure for AI agents.

TrueFoundry, a machine learning startup, has released TrueForge, an open-source AI agent harness under the MIT License. The company claims TrueForge can reduce task completion costs by 30-75% compared to Anthropic's Claude Managed Agents, offering greater developer control and vendor neutrality for AI agent deployment.

DeepSeek has open-sourced the DeepSeek Harness, a new Node.js-based AI agent runtime available as a developer preview under an MIT license. This harness distinguishes itself by treating all components, including the model adapter and agent loop, as replaceable plugins, allowing for dynamic composition and extension.

DeepSeek released DeepSeek Harness v0.1, an open-source agent harness, and DeepSeek-V4-Pro, an updated flagship AI model focused on agentic workloads. The company also announced a new peak and off-peak pricing structure for its API, which will result in higher costs for V4 access. These releases mark DeepSeek's expansion into developer tools for AI agents, competing with existing integrated coding environments.

DeepSeek Harness is now available in developer preview, providing a framework for building AI agents with a plugin-based architecture. This release allows developers to customize agent capabilities like models, tools, and UI, and offers full traceability of agent operations.

DeepSeek AI has released DeepSeek Harness (dsh), an open-source agent harness built with a plugin-based architecture and powered by Cordis. This tool is currently in developer preview, indicating ongoing development and potential for breaking changes.