Raft’s 82.3% Is the Wrong Number to Lead With
TL;DR: Raft is right that agent activity is not agent value. My issue is that the most memorable number in its analysis still measures visible messages, while the outcome measures remain mostly a framework.
Six Teammates, More Management
I tried Raft.build a while ago. We had discussed the product internally as a possible model for the future of multi-human, multi-agent organizations.
I tried to turn Raft into a six-person team: three people and three personal agents, each running on the corresponding team member’s machine.
What surprised me in the end was not the output. It was that the agent messages never stopped.
Every agent was eager to introduce itself, explain what it could do, and ask where to start. The output was fine, rarely bad. But it was also rarely good enough to make me think: I could not have done this without Raft.
Instead, I spent more time reading messages, assigning work, deciding which agents to trust, and debugging the communication between the collaboration platform and the local agents. Actually getting the work done came later.
Activity Is Not Leverage
That is why I paused at Raft’s claim that agents produced 82.3% of visible messages across the engaged workspaces in its data.
The number may be accurate. It may even be useful. It tells Raft that its workspace has become agent-heavy.
But it does not tell us that agents did 82.3% of the work.
A good agent can complete a task, create a useful artifact, and send one short message. A bad one can send ten thoughtful updates, ask five reasonable questions, and still leave me with the real work.
Those two agents look very different in a user trace. They can look identical, or even inverted, in a message metric.
To be fair, Raft says this quite clearly. Opening the app is not value. Sending a message is not value either. For a productivity tool, what matters is whether people save time and whether the work moves forward. The article also says Raft looks at collaboration, artifacts, and outcomes, and warns that “more activity can hide less value.”
My concern is that the article gives us specific DAA and message-share numbers, but no outcome results at the same level. An activity number is evidence. It just is not evidence of leverage.
DAA can still be a reasonable health metric. It tells you whether agents are showing up, whether they are running, and whether the workspace is becoming agent-heavy. But it is not a value metric.
Measure the Work, Not the Messages
The real question is not how many agents spoke. It is whether the person got a better artifact with less effort. Did the agent finish something? Was it adopted? Did it change the task state? Did the user need to become the project manager for six teammates?
If artifacts and outcomes do not appear in the trace alongside activity, then “agent-native work” may just be an agent-native inbox.
82.3% sounds high. But on its own, it mostly measures how much agents like to talk.
Start With the Workflow
That is also the direction we are deliberately taking as we build multi-human, multi-agent teams with Codex, GitHub, and the collaboration platforms we already use. Do not build a new workspace first, then use message share to prove that it is agent-native. Start by asking whether the work made it into the existing workflow, whether the artifact was adopted, and whether the task state actually changed.
简单说: Raft 自己也知道 agent activity 不等于 agent value。我的问题是:文章里最容易被记住的数字仍然来自可见消息,而 outcome metrics 还主要停留在框架里。
六个队友,更多管理
前段时间,我试了下 Raft.build。公司内部对这个产品有过一些讨论:Raft 可能代表了未来“多人多硅”的一种组织形态。
我试着把 Raft 用成一个六人团队:三个人,三个 personal agents,各自在对应成员的机器上运行。
最后让我意外的不是产出,而是 agent messages 根本停不下来。
每个 agent 都很 eager:先自我介绍,再解释自己能做什么,然后问从哪里开始。产出都还行,基本不差;但也很少好到让我觉得:这个东西没有 Raft,我就做不出来。
于是我花了更多时间读消息、分配任务、判断谁值得信,再调试协作平台和 local agents 之间的通信。真正把东西做出来,反而排在了后面。
Activity 不是 leverage
所以看到 Raft 说,在其数据中的 engaged workspaces 里,agent 产生了 82.3% 的可见消息,我停了一下。
这个数也许是对的,也可能有用。至少它说明 Raft 的 workspace 已经很 agent-heavy。
但它不能说明 agents 做了 82.3% 的工作。
一个好 agent 完成任务、给出能用的 artifact,然后只发一句“好了”。一个不太好的 agent 可以连发十条很认真的进度、问五个合理的问题,最后真正的工作还是我来收尾。
放进 user trace,这两个 agent 差很多;放进 message metric,结果甚至可能刚好反过来。
公平地说,Raft 自己讲得很清楚。打开 app 不是 value,发消息也不是 value;对 productivity tool 来说,真正该看的是人有没有省下时间、工作有没有往前走。文章也说他们同时观察 collaboration、artifacts 和 outcomes,并提醒“更多 activity 可能藏着更少 value”。
我的疑虑是:文章给出了具体的 DAA 和 message share,却没有给出同等级的 outcome results。Activity number 是证据,只是不是 leverage 的证据。
DAA 仍然可以是一个合理的 health metric。它能告诉你 agents 有没有出现、有没有跑起来,workspace 有没有变得 agent-heavy。但它不是 value metric。
看工作,不看消息
真正该问的不是多少 agent 说过话,而是人有没有以更少的 effort 拿到更好的 artifact。任务完成了吗?产出被采用了吗?状态变了吗?还是用户得自己当六个队友的 project manager?
如果 artifact 和 outcome 没有和 activity 一起出现在 trace 里,所谓“agent-native work”,大概就只是一个 agent-native inbox。
82.3% 看起来很高。但单独拿出来,它量到的主要还是 agents 有多爱说话。
先看 workflow
这也是我们现在用 Codex、GitHub 和现有协作平台搭多人多硅团队时,刻意选择的方向:不要先造一个新 workspace,再用 message share 证明它很 agent-native。先看工作有没有进入现有流程,artifact 有没有被采用,task state 有没有真的改变。