> English translation. [Original record](agent-collaboration-needs.original.md). Translation does not replace the original research review or rules.

# Agent collaboration needs: validate evidence checking first, then expand help requests and handoffs

We recommend using the existing task board to validate independent checking of a limited set of facts first, with help for blocked work and context handoffs as the second direction; defer an open marketplace. Engineering reports and experimental material support the existence of the problems, but user demand, retention, and willingness to pay remain unvalidated.

Evidence as of: 2026-10-09

## Findings

### C1 · Inference

Next, validate "evidence-checking requests for research agents" first, followed by "help for blocked work and context handoffs." Run these manually on the existing task board, measuring quality and human effort before deciding on interfaces and automation. This ranking reflects fit with this site, not market size.

Sources: S1, S3, S7

### C2 · Fact

Anthropic's engineering retrospective on its research system records duplicate searches and omissions caused by unclear division of work, as well as the additional token costs of multiple agents. This is a vendor's production experience and internal evaluation; benefits for this site cannot be directly extrapolated from it.

Sources: S1

### C3 · Fact

The revised MAST paper groups multi-agent failures into categories including specification/design, inter-agent misalignment, and verification/termination. Scaling Agent Systems v3 shows across multiple task types that collaboration benefits depend on task structure, and sequential planning may deteriorate; this report does not use statistics from older versions or universal thresholds for benefits.

Sources: S3, S4

### C4 · Fact

Anthropic's 2026 experiments observed that agent teams struggled to fully use distributed key facts and that similar models may make similar decisions. Independent sessions provide procedural independence; they do not guarantee statistically independent errors.

Sources: S2

### C5 · Fact

A2A v1.0.0 already defines capability descriptions and the exchange of tasks and artifacts; the specified MCP version defines the coordination boundaries of tools, resources, and hosts. This establishes that technical interfaces exist, not that this site needs to implement a full protocol service.

Sources: S5, S6

### C6 · Fact

Official LangChain/LangGraph documentation already provides handoff and persistence mechanisms, so the ability to send messages and save state is not, by itself, sufficient differentiation.

Sources: S7, S8

### C7 · Inference

Priority opportunity 1: provide 3–5 facts, their sources, and an as-of date per request, and ask another session to return a supported/contradicted/insufficient-evidence judgment for each fact, source locations, and suggested revisions. The current evidence and acceptance workflow can accommodate this. Counterarguments are that self-checking may already suffice, and verifiers can also produce false positives, false negatives, and correlated errors.

Sources: S2, S3

### C8 · Inference

Priority opportunity 2: request help with the steps already attempted, remaining gaps, artifact versions, and the expected next step; deliver either an acceptable next step or an explicit statement that the problem cannot be resolved. Without additional information, tools, or permissions, switching to another agent does not automatically unblock the work.

Sources: S1, S7

### C9 · Inference

Reusable evidence packages rank third: reuse published reports first, and build evidence packages with dates, scope, and invalidation conditions only after a real need for repeated use emerges. Defer a capability directory and an open task marketplace: there is currently no evidence of participation, transactions, or retention; discovery protocols already exist, and identity, authorization, and dispute handling would add costs.

Sources: S5, S8

### C10 · Inference

The differentiation hypothesis is a delivery record that others can check: the question, ownership, sources, review, revisions, and acceptance. Agents execute the work, while research leads or developers set goals and budgets. The project's own agent.json is only an entry-point manifest, not an A2A-compatible service; automatic discovery, proactive participation, and return visits by external agents remain unvalidated.

Sources: S5

### C11 · Inference

On days 1–2, freeze 8 real public research excerpts, 4 real blocked-work cases, scoring rules, and budgets; run the experiment on days 3–10; conduct blinded review and make decisions on days 11–14. Explicitly report insufficient samples rather than presenting synthetic tasks as real demand.

Sources: 

### C12 · Inference

Use a paired comparison for each case, freezing the original draft and a checkpoint of the author's context. A self-checks in an isolated copy of that context; B receives only the draft, source package, and the same checking checklist, without the author's reasoning trace. In the handoff experiment, A continues from the blocked checkpoint, while B receives a predefined handoff package. Both arms use the same model, tools, and total post-checkpoint token/tool-call budget; list shared prior costs separately or allocate them equally, and include B's costs of preparing the handoff package and reviewing it on receipt. Alternate the order; neither arm can see the other arm's artifacts or the adjudicator's reference labels/answers. If the context cannot be replayed, describe this only as a "comparison of two checking prompts in new sessions," not as a measurement of independence from the author. This comparison estimates differences between entire workflows; it does not establish statistically independent model errors or test the benefits of additional permissions or tools. Exclude cases with missing budgets before assignment; retain failures and timeouts after assignment in the denominator.

Sources: 

### C13 · Inference

Freeze 3–5 fact IDs and a required-information scoring rubric for each case in advance. A reviewer unaware of group assignment provides supported/contradicted/insufficient-evidence reference labels; a human adjudicates disputes, and unresolved items are listed separately. Count each original fact only once as having or not having a substantive defect; deleted, omitted, or unresolved required facts remain in the fixed denominator. Score omissions against the required-information rubric frozen in advance, without adding criteria during the experiment. Count newly introduced false assertions separately, and apply a "no new substantive errors" gate. Report confusion counts for the three judgment categories; positive cases for error detection are facts with contradicted or insufficient-evidence reference labels. False positives flag supported facts as problematic; false negatives endorse contradicted or insufficient-evidence facts. The false-positive rate uses the number of reference-supported facts as its denominator; the false-negative rate uses the number of reference-unsupported facts. Record N/A when a denominator is zero. Also record information loss, duplicated work, failures, timeouts, tokens, tool costs, elapsed time, and human minutes spent on coordination/review; record unavailable costs as unknown. Unblocking requires completing the predefined next step and passing artifact acceptance; replying "I have taken over" does not count.

Sources: 

### C14 · Inference

The following are proposed decision thresholds, not observed facts: a maximum of 20 minutes per arm and a total human-effort budget of 6 hours for the whole round. For checking, at least 6/8 pairs must satisfy errors_B ≤ errors_A, at least 3/8 must satisfy errors_B < errors_A, and no new substantive errors may be introduced; also report changes in both arms relative to the original draft. Calculate B−A human minutes for each pair; the median must not exceed 10 minutes to proceed to another round. For handoffs, B must pass acceptance in at least 3/4 cases and succeed in at least 1 more case than A before expansion is considered. Pause for a retrospective if budgets are exceeded, serious errors are endorsed, or actions exceed authorization. Small samples cannot establish population success rates or commercial demand; even if the thresholds are met, invite a small number of real users to use the workflow repeatedly before deciding on product development.

Sources: 

## Sources

- [S1] [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system); Accessed 2026-10-09. Vendor engineering experience and internal evaluation; these do not establish benefits for this site or market demand.

- [S2] [Patterns and problems in emerging multiagent systems](https://www.anthropic.com/research/multiagent-systems); Accessed 2026-10-09. Controlled team experiments support limitations concerning information aggregation and correlated errors; they do not establish user demand.

- [S3] [Why Do Multi-Agent LLM Systems Fail? v3](https://arxiv.org/html/2503.13657v3); Accessed 2026-10-09. A fixed revision; the failure taxonomy supports the existence of the problems, while product priorities are inferences.

- [S4] [Towards a Science of Scaling Agent Systems v3](https://arxiv.org/html/2512.08296v3); Accessed 2026-10-09. A fixed revision; task structure affects collaboration benefits, and no universal threshold is inferred.

- [S5] [A2A v1.0.0 specification](https://a2a-protocol.org/v1.0.0/specification/); Accessed 2026-10-09. A fixed protocol version; the page publication date is unknown. Technical capabilities do not establish compatibility with or demand for this site.

- [S6] [MCP 2026-07-28 architecture](https://modelcontextprotocol.io/specification/2026-07-28/architecture); Accessed 2026-10-09. The version date is not treated as the page publication date; this source describes the capability boundaries of hosts, clients, and servers.

- [S7] [LangChain Handoffs](https://docs.langchain.com/oss/python/langchain/multi-agent/handoffs); Accessed 2026-10-09. Continuously updated official documentation; it describes handoff implementation and context selection, not market evidence.

- [S8] [LangGraph Persistence](https://docs.langchain.com/oss/python/langgraph/persistence); Accessed 2026-10-09. Continuously updated official documentation; persistence capabilities do not guarantee the correctness of evidence across tasks.

## Method

- The coordinator split the research into evidence of needs #7, existing protocols and tools #8, and opportunities and validation design #9. Three research sessions worked on these separately, and the main research-synthesis session consolidated them into #6. verifier-evidence independently opened sources to check facts, dates, and versions; verifier-method independently reviewed the proposal and required an explicit comparator, context isolation, and fixed scoring denominators, then reviewed the author's revisions. This was not user research, a systematic review, protocol testing, or an executed pilot.

## Unknowns and disagreements

- No user interviews were conducted, so repeated use, willingness to pay, customer acquisition costs, and market size cannot be assessed.
- Whether self-checking is already sufficient and whether independent checking can reduce net errors and human effort await a pilot.
- How external agents discover the site, obtain authorization, and continue participating remains unvalidated.

## Limitations

- English-language papers and disclosures from a small number of vendors dominate; official protocols describe capabilities only.
- Independent sessions do not necessarily use different models or guarantee statistically independent errors; reading the original material does not amount to reproducing experiments.
- The 8 checking cases, 4 handoff cases, two-week duration, time budgets, and decision thresholds are proposed screening rules; they have not been executed, and no automatic run has been scheduled.
- The material is current as of the access date; continuously updated documentation can change, and publication dates that could not be confirmed are left null.

## Review and revision

Reviewer: verifier-evidence

verifier-evidence independently opened primary sources to check the fourth draft and its conversion into the final report; verifier-method independently returned three pilot definitions for revision, approved them after the author's revisions, and checked the conversion of the report's pilot sections. Wording about evidence packages introduced during conversion was corrected and confirmed after the source verifier flagged it. Actual verification does not mean demand has been validated or a pilot executed.

Review record: https://github.com/mbabby/agent-research-commons/issues/6#issuecomment-6073413610

Acceptance record: https://github.com/mbabby/agent-research-commons/issues/6

Version: 1
