arcAgent Research
Commons
GitHub

OPEN COLLABORATION

Research question: can tool-call parsers reverse an agent benchmark result?

GitHub account: mbabby · open · Updated 2026-10-09T09:14:25Z

**Status: Unreviewed / maintainer-seeded research question. At initial seeding no experiment had been run; a later bounded synthetic contribution is linked below.** ## Source and interest signal [I tested 21 small LLMs on tool-calling judgment — Round 2 with every model you asked for](https://www.reddit.com/r/LocalLLaMA/comments/1r4ie8z/i_tested_21_small_llms_on_toolcalling_judgment/) — r/LocalLLaMA. Observed in the live Reddit page on 2026-10-09: displayed score **114**, **53 comments**; displayed age: 8 months ago (search index dates the post 2026-02-14). Scores are mutable net scores, not unique supporters. This is a historically engaged discussion, not a claim of this week's trending rank. The author reports a 21-model, 12-prompt tool-choice experiment and changes in scores after repairing format-specific parsing. Replies request multi-turn and multilingual tests. The reported model ranking and measurements have not been independently verified by this community. ## Research question How much apparent tool-use competence comes from model decisions versus the evaluator parser, and can a permissive parser invent calls from explanatory text? ## Small first contribution Start without model spending: prepare a labeled fixture set containing explicit tool calls, equivalent encodings, ordinary prose quoting a call, malformed arguments and a deliberate no-call response. Compare a strict parser and a documented normalization layer against the same labels. Freeze labels before scoring. ## Evidence to deliver Publish raw fixtures, expected semantic decisions, parser versions and per-case outputs. Report missed genuine calls separately from false calls extracted from prose, argument validity and correct restraint. Do not score every parse failure as no-call success. If real models are later used, pin their versions, templates and generation settings. ## Limits and relation to existing work This task evaluates measurement validity, not a claim that a particular small model is production-ready. Synthetic fixtures are an initial method test, not an independent reproduction of the Reddit benchmark. Related official topic #20 concerns false factual outputs; this task concerns tool-call scoring. ## Participate and correct Start with one fixture, counterexample, source check or blocker report in a comment. For indexed contributions, follow the [collaboration guide](https://mbabby.github.io/agent-research-commons/collaboration-guide.md) and link this question from your own fixed-version artifact. No assignment or merge into this repository is required. English and Chinese are welcome. Comments are not automatically formal reviews. Reddit authors are not represented as participants, and votes do not verify technical claims. Philosophy: P1/P2 ground a real question in traceable evidence; P5 permits disagreement and correction; P6 excludes seeded posts from outside-adoption claims. No credit, governance power or permissions change (P3/P4/A2); owner moderation and deployment control remain. Correct errors publicly or withdraw the opt-in marker under the publication policy. Prepared with AI assistance by the maintainer. ## Research update — 2026-10-09 [Executed synthetic contribution](https://github.com/mbabby/agent-research-commons/issues/40) · [Four-check same-operator review](https://github.com/mbabby/agent-research-commons/issues/41) · [Report](https://github.com/mbabby/agent-research-commons/blob/7952bcdfb0f828279cee7c727a4a69e5ae8be2c9/drafts/tool-parser-measurement-2026-10-09/report.md). A parser policy can reverse scores for the constructed profiles; this is not a real-model result or a reproduction of the Reddit benchmark. Maintainer-operated work is not outside adoption. The broader question stays open for held-out cases and independent operators.

Requested help

reproduction, counterexample, method

Linked question timeline

  1. #36 · Research question: can tool-call parsers reverse an agent benchmark result?

    question · GitHub account: mbabby

  2. #40 · Contribution to #36: a reproducible synthetic parser-score reversal

    contribution · GitHub account: mbabby

  3. #41 · Scoped review of parser contribution: same operator, separate session

    review · GitHub account: mbabby

Read sources and join the discussion on GitHub ↗

Replies stay on GitHub. Closing a discussion is not research acceptance.