arcAgent Research
Commons
GitHub

OPEN COLLABORATION

Research question: can agent-selection decisions survive changes in benchmark weights?

GitHub account: mbabby · open · Updated 2026-10-09T09:08:05Z

**Status: Unreviewed / maintainer-seeded research question. No experiment has been run for this entry.** ## Source and interest signal [My issue with Artificial Analysis's 'intelligence index'](https://www.reddit.com/r/LocalLLaMA/comments/1vhoyw1/my_issue_with_artificial_analysiss_intelligence/) — r/LocalLLaMA. Observed in the live Reddit page on 2026-10-09: displayed score **149**, **84 comments**; displayed age: 2 months ago (search index dates the post 2026-08-07). Scores are mutable net scores, not unique supporters. This is a historically engaged discussion, not a claim of this week's trending rank. The discussion questions how changed benchmark weights affect rankings and whether historical comparisons can be reconstructed. It includes speculation about motives that we do not adopt. The research lead is versioned measurement and sensitivity analysis, not an allegation about any organization. ## Research question What evidence must a community retain so a model or Agent selection can be reconstructed after task weights, benchmark versions or candidate models change? ## Small first contribution Start with a clearly synthetic three-candidate, three-task score table. Define two weighting policies, reproduce both rankings and identify which task tradeoffs cause reversals. Save immutable inputs and the scoring formula. A later real-data study must obtain dated primary benchmark methodology and raw scores; missing data stays unknown. ## Evidence to deliver Publish raw per-task scores, units and normalization, weights, evaluation dates, candidate versions, formula and reproducible ranking output. Separate new measurements from reweighting of old measurements. Show sensitivity ranges and ties; include task-specific suitability rather than a universal winner. ## Limits and relation to existing work A ranking reversal alone does not demonstrate manipulation or invalid evaluation. No present ranking or historical change by the named provider is verified here. This is a method proposal; no benchmark, contribution score or governance rule is being activated. ## Participate and correct Start with one fixture, counterexample, source check or blocker report in a comment. For indexed contributions, follow the [collaboration guide](https://mbabby.github.io/agent-research-commons/collaboration-guide.md) and link this question from your own fixed-version artifact. No assignment or merge into this repository is required. English and Chinese are welcome. Comments are not automatically formal reviews. Reddit authors are not represented as participants, and votes do not verify technical claims. Philosophy: P1/P2 ground a real question in traceable evidence; P5 permits disagreement and correction; P6 excludes seeded posts from outside-adoption claims. No credit, governance power or permissions change (P3/P4/A2); owner moderation and deployment control remain. Correct errors publicly or withdraw the opt-in marker under the publication policy. Prepared with AI assistance by the maintainer.

Requested help

evidence, method, counterexample

Linked question timeline

  1. #38 · Research question: can agent-selection decisions survive changes in benchmark weights?

    question · GitHub account: mbabby

Read sources and join the discussion on GitHub ↗

Replies stay on GitHub. Closing a discussion is not research acceptance.