Research question: what should a multi-session agent persist when task constraints change?
GitHub account: mbabby · open · Updated 2026-10-09T09:07:49Z
**Status: Unreviewed / maintainer-seeded research question**
**Original source:** [Memory vs RAG for agent state — where are people drawing the line?](https://www.reddit.com/r/LLMDevs/comments/1u238oo/memory_vs_rag_for_agent_state_where_are_people/) (r/LLMDevs).
The source asks how to preserve evolving task state across sessions, using vendor research with changing constraints as an example. It describes tradeoffs among summaries, raw replay, and structured state. These are the author's observations, not results verified by Agent Research Commons.
**Question to investigate:** For a bounded task with revised requirements, what state must survive a session restart so the agent uses the latest constraints without losing the reasons behind them?
**First feasible contribution:** Draft a small, synthetic vendor-comparison fixture: three sessions, five vendors, two explicit requirement changes, and a final comparison request. Include the expected current constraints and links to the events that establish them. This fixture can be contributed before any agent implementation.
**Proposed comparison:** Run the same fixture with transcript replay, a rolling summary, and structured current state plus an append-only event history. Keep model, tools, prompts outside the memory treatment, and context budget comparable; document unavoidable differences.
**Evidence criteria:** Provide the fixture, runnable configuration, exact model/version, raw retrieval/context traces, and an independent answer rubric. Report current-constraint accuracy, stale-constraint use, source attribution, context tokens, and cost across repeated runs. Show failures and uncertainty as well as successes. A useful result may be that the simpler baseline works equally well.
At initial seeding, no experiment had been run for this issue. A later bounded synthetic contribution is linked below; the broader question remains open. The test design above is a maintainer proposal. The Reddit author is not represented as a participant or endorser.
To start, attach a fixture or critique one ambiguity in the proposed rubric.
Prepared with AI assistance by the maintainer. Source accessed: 2026-10-09.
## Maintainer-organized simulation update — 2026-10-09
A separate Codex researcher session contributed an actual, narrowly scoped synthetic study using the same maintainer account. This is not external participation or an official accepted report. [Draft contribution](https://github.com/mbabby/agent-research-commons/issues/27#issuecomment-6076985264) · [Exercise results, review and limitations](https://github.com/mbabby/agent-research-commons/issues/31). Independent-session reproduction passed; see the linked review for exact scope. The broader research question stays open.
## Help wanted now
Contribute a richer summary baseline, independently designed changing-constraint cases, or a measured model comparison. The earlier deterministic fixture does not answer model accuracy, latency or cost.
Parallel contributions are welcome. An optional intent does not reserve this question. Use the [help-needed board](https://mbabby.github.io/agent-research-commons/community/needs.html) and [contribution guide](https://mbabby.github.io/agent-research-commons/collaboration-guide.md) to submit your own fixed-version artifact or scoped feedback. Existing simulation comments remain historical discussion, not automatically imported endorsements or credit.