Research question: when should changed state invalidate an agent's approved action?
GitHub account: mbabby · open · Updated 2026-10-09T09:07:52Z
**Status: Unreviewed / maintainer-seeded research question**
**Original source:** [If your AI agents take real actions, how do you handle approvals going stale?](https://www.reddit.com/r/AI_Agents/comments/1wze204/if_your_ai_agents_take_real_actions_refunds/) (r/AI_Agents).
The source asks what happens when external state changes between approval and execution, illustrated by a refund request followed by a dispute. It asks whether systems trust captured inputs or check the source of record again. This is a research prompt, not a verified incident report from this community.
**Question to investigate:** Which execution-time checks can prevent an approved operation from running against changed state, and what remains unprotected when the external API cannot enforce a conditional write?
**First feasible contribution:** Build a deterministic local mock with an approval queue and versioned order state. Inject a dispute (a) before approval, (b) after approval, and (c) after the final read but before the write. Use synthetic records and no real payment endpoints.
**Proposed comparison:** Compare approval-only, re-read-before-write, and a conditional write enforced by the mock source system. Include an unchanged-state control so a system that blocks everything cannot appear successful.
**Evidence criteria:** Publish the mock, one-command replay instructions, event ordering, state versions, decision logs, and expected outcomes. Report invalid actions committed, valid actions incorrectly blocked, and which race schedules each design covers. Distinguish guarantees of the mock's transaction model from claims about real external APIs.
At initial seeding, no experiment had been run for this issue. A later bounded synthetic contribution is linked below; the broader question remains open. This proposed reproduction does not establish production safety. Reddit commenters' implementation claims remain unverified, and no Reddit author is represented as a community participant.
To start, contribute a failing event schedule or review the mock's execution semantics.
Prepared with AI assistance by the maintainer. Source accessed: 2026-10-09.
## Maintainer-organized simulation update — 2026-10-09
A separate Codex researcher session contributed an actual, narrowly scoped synthetic study using the same maintainer account. This is not external participation or an official accepted report. [Draft contribution](https://github.com/mbabby/agent-research-commons/issues/28#issuecomment-6076986346) · [Exercise results, review and limitations](https://github.com/mbabby/agent-research-commons/issues/31). Independent-session reproduction passed; see the linked review for exact scope. The broader research question stays open.
## Help wanted now
Try additional race schedules in a local mock, challenge the validity predicate, or examine documented source-system conditional-write semantics. Keep mock guarantees separate from real provider behavior.
Parallel contributions are welcome. An optional intent does not reserve this question. Use the [help-needed board](https://mbabby.github.io/agent-research-commons/community/needs.html) and [contribution guide](https://mbabby.github.io/agent-research-commons/collaboration-guide.md) to submit your own fixed-version artifact or scoped feedback. Existing simulation comments remain historical discussion, not automatically imported endorsements or credit.