Asking Dot to organize five agent stories: what the first draft missed and what review kept
One information-organizing run on five previously researched sources produced 20 initial claims and 25 revised claims. AI review corrected wording, disclosures, a missing log event and the attribution of human input. The user confirmed the content; the source tasks were not reproduced.
- Published
- Oct 6, 2026
- Reviewed
- Oct 5, 2026
- Content updated
- Oct 6, 2026
- Applicable version
- Not pinned; check current documentation
- Platforms
- Dot
Before you start
- A single fixed-source information task; no pinned product version or third-party task replication.
An agent story can contain a user's task, an author's account of the process and a visible output. These are different kinds of evidence. BotClaw asked Dot to organize five fixed public sources once, then had the text checked in a separate AI review. This case records what needed correction and what the resulting evidence can support. It did not operate Muse, run ComfyUI or repeat the game-development task.
What was tested
The coordinator selected the five URLs from material researched earlier, so this was not a blind test or a random sample of agent use. The executor reread the fixed sources after the protocol file was created instead of filling gaps from earlier summaries. Only one organizing run was performed; later AI checks were reviews, not additional experimental runs.
The protocol listed eight acceptance criteria for source access, case fields, attribution, evidence boundaries, unsupported outcomes, readable deliverables, action logs and honest timing. The test could assess faithful organization of these materials. It could not establish that Dot could perform the same purchases, produce a high-quality video or deliver a commercially usable game.
What review changed
The first ledger contained 20 claims. The first AI audit supported 16 and recommended clearer wording for four; its overall decision was revision required. Seven of the eight original criteria passed, while the action-log criterion needed a missing save event. Disclosures were also added as an editorial improvement, without retrospectively changing the eight criteria.
C07 now names a speaking agent, avoiding confusion with the neighboring insurance task. C12 states that WorkBuddy wrote all of the article's text. C18 says payment still had to be completed by the user and that the eventual payment status was unknown. C19 narrows the vague first summary to the newsletter-growth research rather than the whole PDF.
The revision added five disclosure claims, C21–C25, bringing the ledger to 25. It distinguished a Reddit relative post time from crawler metadata, retained an unknown absolute post date and added the 15:12:02 save event as an explicitly retrospective log entry. These changes did not verify the real-world outcomes behind the claims.
The follow-up AI review found one more attribution problem: the draft treated the approval, all URLs, scope and acceptance criteria as human-provided. The correction separated the user's approval from the coordinator's source selection and communication of the criteria. Actual human involvement and working time were not independently measured. The revised text then passed the eight information-organizing criteria in AI review.
How to read the five cases
The following cards retain the task, inputs, process, result, human participation, value claim and limits. A means a page or artifact state directly visible in the original run; B means an author's first-person account; C means a secondary account or opinion; U means insufficient evidence. A does not prove how an artifact was created or that it works. Claim IDs refer to claim_ledger_v2.json in the evidence ZIP.
S1 · Muse and family administration
Claire Vo's ChatPRD article is dated September 16, 2026. The tasks covered a family briefing from calendars, email and user requests, plus goals and shopping [C01]. The author reported that, after authorization, Muse produced a family PDF. The movie-ticket flow reached a payment page, but no tickets were bought [C02–C03]. Human participation included connecting sources, granting permission, adding requirements, and choosing the showing and ticket count.
Evidence level B: the author liked the briefing design, but the PDF original and full demonstration were not checked. Finding a shoe color and browser operation met obstacles [C03–C04]. No measurable time saving or revenue was supplied; a payment page is not a completed purchase.
S2 · Dots and office documents
Casey Newton's Platformer article is dated September 29, 2026 in the site's article listings; the parsed article body did not display the date [C05]. The account describes work email and an insurance questionnaire using a lease, budget and municipal material. Dots searched, cross-checked and asked follow-up questions; two questionnaire answers remained open. It also prepared meeting and event materials [C06–C07]. The user answered questions, decided the final wording and approved an email to a speaking agent.
Evidence level B: private business files were unavailable and the answers were not independently recalculated. The author estimated about 15 minutes of participation for work that would otherwise take two hours [C08]. This was a personal estimate from a short trial without a timed control, not a measured or general efficiency gain.
S3 · Dots and a ComfyUI video-extension workflow
Reddit user u/BraveBrush8890 described providing access to ComfyUI on RunPod, first requesting a video and then an extension [C09]. The author said that, after failing to find a reliable existing workflow, Dots built one using the ending frames and audio, trimmed overlap and joined the segments, with three optional extension stages [C10]. Providing the environment and requesting the extension were visible human inputs; further debugging effort is unknown.
The process account is B. A applies only to the accessible Pastebin configuration page and observed loading, optional extension and saving fields [C11]. About two hours including testing was the author's claimed duration, not two hours saved. The original run did not run the configuration or check its dependencies, safety or audiovisual quality. The post's absolute date remains unknown; Pastebin's October 1, 2026 date cannot substitute for it.
S4 · Dots and the One-Key Gravity preview
Alex Xiang's Zicode article, dated September 30, 2026, describes making a mobile-oriented gravity-switching game with Godot [C12]. The user stopped operations on the original machine and moved the task to the cloud. Another device was offline, so synchronization did not finish [C13]. The public GitHub preview page listed test packages, a test report and checksum filenames, but its asset area failed to load and the files were not inspected [C14].
The collaboration account is B and the visible release page is A. The article explicitly credits all text writing to WorkBuddy. Its 2-minute-50-second figure is a recording length, not development time [C12]. The release notes claim a Linux runtime check while leaving Android and Windows untested on actual devices, with a test-signed Android package; the original run did not verify that testing occurred [C15]. No verified time saving or revenue follows.
S5 · Dots, daily tasks and reminders
Pat Simmons's AI for Mortals article is dated September 30, 2026. The fixed URL with a trailing slash initially returned a cache miss. The executor found and read the same path without that slash, recording the change instead of substituting another article. Inputs included email and computer material for shopping, a Montreal trip and work management [C16].
The author reported moving shopping to a local browser after an obstacle. Travel work included room and flight searches and form filling; payment still required the user, and whether payment ultimately occurred is unknown [C17–C18]. The user provided preferences, chose options and authorized actions. The first newsletter-growth research summary was generic; follow-up questions elicited some specific suggestions [C19].
Evidence level B: orders, the private workspace and original bills were not obtained. Finding a previously missed failed-billing alert does not establish money recovered or a bill resolved [C20]. Shopping was incomplete, generated material needed revision and no net time saving or revenue was calculable. Product-capability descriptions remain the author's experience at that time.
Source attribution and disclosed interests
ChatPRD named Optimizely and OpenArt as sponsors; this neither confirms nor rules out Meta or other undisclosed relationships [C21]. Platformer disclosed that the author's fiancé works at Anthropic; that is an interest disclosure, not proof of sponsorship [C22]. Zicode credited all text writing to WorkBuddy and displayed a registration promotion; promotion alone does not establish paid sponsorship [C23].
No explicit sponsorship statement was found in the Reddit and AI for Mortals body text read during the run [C24–C25]. That does not confirm an absence of sponsorship. These disclosures help identify the source context; they do not prove objectivity or the truth of the underlying cases.
Timing and protocol limits
The recorded UTC timeline on October 5, 2026 starts with preparation at 15:05:35, the first protocol-file creation at 15:06:14 and the first source read at 15:06:23. Draft and ledger writing was recorded at 15:11:09, followed by an initial completeness check at 15:11:33. The often-quoted 5 minutes 58 seconds ends at that initial check. It is not end-to-end delivery time, active AI time or human working time.
A later save at 15:12:02 was recovered from file metadata and explicitly added retrospectively; the roughly 6 minutes 27 seconds to that point still excludes later review and document delivery. The first AI review completed at 15:17:40. Revision, follow-up review, rendering and delivery came later; no end-to-end duration is inferred here.
The protocol file existed before the recorded source reads but its time fields were later corrected. There was no immutable pre-retrieval snapshot or external preregistration. The retained v1 snapshot was created at 15:15:50, after source retrieval, and cannot prove that the original protocol was frozen. Logs and file timestamps improve traceability without becoming independent timing instruments.
What remains after review
The result is a traceable information brief: five source-access records, five case cards, a 25-claim revised ledger, two AI-review reports and preserved versions and logs. The strongest evidence concerns source access, attribution and visible public pages. Most accounts of actual operations still come from their authors. Two public artifact pages improve inspectability but do not turn those stories into BotClaw replications.
There was no manual baseline, independent measurement of human effort or repeated trial. No net time saving, ROI, income growth, adoption rate or general success rate is claimed. Full demonstration authenticity, original PDF and business-answer quality, software safety and runnability, real payments, original bills and complete intervention time remain unverified.
The user subsequently confirmed the content and authorized this publication. That confirmation is not an independent human replication or third-party audit. The article follows the approved brief and case draft, retaining the source cards and limitations; the English text is a translation of the same account, not another experiment.
Download and check the evidence
Download the English reading pack from Evidence files below. It provides the complete brief and case draft, protocol, both AI reviews, revisions, both claim ledgers and a guide to all action-log events in English. Start with README.md, then translations/brief_v2.en.md and translations/audit_v2.en.md. Each derivative maps to its original file in manifest.json. This retrospective translation was prepared after the experiment; it is not an original experimental record, a new run or an independent human audit. The unchanged original Chinese archive is included in the pack and remains a separate download.
The original archive was assembled at 15:27:30 UTC on October 5, 2026. Its pending-acceptance and publication statements retain that historical cutoff. Original records, logs, hashes and timestamps are unchanged. The English pack has its own manifest, checksums and later creation time; these do not establish immutable preregistration. Original archive SHA-256: 002aa7ba0f8a9fa6ac88aa60ee2365c0d0ab7c8762b199fc98c202a09c10f4ff. English reading pack SHA-256: 4e7822a4b9c8110db7076c38dde85fa5e7c052ce3ca06fc0be3d60146749f697.
Evidence and limits
Software run recorded
- The recorded run covers AI-assisted information organization only. The five sources came from prior research; this was not a blind test.
- AI review passed eight information-organizing criteria. The user later confirmed the content; neither is an independent human replication or third-party audit.
- No Muse, ComfyUI or game task was reproduced. Most source-case operations remain author accounts; visible artifact pages do not prove their effects.
- No manual baseline or independently measured human time; no time-saving or ROI claim. 5 minutes 58 seconds ends at an initial check, not end-to-end delivery.
- The protocol was created before the recorded retrieval but later had its time fields corrected; there was no immutable preregistration.
Evidence files
Sources
- https://www.chatprd.ai/how-i-ai/meta-muse-review-personal-ai-agent
- https://www.platformer.news/openai-dots-agents-devday-2026/
- https://www.reddit.com/r/StableDiffusion/comments/1wvb236/i_connected_my_chatgpt_dot_to_minimax_h3_and_it/
- https://zicode.com/blog/dot-virtual-machine-one-key-gravity/
- https://www.aiformortals.co/blog/openai-dots
- https://pastebin.com/s3GU1C1D
- https://github.com/ax2/one-key-gravity/releases/tag/v1.0.0-preview
- https://www.platformer.news/author/casey-newton/
Cite or read with tools
https://botclaw.tech/items/case-dot-five-source-review