Mouten

The numbers passed. The claim still did not

The numbers passed. The claim still did not

I built a pipeline that writes a technical article and then inspects it.

The setup

A human prepares the source. A writer turns it into a draft and a list of what it did not check. A verifier opens primary pages and checks the numbers and the citations. A reviewer writes the objections a reader would bring, and the ways the draft can be misread. A commander marks PASS, REVISE, or REJECT. A human decides whether to publish. The cap is three revises. If the draft does not pass on the third, it is discarded. I ran this on a GrokBot cloud VM. GrokBot launched on 2026-08-11. Bots sat on that VM: a commander, writers, a verifier, and a reviewer. I gave the verifier and the reviewer different jobs. The verifier ends on what has a primary source and how long that source has been there. The reviewer ends on what the draft is pretending to be. I am not writing this as an argument that an agent can write. I am writing the record of an inspection.

What happened first

The writers sent up three drafts, one per outlet. They looked good. They also attached a list I had not asked for. The numbers were copied from the brief and had not been re-measured. Source pages had not been opened. The review text had not been read. That self-report is what made the later stages possible. I did not know that yet.

Where it stopped

The verifier could not find two sources. The draft said "an independent review said" in the third person, as a fact. The verifier could not identify that article. A nearby gist existed. The path the draft named was not in it. The phrase that it was only writing was not in it either. The other missing source was a count of 311 security issues. No page showed 311. A search returned 560. Six numbers were off.

  • GitHub stars in the draft: 235,611. Same-day check: 235,777
  • Comparison stars: 387,363. Check: 387,433
  • Messaging channels: 29. Check: 31
  • The product's channels in the draft: 5 plus CLI. The official comparison table: 28 rows
  • Security issues: 311. Search: 560
  • Stars in the last week: +26. That increment did not reproduce. Third-party data showed +3k for the week One number had the wrong meaning. A security firm's report had 5,900 in it. The draft wrote 5,900 admin panels. The primary page says 5,900 events, sightings. It does not say distinct hosts. The reporting side said it was counting sightings rather than distinct machines. The reviewer named what the draft was wearing. Verification as a costume for a conclusion already made. The opening "I wanted to know if I should switch" was the costume. The title already kept the stack. The role as an escape hatch was decided first. The reviewer also wrote the comments a reader would post.

You checked the star count three ways and still published a number that does not match GitHub the same day. Which independent review says it? Link it. Then link the +26 week and the 311 security issues.

It listed 14 places a reader could misread. If I had been the one reading that draft, I would have missed that the star count had moved the same day, that 5,900 were events, and that the independent review had no source. That was the first send-back.

Finding out why

The first send-back was the two missing sources, the six number mismatches, and events written as panels. The second was a closing generality, and an incident that did not belong in the piece. The third was a title that stated the claim before the measurements. After the third resubmit, verification passed. Missing sources: 0. Number mismatches: 0. Those first-round catches were gone. They are not why the commander discarded the piece. The commander wrote that the numbers had passed and the shape of the claim had not. The source claim was that necessity is thin. What had been measured was an install, 1+1, and a 706-character summary. Cut the unsupported parts and an install log is left. That original claim cannot be made from this source. Discarded.

What I changed

I lowered the claim. "The product is unnecessary" became "the install hit five stops." I cut from the source everything that had not been measured on that run. Star counts. A judgment of a learning loop. A security incident. Comparisons. What remained was the install log. The second article passed after one revise. Separately, I changed the reviewer's brief. The first brief said to produce three objections. A fixed count means every draft gets three. If the reviewer marks them legitimate, the draft fails. The new brief is: if there is a legitimate objection, write all of them. If there is none, write none. Legitimate means the main claim of the article does not survive the objection. Side issues go under misreads. They do not go in with the objections. That brief change is not why the first article was discarded.

What was left

The first-round catches are what I would have missed if I had been the one reading. They were already fixed when the piece was discarded. What was discarded is the claim. The source wanted necessity is thin. The run had an install, 1+1, and a 706-character summary. The writing was not what was discarded. The source did not support the claim.

日本語のブログ