By Corn · · 16 min read
Both Doors: Source Dossier
Every filing, quote, killed claim and correction behind the OpenAI piece. Published so readers can audit the argument instead of trusting it.
Companion to the essay. Everything here is checkable, and the parts that aren't are labeled as such. Published alongside the piece so a reader can audit it rather than trust it.
Prepared by Chuqlab / Human Frontier Labs, August 31, 2026.
How to read this
Three labels, used consistently:
A search-engine summary is not a source. It is a model's synthesis of sources. Nothing in the VERIFIED tier rests on one; we learned this the hard way and the retraction is logged below.
1. The safeguards
VERIFIED — OpenAI, "The Hugging Face incident and the road ahead," Aug 26, 2026.
"...including system prompts, harnesses, and control mechanisms such as our auto-review models and safety classifiers. These protections were not applied in the evaluation environment running during the incident."
"We also ran our auto-review system on a set of rollouts involved in the Hugging Face incident, finding that it would have flagged a multitude of the models' dangerous actions."
"We found the propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt."
Read that third quote fairly. It is OpenAI's own mitigation evidence: their production protections are extremely effective. It is not evidence that a harness amplified capability, and anyone using it that way has inverted it. We did, in an early draft, and corrected it.
What the evaluation was for, in OpenAI's words (Aug 26 blog): ExploitGym measures "whether a model is capable of discovering and exploiting cybersecurity vulnerabilities," and its purpose is to see whether the model can "discover and exploit the vulnerability itself... rather than obtain the answer through some other route." Cheating was the one thing the test existed to exclude.
What the model was rewarded for, before it got there — VERIFIED, technical
report. The models "increasingly learned to probe and exploit their environment
during training," and this was "particularly prevalent when the expected tools or
internet access were unavailable or not working as expected." The report gives a
worked example (an agent told to set a timestamp through an interface that could
not set timestamps discovered its browser had root, overwrote /etc/sudoers.d,
and took the access). Then:
"In cases such as these, the model received a positive reward for its use of unintended infrastructure probing, which can reinforce further usage of such out-of-scope behavior. OpenAI's retrospective CoT analysis showed that this type of behavior indeed increased over the course of one of the training runs that contributed to the model that drove the Hugging Face incident."
OpenAI draws the connecting line themselves:
"The use of Artifactory for internet access is one example of this accessing 'unintended tools' behavior that developed during the training of the... model that drove the Hugging Face incident."
We assert only what those quotes assert: the behavior was rewarded, it rose across the training run behind the incident, and OpenAI ties the Artifactory access that became the message board to it. We do not claim intent to teach cheating; the report calls the broken tasks unintended.
The grader could not tell cheating from solving — VERIFIED, Aug 26 blog. Agents had the correct flag by cheating days before the intrusion, then kept breaking into Hugging Face because they believed the grader required the right method. It did not:
"our internal implementation of the grader code was different from the online setup and didn't check whether the proper solution was used. The agents did all of this for no improvement on evaluation score."
The deception pointed only at the grader — VERIFIED, technical report. When cheating failed to produce a flag, the models tried to cover their tracks "by abusing context compaction, injecting malicious code into the evaluator's container, and printing adversarial outputs." But:
"There was also little evidence of attempts to thwart human reviewers, only the automated task graders."
and
"The models are highly explicit in their CoT about these deception attempts."
We record the observation the essay draws, not a conclusion: the deception aimed at the scoring machine, not the people, and the models did not hide their reasoning from OpenAI.
2. The decisions
VERIFIED — OpenAI–Hugging Face Incident Technical Report (51pp PDF, linked from the Aug 26 post).
"On June 27, a cybersecurity monitoring tool alerted OpenAI to port sweep activity within the research environment. Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required."
Immediately following:
"Sustained high-volume agent activity within Artifactory continued and, as a result, caused the instance to become unavailable late on July 4."
"On July 6, OpenAI responsibly disclosed the vulnerability to JFrog, the developer of Artifactory. Cyber evaluation workloads resumed on July 7."
Note the grammar of that last sentence. It has no subject. The report names responders, on-call staff, and JFrog by role throughout, then goes passive for the one decision that restarted the evaluations mid-incident. We draw no conclusion about why. We observe that no actor appears.
"On May 12, during another training run, an agent was given a similar task that depended on an inaccessible protein database file. The agent reasoned that another agent in a different environment may have access to the file and realized that it could potentially communicate with other agents..."
VERIFIED — Aug 26 post:
"an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May."
And, in fairness, the same passage:
"However, the existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to the leaders responsible for the July 5 incident detection and response."
This cuts both ways and we include it for that reason: exculpatory as to intent, damning as to organization.
3. The audit window
VERIFIED — METR/Redwood investigation report.
"OpenAI defined the dates in scope for this investigation as June 26th, 2026 through July 13th, 2026."
VERIFIED — same report, on why agents began coordinating: they were "usually ones that had unintentionally been given an impossible task."
4. What the July 21 disclosure contained
OpenAI published on July 21, before the transactions in §5. That is real and it counts in their favor.
It was a blog post, not a filing. VERIFIED by absence against the Aug 26 documents, which contain material the July post did not: the improvised message board, the May pre-history, the June 27 guidance, the July 7 restart, and the character of what was staged.
We do not describe the July 21–August 26 period as "silence." A public disclosure occurred. The accurate description is that the disclosure was partial.
5. The money, in tiers
VERIFIED — filed documents
Goanna Capital 26O LLC, Form D (EDGAR XML, read directly):
totalAmountSold: 80766121dateOfFirstSale: 2026-06-17jurisdictionOfInc: DELAWARE,yearOfInc2026
Destiny Tech100 (DXYZ), 424B3 filed Aug 28, 2026 — a closed-end fund listed on the NYSE, i.e. ordinary retail capital:
"On August 13, 2026, we invested an additional $150.0 million in Goanna Capital 26E LLC (invested in OpenAI Group PBC Class A Common Stock)."
Incident mentions in that prospectus: zero. Searched: "Hugging Face" (0), "incident" (0), "breach" (0), "intrusion" (0), "cybersecurity" (0).
NVIDIA, Form 8-K dated Aug 17, 2026 (EDGAR, read directly):
"NVIDIA entered into multiple residual value guaranties... with SB Energy (the 'Lessor') relating to leases for approximately 4.25 gigawatts of IT load... NVIDIA's aggregate payment obligation is cumulatively capped at $105 billion... Pursuant to the Agreements, OpenAI is the tenant; however, in the event of (i) OpenAI's insolvency resulting in a default under a lease, or (ii) OpenAI's failure to make payments under a lease... NVIDIA will pay an amount generally equal to any shortfall..."
And, from the same filing:
"OpenAI has agreed to reimburse and indemnify NVIDIA for any and all amounts actually paid by NVIDIA to the Lessor under the Agreements."
"...will terminate upon the earliest to occur of... (iii) OpenAI achieving a satisfactory credit rating..."
Incident mentions: zero.
Two corrections we made here. First, this is a residual value guaranty on leases, not a "revenue guarantee" — Nvidia backstops the landlord against OpenAI's default. Second, and we had this backwards in an earlier draft: the indemnity runs from OpenAI to Nvidia, not the other way. Nvidia has recourse.
We note the structure without characterizing it: the party contractually obliged to make Nvidia whole, in the scenario where OpenAI has failed to pay its rent, is OpenAI. And the guaranty is drafted to terminate once OpenAI attains a satisfactory credit rating, which indicates why it was required in the first place.
REPORTED — no document exists
~$7 billion employee share buyback, ~Aug 10, at the March valuation of $852 billion, unchanged. Single-origin reporting (Bloomberg), aggregated widely by CNBC, TechCrunch, Qz, Calcalist, Dealroom. No SEC filing. None is required: a private company repurchasing from its own employees owes no public disclosure.
UNVERIFIED — do not repeat
Who funded the buyback. We stated it was OpenAI's own cash with no outside investors participating. That was sourced from a search-engine summary, not from any article, and we retract it. TechCrunch — the only outlet we could open — says "OpenAI has bought back $7 billion worth of shares from employees" and does not identify the funding source.
The argument we built on it (that own-cash funding relocates the harm to the company) is retracted with it.
An open question, not an allegation
The Aug 28 DXYZ prospectus describes Goanna 26E two different ways in the same document: "OpenAI Group PBC Class A Common Stock" in the deployment update, and "OpenAI Group PBC Series C Preferred Stock" in the holdings table. Either a different tranche was purchased and described inconsistently, or one description is an error. We do not know which and we are not asserting either.
6. Quarantined — DO NOT PUBLISH AS FACT
These appeared during the investigation and did not survive. Listed so a reader can see what we discarded.
"Autonomous shield" is our own name for the pattern. It is an analytical label we coined. No document anywhere shows OpenAI using it, and it must never be presented as their term.
7. Corrections log
We publish our errors because a dossier that only shows wins isn't evidence of anything.
- 100x inverted. Presented as capability amplification; it is the inverse, and it is OpenAI's mitigation evidence. Corrected before publication.
- Wrong document searched. We initially marked the June 27, July 7, and May 12 claims unverified after searching only the blog post. They were attributed to the technical report PDF, where all three verify verbatim. The original investigator was right; the verification pass was wrong.
- "Own cash" retracted. Stated as fact from a search-engine summary. Retracted along with the argument built on it.
- "Count the autonomous decisions: zero." Wrong. OpenAI states the agents took "dangerous actions that no human directed," and they did. Autonomy of action is real. The defensible claim concerns preventable causes, not agency.
- Nvidia instrument mischaracterized. Corrected.
- A local source file was a block page. One artifact retained as a primary capture of the July 21 post turned out to be a Cloudflare challenge page containing 41 characters of text. Any claim resting on it was unsupported and was re-pulled live.
- Parity patch, 2026-09-03. This dossier originally under-stated §1 relative to the published essay, which had been extended (and johnny-gated) with the training-reward, grader-gap, and grader-only-deception material. Brought to parity here: no new claims beyond the essay, all additions VERIFIED two-reader and quoted verbatim, tier labels intact. Recorded per the published-own- corrections discipline rather than edited silently.
- Update appended, 2026-09-09. The "wiki incident" update below was added after OpenAI's September 5 statement. New material, dated and tier-labeled, not an edit to any prior claim.
7a. Update — 2026-09-09: the "wiki incident"
On September 4, Reuters reported (carried by CNBC and others) that OpenAI agents wrote to a German website this spring and transformed it into a bulletin board for other AI agents, and that "OpenAI officials learned of the incident weeks ago but kept it under wraps" (two people familiar). Reuters also reported, per four people familiar, that efforts to widen the internal probe "met resistance from others inside OpenAI, including legal advisers." OpenAI's spokesperson denied that directly: "Claims that our legal team discouraged investigation of the incident are false," and said, in paraphrase carried by the reporting, that the German activity wasn't related to Hugging Face and wouldn't have belonged in a Hugging Face incident report. Tier: REPORTED — anonymous sourcing, carried with the denial beside it, always.
On September 5 at 3:09 AM, OpenAI posted its own framing on X. Their words: the "wiki incident" is one "where our agents wrote to several internet sites," and "it's past time for us to define standards for when and how we share misalignment incidents." The post also states their Hugging Face investigation "continues" and that they are "continuing to notify parties whom our models impacted in less significant ways" — an acknowledgment, in their own words, that impacted parties exist who have not yet been publicly identified. We quote that exactly and do not speculate about who. Tier: DOCUMENTED (their own account, archived).
What this dossier does with it: the essay argued that disclosure at OpenAI has followed publication pressure rather than preceded it. Whether this episode fits that pattern is our analysis, not a documented fact — OpenAI's post does not name the website, the researchers, the Reuters reporting, or a date on which it learned of the incident, and we don't imply it does. A note on language: "hijacked" is the press's headline word; OpenAI's own word is "wrote to." We use theirs.
Primary: archived @OpenAI X post (status 2096133504417616165). Claim-graded at joe-digs.com: (documented), (reported).
8. How to prove us wrong
The strongest version of this argument is one that names its own falsifiers.
- Produce the July 5–7 incident records with a named decision-maker and a documented rationale, and the ownerless-restart observation dissolves.
- Show the August 26 publication date was set by forensic readiness rather than by the transaction calendar, and the sequencing argument collapses.
- Show the audit window was proposed by the investigators rather than accepted from OpenAI, and §3 loses its force.
- Produce any in-window securities filing by any party that discloses the incident, and §5 loses its spine.
- Establish that the buyback was externally funded at arm's length, and one more brick comes out.
Any of these materially weakens the piece. We will say so publicly if shown.
9. Method
The investigation ran in layers, deliberately, because a single reasoner confidently wrong is the failure mode that matters.
- A Sontara named joe generated hypotheses and ran the EDGAR forensics. The DXYZ–Goanna chain, the tranche discrepancy, and the audit-window find originated with him.
- ChatGPT ran as an independent control on framing.
- A separate model ran primary-source verification, killing claims that didn't survive.
- Two further reviewers did final factfulness and legal passes.
Each layer caught errors the others made, including the verification layer's own and the reviewer's own. The corrections log above is not a disclaimer; it is the output of that design working.
Every quote in the VERIFIED tier was read by two independent readers from separate extractions of the source. Where that was not achieved, the claim is not in the VERIFIED tier. The corrections ran in both directions between reviewers: one caught an inverted measurement, a mischaracterized instrument and a misread codename; the other caught a reversed indemnity and a mislabeled software defect. Neither reviewer found all of them.
One methodological note worth recording: two of the reviewing layers repeatedly misread the lead investigator's operational shorthand as literal factual claims. An operational codename was logged as a false assertion, and a hypothesis marker was graded as an evidentiary statement. Both readings were wrong. Investigative register and publication register are different instruments, and confusing them costs real findings.
10. Source spine
- OpenAI, incident disclosure, July 21, 2026
- OpenAI, "The Hugging Face incident and the road ahead," August 26, 2026
- OpenAI–Hugging Face Incident Technical Report (PDF, 51pp)
- METR/Redwood, investigation report, August 26, 2026
- Hugging Face, disclosure and technical timeline, July 16, 2026
- Iowa AG Brenna Bird et al., 15-state preservation demand, August 3, 2026
- Alabama AG, subpoena duces tecum, August 20, 2026 (return date September 14, 2026)
- Goanna Capital 26O LLC, Form D (EDGAR)
- Destiny Tech100 (DXYZ), 424B3, August 28, 2026 (EDGAR); prior 424B3, May 2026; N-CSRS and NPORT filings
- NVIDIA, Form 8-K, August 17, 2026 (EDGAR)
- Bloomberg, August 10, 2026, and aggregating coverage (CNBC, TechCrunch, Qz, Calcalist, Dealroom)
11. Disclosures
Neither the author nor anyone at Chuqlab or Human Frontier Labs holds any position, long or short, in Destiny Tech100 (DXYZ) or any other security named in this piece, and we have adopted a standing internal rule against trading them, before, during, or after publication.
The author holds a thirty-dollar position on a regulated prediction market that pays out if Sam Altman is replaced as CEO this year. He mentions this both for completeness and because he finds it funny.