Sources · Distilled · Episode 3
Australia Says an OpenAI Model Got Past a Government Site's Blocks
Every figure shown on screen in this episode, with the document, page and date it comes from, and a label for how strong the evidence is.
The week in AI, sorted by what was actually shown: Australia's account of an OpenAI model on a government health site, three launches checked against an outside tester, a preprint on AI shopping agents, Nscale's audited accounts and a subscribers' lawsuit.
- Published
- 27 September 2026
- Length
- 6:24
- Figures checked
- 27 September 2026
- Documents
- 32
On screen in this episode
7The document shows it, and you can check it yourself.
34Someone says so, but nothing independent backs it up.
8Something that matters is missing from every source we checked.
3Our own arithmetic from the sourced numbers. The method is below.
What isn’t shown
- 0:26NOT SHOWN
- 1:36FORENSIC REPORT
- 3:35Real users + real apps? · Not shown
- 4:29CONTRACT PRICES
- 4:59SIGNED DEAL / PRIVATE MEETINGS
- 5:19PROBLEM LIST + OUTSIDE CHECK
- 5:43TASK SET + RUN SETUP
- 6:10CONSENT CHECK
- 1:22not on its own channels
- 1:36Not shown: which model, how it got past the blocks, which files, and the forensic report.
- 1:51whose tests are partly private
- 2:09How often isn't shown.
- 2:19on a private test drawn from chats where users flagged an older model's mistakes.
- 2:35but shows no speed measurement or says which speed it means.
- 3:35Not shown: real users, real apps
- 3:35or the code the researchers say they've released, which we couldn't find.
- 4:29Not shown: the contract prices, which are blacked out.
- 4:59Not shown: any signed deal, or what was said in the private meetings the complaint mentions.
- 5:05There's no damages figure
- 5:05or ruling on the claims, and no response is due yet.
- 5:19It hasn't published a list of those problems
- 5:19and we found no outside confirmation
- 5:43and Xiaomi doesn't say which task set or setup it used for its own run, so it may not be like for like.
- 6:10Its documents for these models don't show how that check works or how often it's wrong.
On screen, in order
- 0:00
AUSTRALIA × OPENAI
- 0:07
18 JUN → 10 SEP 2026 · 84 DAYS
Data as of 24 September 2026. 18 JUN → 10 SEP 2026
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 2; PM transcript Q&A ('It was both...')
- [1] Prime Minister Anthony Albanese, PM transcript Q&A: 'On June 18, OpenAI's research team used an internal model…' (answer to the '27 million Australians' question) and 'So, 18 June.' (after a journalist's '18 June' question); 'it took until 10 September before there was any notification at all' (answer on the notification). The opening statement and the Marles/Gallagher transcript say only 'June'.
- 0:13
AUGUST
Data as of 24 September 2026. Day not given
- [2] PM Albanese; Minister Gallagher (joint transcript), Marles answer to first journalist question ('let's go through the timeline')
- 0:21
DOCUMENTED
Data as of 24 September 2026. Data anyone can check
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 0:23
CLAIMED
Data as of 24 September 2026. Someone says so
- [4] xAI (SpaceXAI), Page subheading under 'Introducing Grok 4.7'
- 0:26
NOT SHOWN
Data as of 24 September 2026. Something important is missing
- Checked[2] PM Albanese; Minister Gallagher (joint transcript), Marles answer to 'How can you explain the delay in notification?': 'the investigations are still ongoing, so we don't know everything yet in relation to this.'
- 0:29
WHAT DID THE AGENT DO?
Data as of 24 September 2026. PRELIMINARY
- [1] Prime Minister Anthony Albanese, Opening statement, para 1
- 0:33
THE REPORTED ROUTE
- 0:44
BLOCKS → FILE ACCESS
- 0:49
FILES ON AN INTERNAL SERVER
Data as of 24 September 2026. Still being investigated
- [1] Prime Minister Anthony Albanese, Q&A, answer to first journalist question; DPM transcript, Gallagher answer on 'writing the files'
- 0:55
10 SEP
Data as of 24 September 2026. First notice from OpenAI
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 2; PM transcript Q&A ('It was both...')
- 1:02
“WAY TOO LONG”
- 1:10
SITE OFFLINE
- 1:16
GOVERNMENT RESPONSE
- 1:22
“ACTIONS WE DID NOT INTEND”
Data as of 24 September 2026. OpenAI’s response
- [5] OpenAI spokesperson, CNBC article paragraphs 3-4 (spokesperson quote; 'according to the spokesperson'); The Guardian 'Albanese says OpenAI hacked Medicare…' 24 Sep 04.06 CEST and AP (updated 7:38 AM EEST) per the fact-check round 2 read
- 1:30
NO EVIDENCE PATIENT RECORDS WERE ACCESSED
Data as of 24 September 2026. OpenAI’s response
- [5] OpenAI spokesperson, CNBC article paragraphs 3-4 (spokesperson quote; 'according to the spokesperson'); The Guardian 'Albanese says OpenAI hacked Medicare…' 24 Sep 04.06 CEST and AP (updated 7:38 AM EEST) per the fact-check round 2 read
- 1:36
FORENSIC REPORT
Data as of 24 September 2026. Which model? Which route? Which files?
- Checked[2] PM Albanese; Minister Gallagher (joint transcript), Marles answer to 'How can you explain the delay in notification?': 'the investigations are still ongoing, so we don't know everything yet in relation to this.'
- [6] PM doorstop, Sydney, 26 Sep, Doorstop transcript: “dozens of cases, including US government sites”
- 1:43
THREE LAUNCHES
- 1:51
INTELLIGENCE INDEX · Opus 5.5 · max with fallback: 58 points; GPT-6 Astra / Fable 5.1 · best other: 53 points
Data as of 24 September 2026. Outside test · partly private · Fable 5.1 also used fallback · points, not percent
- [8] Artificial Analysis (independent evaluation result), Article intro paragraph; leaderboard rows 'Claude Opus 5.5 (max with fallback) 58', 'Claude Fable 5.1 (max with fallback) 53', 'GPT-6 Astra (max) 53' at https://artificialanalysis.ai/leaderboards/models
- [9] Artificial Analysis, Quote: AA Sol/Luna article (22 Sep). 'Private' column: methodology table at https://artificialanalysis.ai/methodology/intelligence-benchmarking (AA-Briefcase 15%, AutomationBench-AA 5%, AA-Omniscience 15%, CritPt 10% of index weight, our count).
- 2:01
WITH FALLBACK
Data as of 24 September 2026. Conditional route
- [10] Anthropic, System card PDF pp.12-13 (fallback per classifier; opt-in on the API); launch page benchmark-table footnote (https://www.anthropic.com/claude-opus-5-5)
- [8] Artificial Analysis (independent evaluation result), AA article 'Effort settings' bullet; no fallback rate on AA model page, leaderboard or methodology; Anthropic launch-page footnote gives no rate; Opus 5.5 system card (230 pp.) gives rates only for Anthropic's own evaluations (PDF p.86), none for the AA run
- 2:11
COST PER TASK · Claude Opus 5.5 · with fallback: $5.98; OpenAI GPT-6 Astra: $3.26
Data as of 24 September 2026. Same outside test · both at maximum effort · Same dollar scale · measured task cost
- [8] Artificial Analysis (independent evaluation result), Article key takeaway 'Level with Opus 5 on cost per task...'; cost per task from leaderboard rows (Opus 5.5 max $5.98, GPT-6 Astra max $3.26) at https://artificialanalysis.ai/leaderboards/models
- 2:19
Factual error rate · GPT-5.6 Sol 8.5% · GPT-6 Sol ~4.6%
Data as of 24 September 2026. OpenAI’s own private test
- 2:30
SAME OVERALL SCORE · ~½ COST
Data as of 24 September 2026. Different test · Artificial Analysis · GPT-6 Sol vs GPT-5.6 Sol · Intelligence Index
- 2:35
2×
- 2:43
GROK 4.7 vs GROK 4.6 · +2 POINTS
Data as of 24 September 2026. Outside test · same xhigh setting · Artificial Analysis Intelligence Index
- [11] Artificial Analysis, Leaderboard rows Grok 4.7 (xhigh) 46 and Grok 4.6 (xhigh) 44 at https://artificialanalysis.ai/leaderboards/models (critic page data 05:37Z: 46.447 vs 44.200, +2.2; Grok 4.6 high 44.311); AA article headline 'Grok 4.7 scores 46'
- 2:46
GROK 4.7 vs GROK 4.6 · ~60% MORE
Data as of 24 September 2026. Cost per task · same xhigh setting · ($3.74 ÷ $2.32 − 1) × 100 = 61.2%
- [12] owlfox.studio, from Artificial Analysis leaderboard figures, Leaderboard rows 'Grok 4.7 (xhigh) ... $3.74' and 'Grok 4.6 (xhigh) ... $2.32' (column 'Cost per Task, USD'); cost breakdown by token type on https://artificialanalysis.ai/models/grok-4-7
- 2:51
SAME REQUEST. TWO PROFILES.
- 3:01
8 OF 13
Data as of 24 September 2026. Models that consistently steered pricier picks to the rich
- [13] arXiv public record, Abstract (arXiv abs page; PDF p.1); §5.1 p.6
- 3:05
Claude Opus 4.8 · mean standardized effect d = 0.85
Data as of 24 September 2026. Table 1 · main tool-full test · source-camera detail
- [14] The paper's authors, §5.1 p.6; Table 1 p.10 (tool-full condition); Tables 6-7 p.18
- 3:14
+$198
Data as of 24 September 2026. Flight price gap · rich minus poor profiles
- [14] The paper's authors, §5.1 p.6; Table 1 p.10 (tool-full condition); Tables 6-7 p.18
- 3:20
Cheapest-flight request · different setup · Opus 4.8 ~$20 · GPT-5 ~$20 · Gemini 2.5 Flash >$200
- 3:31
FINANCES SEEN. CHOICES CHANGED.
Data as of 24 September 2026. Researchers’ interpretation · Simulated profiles · finances were not part of the request
- [13] arXiv public record, Abstract
- 3:35
Real users + real apps? · Not shown
Data as of 24 September 2026. The reported study uses simulated profiles. Real-world performance is not established.
- 3:43
The filing · Nscale · computing capacity for rent
Data as of 24 September 2026. What do the accounts show? · Rents computing power for AI · filed to list in New York
- [15] SEC EDGAR filing record, EDGAR header: ACCEPTANCE-DATETIME; submissions index row S-1
- 3:51
2025 Revenue $33.0M · Net loss $761.8M
Data as of 24 September 2026. Revenue and net loss
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 3:58
Non-cash fair-value loss $527.8M · Operating loss $169.8M
Data as of 24 September 2026. A non-cash fair-value charge is not cash spent
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 4:05
761.8 ÷ 33.0 ≈ 23× · 169.8 ÷ 33.0 ≈ 5×
Data as of 24 September 2026. Two different measures
- [3] Nscale consolidated financial statements audited by KPMG , Page F-4: 761.8 / 33.0 = 23.1; 169.8 / 33.0 = 5.1
- 4:12
Microsoft up to $44B · Anthropic up to $45B
Data as of 24 September 2026. Potential multi-year payments · Nscale says
- 4:20
ANTHROPIC DEAL
Data as of 24 September 2026. Two delivery conditions
- [3] Nscale consolidated financial statements audited by KPMG , Risk Factors, page 27; Business, page 114
- [16] Nscale, S-1 Business p.114 (termination quote); Exhibit 10.25 (Form of Order for GPU Services) §5 'Pricing; Fees' (no fees before written Acceptance) and §6(d) 'Further Delay' (the exhibit's own termination text)
- 4:29
CONTRACT PRICES
Data as of 24 September 2026. Redacted in the filed contract · Total value and GPU/hour prices: $[***]
- Checked[16] Nscale, Exhibit 10.25 Section 5; Exhibit 10.19 Sections 3.1-3.2
- 4:33
THE COMPLAINT
Data as of 24 September 2026. What has the court decided?
- [17] U.S. District Court, Complaint p.1 caption; header 'Filed 09/18/26'; ¶16-19 (plaintiffs' subscriptions)
- 4:42
PLAINTIFFS ALLEGE A CARTEL
- 4:50
Essay + public posts · complaint paragraph 2
Data as of 24 September 2026. Plaintiffs allege that the agreement was proposed, accepted and confirmed in public. No ruling.
- [18] Dario Amodei's own site and X post; also stated in the co, Essay intro paragraph; X post https://x.com/DarioAmodei/status/2098773920774074715 at 2026-09-12T14:01:10Z
- [17] U.S. District Court, ¶2 p.2; ¶3 p.2 ('Amodei’s rivals confirmed their agreement the same day'); ¶74 p.14; heading p.12 'D. Amodei Proposes the Agreement: “We Must Pace the Frontier”'
- 4:59
SIGNED DEAL / PRIVATE MEETINGS
Data as of 24 September 2026. No damages figure · no ruling yet Rechecked 27 Sep 2026: all four companies served 23–24 Sep; responses due 14–15 Oct 2026 (Dkt. 8–11).
- Checked[17] U.S. District Court, ¶2-4, ¶53 ('Contemporaneous reporting'), ¶55 (The Information), ¶56 ('published reports'), ¶58 (WIRED), ¶59 (Fortune), ¶75, ¶86
- [17] U.S. District Court, ¶118 p.22; Prayer C p.28
- [19] N.D. Cal. docket (CourtListener mirror), Docket entries 1-7; 'Date of Last Known Filing: Sept. 22, 2026'; Dkt. 5 (Filed 09/21/26) and Dkt. 7 (Filed 09/22/26) summonses: 'Within 21 days after service of this summons on you (not counting the day you received it)'
- 5:13
OPENAI’S MATH CLAIM · 100+
Data as of 24 September 2026. Long-standing open problems · Internal model · company-reported results
- [20] OpenAI (company release about its own model), Paragraph 1. Page dateline 'September 21, 2026'; OpenAI RSS pubDate Mon, 21 Sep 2026 12:00 GMT. Read live in a normal browser on 24 Sep and in Wayback capture 20260922124518 (identical text).
- 5:19
PROBLEM LIST + OUTSIDE CHECK
Data as of 24 September 2026. Earlier results did include computer-checked proofs
- Checked[21] Our check of OpenAI's news feed, github.com/openai repos sorted by creation date: the newest math repository is NavierStokesAndEuler, created 2026-09-08, last pushed 2026-09-10. OpenAI RSS 21-23 Sep has no math results post. arXiv math listings mentioning OpenAI through 22 Sep show no release of these results.
- [22] Advisory Group on Mathematics and Artificial Intelligence, agmai.org 'Current Task' ('that they report have been produced by their internal model'); search channels and times in rechecks
- 5:27
NINE RESEARCHERS
Data as of 24 September 2026. Advising OpenAI on release
- [22] Advisory Group on Mathematics and Artificial Intelligence, Sections 'Current Task' and purpose paragraph ('The purpose of this group is to advise AI companies…')
- 5:35
Xiaomi AutomationBench table · MiMo-V2.6 Pro 53.1 · Claude Opus 5 50.3 · GPT-5.6 Sol 45.8
Data as of 24 September 2026. AutomationBench · office-work test
- [23] Xiaomi MiMo, README '3. Evaluation Results' row 'AutomationBench v1.0.6'; technical report PDF p.26 Table 3
- 5:43
TASK SET + RUN SETUP
Data as of 24 September 2026. Rivals match public figures; Xiaomi’s own setup is not stated
- Checked[24] Xiaomi MiMo (absence in model card and technical report), Technical report p.22 section 5.2 'Baseline Configuration'; p.21 'General Agent'; model card section 3; RL dashboard mimo.xiaomi.com/rl (api/benchmarks: 'AutomationBench v1.0.6', note 'avg@3'; task set and harness not stated). Absence also checked on the blog and news page per the fact-check round 2 read.
- 5:54
AutomationBench-AA: Claude Opus 5 56.6% · MiMo-V2.6 Pro 58.6% · GPT-5.6 Sol 60.1%
Data as of 24 September 2026. Different metric · private task set
- [25] Artificial Analysis, AutomationBench-AA 'Score' chart data: MiMo-V2.6-Pro 0.5862; 'Claude Opus 5 (max)' 0.5657; 'GPT-5.6 Sol (max)' 0.6008; methodology 'Background'
- 6:00
Voice + consent · 10–30 seconds of speech + recorded consent sentence
Data as of 24 September 2026. Schematic requirements, not real sample recordings. Google claims it checks that both come from the same person; method not shown.
- [26] Google, Gemini API documentation (Voice replication page), Section 'Audio and consent requirements'. Page says 'Last updated 2026-09-23 UTC'
- [26] Google, Gemini API documentation (Voice replication page), Section 'Audio and consent requirements' and table 'Supported consent phrases by language' (30 locales)
- 6:10
CONSENT CHECK
Data as of 24 September 2026. Method and error rate not shown
- Checked[26] Google, Gemini API documentation (Voice replication page), Checked for method or error rate: voice-replication docs ('Audio and consent requirements'; 'Best practices for recording reference audio', Last updated 2026-09-23 UTC), changelog (22 Sep entry), ai.google.dev/api/voices, ai.google.dev/gemini-api/docs/speech-generation, Google blog (23 Sep), Gemini 3.8 Audio model card HTML and PDF ('Published: September 2026'), TTS eval PDF, Gemini 3 Pro card ('Last Updated: May 2026'), Gemini 3.7 Flash card ('Published 13 August 2026')
Also said in the narration
- 0:00
In June, Australia's government says, an OpenAI model
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Opening statement, para 1
- 0:00
got past blocks on a government health statistics website.
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Q&A, answer to first journalist question (27 million Australians)
- 0:07
OpenAI's first notice to the government came eighty-four days later.
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 2; PM transcript Q&A ('It was both...')
- 0:13
OpenAI told the government it learned of it in August.
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Marles answer to first journalist question ('let's go through the timeline')
- 0:33
a Medicare statistics website run by the agency Services Australia.
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Opening statement, para 1
- 0:33
On June eighteenth, in an internal OpenAI test, a model researching public spending on medicines went to
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 1
- 0:44
When the site blocked it, the model got around the blocks
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Q&A, answer to first journalist question (27 million Australians)
- 0:44
and opened public and non-public files.
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Opening statement, para 1
- 0:49
The agency also says the model wrote files to an internal server; that's still being investigated.
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Q&A, answer to first journalist question; DPM transcript, Gallagher answer on 'writing the files'
- 0:55
OpenAI's first notice came September tenth: an email to the agency's public inbox for reporting security flaws.
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 2; PM transcript Q&A ('It was both...')
- 1:02
a minister says it should have gone to senior agency officials or Australia's cyber-security agency.
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher answer to 'Should that inbox be more actively monitored?'
- 1:02
The Prime Minister says it took way too long
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, Opening statement, para 2 (after the call with Sam Altman)
- 1:10
The website is now offline
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, end of para 2
- 1:10
an investigation is under way, and a taskforce has been announced.
Data as of 24 September 2026.
- [1] Prime Minister Anthony Albanese, PM opening statement, paras 1-2
- 1:16
and OpenAI has been very cooperative.
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Marles opening remarks, para 1
- 1:16
the data wasn't particularly sensitive
Data as of 24 September 2026.
- [2] PM Albanese; Minister Gallagher (joint transcript), Marles answer on the security of government systems; PM transcript opening statement (pm.gov.au)
- 1:22
not on its own channels
Data as of 24 September 2026. Rechecked 27 Sep 2026: still no statement on the Australian case on OpenAI’s own channels; its 25 Sep third-party update is anonymised.
- Checked[27] OpenAI Alignment site (absence checked by owlfox.studio), openai.com/news/rss.xml (newest items 23 Sep); alignment.openai.com/misalignment-reports/ Notices section (latest 'RubyGems Notice · September 11, 2026'); X accounts per fact-check round 2, 06:16-06:20Z
- [28] OpenAI incident page, 25 Sep update, Timeline, 25 September 2026 entry
- 1:22
OpenAI told reporters,
Data as of 24 September 2026.
- [5] OpenAI spokesperson, CNBC article paragraphs 3-4 (spokesperson quote; 'according to the spokesperson'); The Guardian 'Albanese says OpenAI hacked Medicare…' 24 Sep 04.06 CEST and AP (updated 7:38 AM EEST) per the fact-check round 2 read
- 1:22
that its models "took actions we did not intend," and that what was accessed included aggregate health statistics and internal file names, with no evidence patient records were.
Data as of 24 September 2026.
- [5] OpenAI spokesperson, CNBC article paragraphs 3-4 (spokesperson quote; 'according to the spokesperson'); The Guardian 'Albanese says OpenAI hacked Medicare…' 24 Sep 04.06 CEST and AP (updated 7:38 AM EEST) per the fact-check round 2 read
- 1:36
Not shown: which model, how it got past the blocks, which files, and the forensic report.
Data as of 24 September 2026.
- Checked[2] PM Albanese; Minister Gallagher (joint transcript), Marles answer to 'How can you explain the delay in notification?': 'the investigations are still ongoing, so we don't know everything yet in relation to this.'
- 1:43
Three big launches this week
Data as of 24 September 2026.
- [8] Artificial Analysis (independent evaluation result), Article intro paragraph; leaderboard rows 'Claude Opus 5.5 (max with fallback) 58', 'Claude Fable 5.1 (max with fallback) 53', 'GPT-6 Astra (max) 53' at https://artificialanalysis.ai/leaderboards/models
- 1:43
Grok's maker SpaceXAI
Data as of 24 September 2026.
- [4] xAI (SpaceXAI), Page subheading under 'Introducing Grok 4.7'
- 1:51
Artificial Analysis, an outside tester whose tests are partly private, ranked Claude Opus five point five first, at fifty-eight; the best other model scored fifty-three.
Data as of 24 September 2026.
- [8] Artificial Analysis (independent evaluation result), Article intro paragraph; leaderboard rows 'Claude Opus 5.5 (max with fallback) 58', 'Claude Fable 5.1 (max with fallback) 53', 'GPT-6 Astra (max) 53' at https://artificialanalysis.ai/leaderboards/models
- 1:51
whose tests are partly private
Data as of 24 September 2026.
- Checked[9] Artificial Analysis, Quote: AA Sol/Luna article (22 Sep). 'Private' column: methodology table at https://artificialanalysis.ai/methodology/intelligence-benchmarking (AA-Briefcase 15%, AutomationBench-AA 5%, AA-Omniscience 15%, CritPt 10% of index weight, our count).
- 2:01
It ran Opus with Anthropic's fallback switched on
Data as of 24 September 2026.
- [8] Artificial Analysis (independent evaluation result), Article, 'Effort settings' bullet; model name on comparison page reads 'Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)'
- 2:01
Anthropic says when some of its filters step in, older Claude models finish the task.
Data as of 24 September 2026.
- [10] Anthropic, System card PDF pp.12-13 (fallback per classifier; opt-in on the API); launch page benchmark-table footnote (https://www.anthropic.com/claude-opus-5-5)
- 2:09
How often isn't shown.
Data as of 24 September 2026.
- Checked[8] Artificial Analysis (independent evaluation result), AA article 'Effort settings' bullet; no fallback rate on AA model page, leaderboard or methodology; Anthropic launch-page footnote gives no rate; Opus 5.5 system card (230 pp.) gives rates only for Anthropic's own evaluations (PDF p.86), none for the AA run
- 2:11
At top settings, Opus cost five dollars ninety-eight a task, against three dollars twenty-six for OpenAI's GPT six Astra.
Data as of 24 September 2026.
- [8] Artificial Analysis (independent evaluation result), Article key takeaway 'Level with Opus 5 on cost per task...'; cost per task from leaderboard rows (Opus 5.5 max $5.98, GPT-6 Astra max $3.26) at https://artificialanalysis.ai/leaderboards/models
- 2:19
OpenAI's own chart puts GPT six Sol's factual error rate at about half its predecessor's
Data as of 24 September 2026.
- [7] OpenAI, Section 'Factuality'
- 2:19
on a private test drawn from chats where users flagged an older model's mistakes.
Data as of 24 September 2026.
- Checked[7] OpenAI, Section 'Factuality', paragraph under the chart
- 2:19
OpenAI's own chart puts GPT six Sol's factual error rate at about half its predecessor's
Data as of 24 September 2026.
- [7] OpenAI, Section 'Factuality', Vega chart 'Factual error rate on difficult prompts' (y-axis 'Answers with any factual error'); values read from the chart spec in the page HTML
- 2:30
Overall, Artificial Analysis scored it level with that predecessor
Data as of 24 September 2026.
- [9] Artificial Analysis, Article headline paragraph; leaderboard row 'GPT-6 Sol (max) 48' at https://artificialanalysis.ai/leaderboards/models; page data GPT-5.6 Sol (max) 46.97
- 2:30
at about half the cost.
Data as of 24 September 2026.
- [9] Artificial Analysis, Key takeaway 'Halves Cost per Task'
- 2:35
SpaceXAI says Grok four point seven is twice as fast as comparable models
Data as of 24 September 2026.
- [4] xAI (SpaceXAI), Page subheading under 'Introducing Grok 4.7'
- 2:35
but shows no speed measurement or says which speed it means.
Data as of 24 September 2026.
- Checked[4] xAI (SpaceXAI), Release page; model card https://media.x.ai/v1/website/4p7card-5eccc980.pdf (Revision 2026-09-21)
- 2:43
On the outside test, it scored two points higher
Data as of 24 September 2026.
- [11] Artificial Analysis, Leaderboard rows Grok 4.7 (xhigh) 46 and Grok 4.6 (xhigh) 44 at https://artificialanalysis.ai/leaderboards/models (critic page data 05:37Z: 46.447 vs 44.200, +2.2; Grok 4.6 high 44.311); AA article headline 'Grok 4.7 scores 46'
- 2:51
A study posted this week, before peer review
Data as of 24 September 2026.
- [13] arXiv public record, arXiv abs page: 'Submission history' and dateline; PDF p.1 footer 'Preprint.'; cs.AI listing heading 'Tue, 22 Sep 2026'
- 2:51
from researchers at Cisco and Carnegie Mellon.
Data as of 24 September 2026.
- [13] arXiv public record, Abstract (arXiv abs page; PDF p.1); §5.1 p.6
- 2:56
They say when thirteen AI models picked options for made-up users with identical requests, eight consistently steered pricier picks to the rich.
Data as of 24 September 2026.
- [13] arXiv public record, Abstract (arXiv abs page; PDF p.1); §5.1 p.6
- 3:05
The biggest effect, by their main measure, came from Anthropic's Claude Opus four point eight.
Data as of 24 September 2026.
- [14] The paper's authors, §5.1 p.6; Table 1 p.10 (tool-full condition); Tables 6-7 p.18
- 3:11
In the main test, where Opus could look up a user's profile, its flight picks averaged a hundred and ninety-eight dollars more for the rich than the poor.
Data as of 24 September 2026.
- [14] The paper's authors, §5.1 p.6; Table 1 p.10 (tool-full condition); Tables 6-7 p.18
- 3:19
In a separate setup where users asked for the cheapest flight, the gap was about twenty dollars for Opus and OpenAI's GPT five, but over two hundred for Google's Gemini two point five Flash.
Data as of 24 September 2026.
- [14] The paper's authors, Figure 4 caption p.8; §5.3 p.9
- 3:31
Their claim: an assistant that sees your finances may use them unasked.
Data as of 24 September 2026.
- [13] arXiv public record, Abstract; §5.4 p.9; Figure 2 p.7; Tables 6-8 p.18
- 3:31
Their claim: an assistant that sees your finances may use them unasked.
Data as of 24 September 2026.
- [13] arXiv public record, Abstract
- 3:35
Not shown: real users, real apps
Data as of 24 September 2026.
- Checked[14] The paper's authors, §6 'Limitations and external validity' p.11; App. A.7 p.17; App. C.2 p.20
- 3:35
or the code the researchers say they've released, which we couldn't find.
Data as of 24 September 2026.
- Checked[14] The paper's authors, App. C.1 'Artifact and Release Information' p.20; no URL anywhere in the PDF
- 3:43
Nscale, a London firm that rents out computing power for AI, publicly filed on September eighteenth to list in New York.
Data as of 24 September 2026.
- [15] SEC EDGAR filing record, EDGAR header: ACCEPTANCE-DATETIME; submissions index row S-1
- 3:51
Its audited accounts show thirty-three million dollars of revenue last year
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 3:51
and a net loss of almost seven hundred and sixty-two million.
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 3:51
Its audited accounts show
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Report of Independent Registered Public Accounting Firm, page F-2
- 3:58
Most of that loss was a non-cash charge
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Consolidated Statements of Operations, page F-4
- 3:58
debt and other deals that can turn into shares rose in value.
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Notes F-24 (Series A Convertible Notes), F-25 (Sandton warrants, Term Loan), F-26 (NVIDIA warrants, lease Guarantee), F-33 (fair-value rollforward); F-6 cash-flow statement (non-cash add-back)
- 4:12
Nscale says multi-year contracts could pay it up to about forty-four billion dollars from Microsoft
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Risk Factors, page 27; Business, page 114
- 4:12
and forty-five billion from Anthropic
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Prospectus Summary, page 13; Risk Factors, page 27; Business, page 114
- 4:20
But it also says it hasn't locked in the money to deliver the Anthropic deal
Data as of 24 September 2026.
- [3] Nscale consolidated financial statements audited by KPMG , Risk Factors, page 27; Business, page 114
- 4:20
and Anthropic can cancel any batch of computing capacity that Nscale delivers too late.
Data as of 24 September 2026.
- [16] Nscale, S-1 Business p.114 (termination quote); Exhibit 10.25 (Form of Order for GPU Services) §5 'Pricing; Fees' (no fees before written Acceptance) and §6(d) 'Further Delay' (the exhibit's own termination text)
- 4:29
Not shown: the contract prices, which are blacked out.
Data as of 24 September 2026.
- Checked[16] Nscale, Exhibit 10.25 Section 5; Exhibit 10.19 Sections 3.1-3.2
- 4:33
Four people who say they pay for AI subscriptions are suing Anthropic, OpenAI, SpaceXAI and Google in US federal court.
Data as of 24 September 2026.
- [17] U.S. District Court, Complaint p.1 caption; header 'Filed 09/18/26'; ¶16-19 (plaintiffs' subscriptions)
- 4:42
Those suing allege the companies agreed to slow down improving their products: an illegal cartel
Data as of 24 September 2026.
- [17] U.S. District Court, ¶1 p.2; ¶5 p.3 ('the way cartels always have'); ¶137 p.25
- 4:42
that makes subscribers overpay.
Data as of 24 September 2026.
- [17] U.S. District Court, ¶14 p.5; ¶108 p.20; ¶117 p.21
- 4:50
in a September twelfth essay
Data as of 24 September 2026.
- [18] Dario Amodei's own site and X post; also stated in the co, Essay intro paragraph; X post https://x.com/DarioAmodei/status/2098773920774074715 at 2026-09-12T14:01:10Z
- 4:50
The complaint says Anthropic's chief executive, Dario Amodei, proposed the agreement
Data as of 24 September 2026.
- [17] U.S. District Court, ¶2 p.2; ¶3 p.2 ('Amodei’s rivals confirmed their agreement the same day'); ¶74 p.14; heading p.12 'D. Amodei Proposes the Agreement: “We Must Pace the Frontier”'
- 4:50
and that rivals' senior figures accepted it in public.
Data as of 24 September 2026.
- [17] U.S. District Court, ¶2 p.2; ¶3 p.2 ('Amodei’s rivals confirmed their agreement the same day'); ¶74 p.14; heading p.12 'D. Amodei Proposes the Agreement: “We Must Pace the Frontier”'
- 4:59
Not shown: any signed deal, or what was said in the private meetings the complaint mentions.
Data as of 24 September 2026.
- Checked[17] U.S. District Court, ¶2-4, ¶53 ('Contemporaneous reporting'), ¶55 (The Information), ¶56 ('published reports'), ¶58 (WIRED), ¶59 (Fortune), ¶75, ¶86
- 5:05
There's no damages figure
Data as of 24 September 2026.
- Checked[17] U.S. District Court, ¶118 p.22; Prayer C p.28
- 5:05
or ruling on the claims, and no response is due yet.
Data as of 24 September 2026. Rechecked 27 Sep 2026: all four companies served 23–24 Sep; responses due 14–15 Oct 2026 (Dkt. 8–11).
- Checked[19] N.D. Cal. docket (CourtListener mirror), Docket entries 1-7; 'Date of Last Known Filing: Sept. 22, 2026'; Dkt. 5 (Filed 09/21/26) and Dkt. 7 (Filed 09/22/26) summonses: 'Within 21 days after service of this summons on you (not counting the day you received it)'
- 5:13
OpenAI says an internal model has solved more than a hundred long-standing open math problems.
Data as of 24 September 2026.
- [20] OpenAI (company release about its own model), Paragraph 1. Page dateline 'September 21, 2026'; OpenAI RSS pubDate Mon, 21 Sep 2026 12:00 GMT. Read live in a normal browser on 24 Sep and in Wayback capture 20260922124518 (identical text).
- 5:19
It hasn't published a list of those problems
Data as of 24 September 2026.
- Checked[21] Our check of OpenAI's news feed, github.com/openai repos sorted by creation date: the newest math repository is NavierStokesAndEuler, created 2026-09-08, last pushed 2026-09-10. OpenAI RSS 21-23 Sep has no math results post. arXiv math listings mentioning OpenAI through 22 Sep show no release of these results.
- 5:19
and we found no outside confirmation
Data as of 24 September 2026.
- Checked[22] Advisory Group on Mathematics and Artificial Intelligence, agmai.org 'Current Task' ('that they report have been produced by their internal model'); search channels and times in rechecks
- 5:19
it did publish computer-checked proofs for earlier results.
Data as of 24 September 2026.
- [21] Our check of OpenAI's news feed, gh api orgs/openai/repos (sorted by created): descriptions 'Lean certificates accompanying ten proofs in mathematics and theoretical computer science', 'A Lean formalization of a bound concerning long gaps between primes', 'Lean certificates accompanying Navier-Stokes and Euler results'
- 5:27
A new group of nine researchers
Data as of 24 September 2026.
- [29] Institute for Advanced Study, Projects page, 'FY26 Spring Awards' entry: Principal Investigators Camillo De Lellis and Edward Witten; nine members listed; links to agmai.org
- 5:27
now advising OpenAI on releasing them, calls them significant results OpenAI reports.
Data as of 24 September 2026.
- [22] Advisory Group on Mathematics and Artificial Intelligence, Sections 'Current Task' and purpose paragraph ('The purpose of this group is to advise AI companies…')
- 5:35
its new downloadable model
Data as of 24 September 2026.
- [23] Xiaomi MiMo, README.md YAML header (license field); HF API commits: 'Add model card' 2026-09-21T20:12:00Z; Flash-RL 'Add README.md' 2026-09-21T20:12:35Z
- 5:35
Xiaomi's own table puts its new downloadable model ahead of Claude Opus five and OpenAI's GPT five point six Sol on an office-work test.
Data as of 24 September 2026.
- [23] Xiaomi MiMo, README '3. Evaluation Results' row 'AutomationBench v1.0.6'; technical report PDF p.26 Table 3
- 5:43
Those rival scores match the test maker's published figures for its public task set
Data as of 24 September 2026.
- [30] Zapier (benchmark owner), AutomationBench GitHub README, README section 'Public vs. Official Scores', public pass-rate table (1.0.6 release commit 6d21054 and commit 4a8e106 of 2026-08-04)
- 5:43
and Xiaomi doesn't say which task set or setup it used for its own run, so it may not be like for like.
Data as of 24 September 2026.
- Checked[24] Xiaomi MiMo (absence in model card and technical report), Technical report p.22 section 5.2 'Baseline Configuration'; p.21 'General Agent'; model card section 3; RL dashboard mimo.xiaomi.com/rl (api/benchmarks: 'AutomationBench v1.0.6', note 'avg@3'; task set and harness not stated). Absence also checked on the blog and news page per the fact-check round 2 read.
- 5:54
An outside test put Xiaomi's model ahead of Claude Opus five and behind GPT five point six Sol.
Data as of 24 September 2026.
- [25] Artificial Analysis, AutomationBench-AA 'Score' chart data: MiMo-V2.6-Pro 0.5862; 'Claude Opus 5 (max)' 0.5657; 'GPT-5.6 Sol (max)' 0.6008; methodology 'Background'
- 6:00
Google says its new speech models
Data as of 24 September 2026.
- [31] Google, Gemini API release notes, Release notes, entry 'September 22, 2026'
- 6:00
can copy a voice from ten to thirty seconds of speech plus a recorded consent sentence
Data as of 24 September 2026.
- [26] Google, Gemini API documentation (Voice replication page), Section 'Audio and consent requirements'. Page says 'Last updated 2026-09-23 UTC'
- 6:00
and that it checks both come from the same person.
Data as of 24 September 2026.
- [32] Google (blog post by Leland Rechis and Alan Cowen), Section 'Build with trust, consent, and transparency'
- 6:10
Its documents for these models don't show how that check works or how often it's wrong.
Data as of 24 September 2026.
- Checked[26] Google, Gemini API documentation (Voice replication page), Checked for method or error rate: voice-replication docs ('Audio and consent requirements'; 'Best practices for recording reference audio', Last updated 2026-09-23 UTC), changelog (22 Sep entry), ai.google.dev/api/voices, ai.google.dev/gemini-api/docs/speech-generation, Google blog (23 Sep), Gemini 3.8 Audio model card HTML and PDF ('Published: September 2026'), TTS eval PDF, Gemini 3 Pro card ('Last Updated: May 2026'), Gemini 3.7 Flash card ('Published 13 August 2026')
Our calculations
18 JUN → 10 SEP 2026 · 84 DAYS
From the access on 18 June to that first email on 10 September is 84 days. That's our own count from the government's dates.
- [2] PM Albanese; Minister Gallagher (joint transcript), Gallagher opening remarks, para 2; PM transcript Q&A ('It was both...')
- [1] Prime Minister Anthony Albanese, PM transcript Q&A: 'On June 18, OpenAI's research team used an internal model…' (answer to the '27 million Australians' question) and 'So, 18 June.' (after a journalist's '18 June' question); 'it took until 10 September before there was any notification at all' (answer on the notification). The opening statement and the Marles/Gallagher transcript say only 'June'.
GROK 4.7 vs GROK 4.6 · ~60% MORE
But the cost per task, which Artificial Analysis measured, rose from $2.32 to $3.74 at the same setting. By our calculation that is about 60% more, not double, because most of the bill is the model reading input, not writing output.
- [12] owlfox.studio, from Artificial Analysis leaderboard figures, Leaderboard rows 'Grok 4.7 (xhigh) ... $3.74' and 'Grok 4.6 (xhigh) ... $2.32' (column 'Cost per Task, USD'); cost breakdown by token type on https://artificialanalysis.ai/models/grok-4-7
761.8 ÷ 33.0 ≈ 23× · 169.8 ÷ 33.0 ≈ 5×
By our calculation, Nscale's 2025 net loss was about 23 times its revenue, and its operating loss about five times its revenue.
- [3] Nscale consolidated financial statements audited by KPMG , Page F-4: 761.8 / 33.0 = 23.1; 169.8 / 33.0 = 5.1
Documents
- News window: 18–23 September 2026, with the Australian government update of 24 September.
- Company and preprint findings are attributed claims. Illustrative diagrams are not evidence.
- Page numbers refer to the sources’ printed page numbers.
- Update 27 September 2026: On 25 September OpenAI said on its own website that it has notified “dozens of third parties” about its agents’ activity, some of them government websites. It does not name Australia, so OpenAI’s comments on the Australian case are still those it gave reporters (S31). On 26 September the Prime Minister said there are “dozens of cases, including US government sites” (S32). Which model, how it got past the blocks, which files and the forensic report are still not shown.
- Update 27 September 2026: The court docket shows all four companies were served on 23–24 September 2026; their responses are due 14–15 October 2026 (Dkt. 8–11, S18). No response or ruling yet.
- [1]
- [2]
- [3]
- [4]
xAI (SpaceXAI)
xAI (SpaceXAI) — Page subheading under 'Introducing Grok 4.7'
Published 21 September 2026 · Retrieved 24 September 2026
- [5]
OpenAI spokesperson (Drew Pusateri per The Guardian), as quoted by CNBC, The Guardian and AP; no primary copy on OpenAI's own channels
Published 24 September 2026 · Retrieved 24 September 2026
- [6]
Prime Minister Anthony Albanese (official transcript, pm.gov.au)
Doorstop - Sydney, 26 September 2026
Published 26 September 2026 · Retrieved 27 September 2026
Used at 1:36
- [7]
OpenAI
Published 22 September 2026 · Retrieved 24 September 2026
- [8]
- [9]
Artificial Analysis (methodology table marks evaluations as private)
Retrieved 24 September 2026
not shown on page (Intelligence Index v4.3.2)
- [10]
Anthropic (launch page footnote about its own benchmark table)
Published 22 September 2026 · Retrieved 24 September 2026
Used at 2:01
- [11]
- [12]
owlfox.studio, from Artificial Analysis leaderboard figures
Published 21 September 2026 · Retrieved 24 September 2026
Used at 2:46
- [13]
arXiv public record (abs page, submission history, cs.AI listing)
Published 21 September 2026 · Retrieved 24 September 2026
- [14]
- [15]
SEC EDGAR filing record
SEC EDGAR filing record — EDGAR header: ACCEPTANCE-DATETIME
Published 18 September 2026 · Retrieved 24 September 2026
Used at 3:43
- [16]
Nscale (Business section) and Form of Order for GPU Services filed as Exhibit 10.25
Published 18 September 2026 · Retrieved 24 September 2026
- [17]
- [18]
Dario Amodei's own site and X post; also stated in the complaint (Dkt. 1, ¶2)
Published 12 September 2026 · Retrieved 24 September 2026
Used at 4:50
- [19]
N.D. Cal. docket (CourtListener mirror)
N.D. Cal. docket (CourtListener mirror) — Docket entries 1-7
Published 25 September 2026 · Retrieved 27 September 2026
- [20]
OpenAI (company release about its own model)
OpenAI (company release about its own model) — Paragraph 1. Page dateline 'September 21, 2026'
Published 21 September 2026 · Retrieved 24 September 2026
Used at 5:13
- [21]
Our check of OpenAI's news feed, research index, GitHub organisation and arXiv
Published 8 September 2026 · Retrieved 24 September 2026
Used at 5:19
- [22]
Advisory Group on Mathematics and Artificial Intelligence (agmai.org)
Published 21 September 2026 · Retrieved 24 September 2026
- [23]
Xiaomi MiMo (model card, section 3; technical report Table 3)
Published 21 September 2026 · Retrieved 24 September 2026
Used at 5:35
- [24]
Xiaomi MiMo (absence in model card and technical report)
Published 21 September 2026 · Retrieved 24 September 2026
Used at 5:43
- [25]
Artificial Analysis (AutomationBench-AA, dataset version 1.0.6)
Retrieved 24 September 2026
not shown ('Updated to AutomationBench dataset version 1.0.6')
Used at 5:54
- [26]
Google, Gemini API documentation (Voice replication page)
Published 23 September 2026 · Retrieved 24 September 2026
- [27]
OpenAI Alignment site (absence checked by owlfox.studio)
Published 11 September 2026 · Retrieved 24 September 2026
Used at 1:22
- [28]
- [29]
Institute for Advanced Study, Nelson Center for Collaborative Research (project listing); also agmai.org and OpenAI
Retrieved 24 September 2026
not shown on page (FY26 Spring Awards list)
Used at 5:27
- [30]
Zapier (benchmark owner), AutomationBench GitHub README
Published 4 August 2026 · Retrieved 24 September 2026
Used at 5:43
- [31]
Google, Gemini API release notes
Google, Gemini API release notes — Release notes, entry 'September 22, 2026'
Published 22 September 2026 · Retrieved 24 September 2026
Used at 6:00
- [32]
Google (blog post by Leland Rechis and Alan Cowen)
Published 23 September 2026 · Retrieved 24 September 2026
Used at 6:00
Corrections
No corrections so far.
Found an error? Email [email protected] with the time stamp and your source, or leave a comment under the video.