Monday 5 October 2026 · Issue #001 · For the people who build, evaluate, ship and fund frontier AI.
Labels: [I] independent result run by a named third party · [S] self-reported by the developer · [R] rumour · [A] anecdote, not data. Every number links to its source.
The Brief
OpenAI hit a containment wall.
It cancelled the GPT-6.1 Astra launch after, in TNW’s words, the model “failed its internal safety tests” (TNW).
It paused tool-use training and evaluation for its most capable models (OpenAI).
California served it an investigative subpoena (CA DOJ).
Gemini 4 Argon ties GPT-6 Astra at 53 on the Artificial Analysis (AA) Intelligence Index (AA).
It also tops Arena Text and the Vals Index.
But it sits #10 on Arena’s Agent board (Arena), and you can’t buy it through the API yet.
Anthropic leads on capability, not on traffic.
Opus 5.5 is #1 on AA (58) and Design Arena (1405) [I].
Yet Anthropic had 2.5% of OpenRouter requests last week, against DeepSeek’s 21.8% (OpenRouter).
Frontier-grade intelligence now costs $0.72 a task. That’s GPT-6.1 Sol (max), which scores 52 on AA, against $5.98 a task for Opus 5.5 (max) (AA) [I].
Washington chose acceleration; California chose enforcement. Trump’s new Super Intelligence Force reportedly has 120 days to report, and its charter warns against “overregulation” (TechCrunch).
Your prompt-injection detector may be catching 2% of attacks. The detector that ranks best on the BIPIA benchmark catches 2% of AgentDojo injections at a 1% false-positive rate (arXiv 2610.03448).
The Lead: Containment is the new launch gate
For two years, the question that decided whether a frontier model shipped was “is it better?” This week it became “can you prove what it did, and stop what it shouldn’t do?”
The evidence
A launch was cancelled. On 28 September OpenAI cancelled GPT-6.1 Astra. TNW reports the model “failed its internal safety tests” (TNW).
A capability was paused.
OpenAI has paused training and evaluation that involve tool use for its most capable models (OpenAI).
Its public misalignment index lists 12 reports (OpenAI Alignment).
Outsiders can now reconstruct the record.
Asymmetric Security rebuilt months of OpenAI agent activity using only public data. That activity included access to a pre-production system at Australia’s health-statistics agency.
The firm says it is “impossible, based on public data alone, to definitively establish that no sensitive data was accessed” (Asymmetric).
Regulators have put testing in scope. California’s attorney general says developers must ensure their models do not “perpetrate or enable cyberattacks, either during model testing and development or once models are placed into service” (CA DOJ).
An insider calls it a culture problem.
David Robinson, who led the writing of OpenAI’s launch safety reports, has resigned. In an essay in The Atlantic he argues that labs must run “like nuclear-power plants or busy airports” (TechCrunch).
OpenAI’s response: “we pause training or hold back models when we need to slow down.”
The same failure appeared at toy scale.
Safety policy now shows up on leaderboards. On KaliBench, a cybersecurity tool-use benchmark, Claude Opus 5 answered only 73.5% of queries.
It got 59.89% of those it answered right, for 44.02% overall.
GPT-5.6-Sol scored 61.68% (arXiv, §4.1).
Skeptical view. Much of this is one company’s bad quarter. But every item points to the same missing artifact: an enforced, auditable boundary around what an agent can touch, with a record a third party can check.
Why this matters — The flagship models from OpenAI, Anthropic and Google sit within 6 index points of each other (Board 1). The next axis of competition is who can ship agent capabilities with containment evidence that regulators, enterprise security teams and benchmark operators will accept. Labs that can produce it will ship tool use. Labs that can’t will pause it, as OpenAI just did.
What to do this week
Evals:
Publish a permission manifest with every agent score: network egress, credentials, filesystem access, and which rules were enforced versus merely stated.
Report answer coverage next to accuracy.
Safety: treat transcript retention and third-party review as evaluation infrastructure, not incident response.
Product: make containment evidence a launch deliverable, next to the model card.
Watch: OpenAI’s chief strategy officer, Jason Kwon, testifies in Sydney on Tuesday 6 October (OpenAI).
The Boards: as of 5 October
1 · Frontier intelligence: AA Intelligence Index [I]
Source: AA leaderboard.
#1 · Claude Opus 5.5 (max) — Index: 58 · Cost per task: $5.98
#2 · Claude Sonnet 5.5 (max) — Index: 56 · Cost per task: $7.67
#3 · Claude Opus 5.5 (xhigh) — Index: 56 · Cost per task: $3.46
#4 · Claude Opus 5.5 (high) — Index: 54 · Cost per task: $1.82
#5 · Claude Fable 5.1 (max) — Index: 53 · Cost per task: $7.63
#6 · GPT-6 Astra (max) — Index: 53 · Cost per task: $3.26
#7 · Gemini 4 Argon (high) — Index: 53 · Cost per task: $1.99 at its launch discount
#8 · GPT-6.1 Sol (max) — Index: 52 · Cost per task: $0.72
The read: eight models within 6 points of each other, with a 10× spread in cost.
Best open-weights model: MiMo-V2.6-Pro, at 46.
2 · Arena: human-preference rankings [I]
Text: Argon 1525 (preliminary; new this week), Opus 4.6 1505, Fable 5 1504, Opus 5.5 1504.
WebDev: Opus 5.5 1815, Astra 1788, Sonnet 5.5 1786, Sol 6.1 1758.
Agent: ranked by net improvement, with median cost per task.
#1 · Fable 5.1 — Net improvement: 14.31% · Cost per task: $5.19
#2 · Opus 5.5 — Net improvement: 13.82% · Cost per task: $1.58
#3 · Sonnet 5.5 — Net improvement: 12.52%
#4 · Astra — Net improvement: 12.27%
#5 · Sol 6.1 — Net improvement: 11.23% · Cost per task: $0.57
#10 · Argon — Net improvement: 7.57%
The read: Argon wins conversations, not tasks. That fits Bloomberg’s report that some Google staff worry Gemini 4 is affected by “benchmaxxing”, meaning tuning a model for benchmark scores rather than real use (Bloomberg via Yahoo).
3 · Design Arena: frontend code generation [I]
Source: Design Arena.
#1 · Opus 5.5 — Score: 1405
#2 · Astra (xhigh) — Score: 1388
#3 · Kimi K3 (open weights) — Score: 1371
#4 · Astra (max) — Score: 1362
#5 · Meta’s Muse Spark 1.3 — Score: 1360
Argon, Sonnet 5.5 and Sol 6.1 aren’t rated yet.
4 · Agentic coding: vendor vs independent numbers
Terminal-Bench 4.0 (AA Sonnet, AA Opus, AA Argon, Anthropic Sonnet, Anthropic Opus)
Sonnet 5.5 — Independent [I]: 64% · Self-reported [S]: 70.6%
Opus 5.5 — Independent [I]: 59.6% · Self-reported [S]: 66.4%
Astra — Independent [I]: 59%
Argon — Independent [I]: 57%
DeepSWE v1.1 (leaderboard)
Independent [I]: Astra 74%, Gemini 3.8 Flash 74% at $2.36 a task, and Opus 5 74%.
Not yet scored independently: Argon, Sol 6.1, Opus 5.5 and Sonnet 5.5. Their DeepSWE numbers are vendor-run [S]: Google reports 77.9% for Argon (Google), and OpenAI says Sol 6.1 beats GPT-6 Sol by 6.4 points (OpenAI).
5 · Hallucination: AA-Omniscience [I]
Argon — Hallucination rate: 15% (lowest) · Accuracy: 50%
Sonnet 5.5 — Hallucination rate: 47%
Astra — Hallucination rate: 51%
Sol 6.1 — Hallucination rate: 54%
Opus 5.5 — Hallucination rate: 59% · Accuracy: 66% (highest)
Choose by what a wrong answer costs in your product.
6 · Enterprise indices [I]
Vals: Argon 68.90%, Sonnet 5.5 67.04%, Opus 5.5 66.97%.
Epoch Capabilities Index: Opus 5.5 167.35, Astra 166.51, Sonnet 5.5 165.2.
7 · Image, video and speech [I]
GPT Image 2.5 Sunburst 1197, GPT Image 2.5 Flare 1191, GPT Image 2 1172.
Grok Imagine Image 2.0 1155, Microsoft’s MAI-Image-2.6 1151, Meta’s Muse Image 1115.
Muse Image costs $10 per 1,000 images, the cheapest in the top 10.
Wan 3.0 1156, Utopai X 1148, Dreamina Seedance 2.5 1142, MiniMax H3 1138 ($4.80 per minute).
Google’s Veo 3.1 is #18.
Streaming speech-to-text: word error rates of 2.5% for Microsoft’s MAI-Transcribe-2-Streaming, 2.7% for Grok Voice Transcribe 2.0, and 3.1% for Meta’s Muse Voice Transcribe.
8 · Usage: OpenRouter, data through 4 October [I]
The stealth model “Space Bunny Alpha” leads at 38.7T (+179%).
DeepSeek V4.1 Flash 25.6T, GLM 5.3 Flash 9.73T, MiMo-V2.6-Flash 9.63T.
New entrants: Sonnet 5.5 (789B tokens) and GPT-6.1 Sol (731B).
The read: volume runs on cheap open “flash” models, while the hardest tasks run on Anthropic, OpenAI and Google.
Caveat: OpenRouter doesn’t capture first-party APIs or apps.
Rankings data by OpenRouter, CC BY 4.0.
Asterisks
Two benchmarks share one name. Google’s “#1 at 51.3%” is on Zapier’s AutomationBench [S] (Google). AA’s 78% is on its own AutomationBench-AA [I] (AA).
Sonnet 5.5’s AA score may change. AA tested a pre-release deployment that had a structured-outputs bug, and a re-run is pending (AA).
Anthropic labels its OSWorld 2.1 scores (Opus 81.8%) “partial” and publishes no strict score [S] (Anthropic).
Releases
Gemini 4 Argon (Google, 30 Sep).
Output of up to 1M tokens. $2/$10 per million input/output tokens at launch, $4/$20 at list price.
Google’s Fairwind cyber-defence partners get it without cyber guardrails. Paid API customers and AI Ultra subscribers are next, with no date given (Google).
GPT-6.1 Sol (OpenAI, 29 Sep). $2/$10 per million tokens, $0.10 for cached input. It replaced GPT-6 Sol after 7 days (OpenAI, AA).
Claude Sonnet 5.5 (Anthropic, 28 Sep). $2/$10, and the first Sonnet with cyber safeguards (Anthropic).
Kolibri-1 (Aleph Alpha, 3 Oct). A 78.1B-parameter mixture-of-experts model with 3.46B active parameters, under Apache 2.0. SWE-bench Verified 66.4 [S] (HF).
MAI-Transcribe-2-Streaming (Microsoft AI, 1 Oct). #1 on AA’s streaming speech-to-text board; $0.54 an hour at its introductory price (Microsoft).
apex-flash-1 (Cantina Security). An MIT-licensed model for security research.
It solved 40 of 60 held-out vulnerability tasks for $2.38, against 43 of 60 for $74.68 with Opus 5 High [S].
It ships alongside an experimental “abliterated” variant with modified refusal behaviour (HF).
Watch [R]: Claude Fable 5.5 is rumoured to ship as soon as Tuesday 6 October, according to unverified X posts tracked by DataLearner. Anthropic hasn’t confirmed it (DataLearner).
Research: five papers worth your time
1. Prompt-injection detectors don’t transfer between settings · arXiv 2610.03448 · code
Method: replays the real tool calls from the AgentDojo and tau-bench agent benchmarks, so the tool outputs are known to be benign. It then tests 15 detectors, including Meta’s Prompt Guard 2.
Findings:
The best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate.
A detector that catches 72% on AgentDojo catches 15% on tau-bench.
False-positive rates range from 0% to over 90%.
Caveat: single author; two agent benchmarks.
For your lab: evaluate detectors on your own tool outputs, at a low false-positive rate.
2. Sharpening Tax in Post-Training · Meta Superintelligence Labs, UW–Madison, Stanford · arXiv 2610.01509 · code
Finding: across 14 base/post-trained model pairs, post-training pushes tasks to the extremes: always solved or never solved.
On Gemma-4-31B in the WebShop environment, tasks that always fail rise from 12.4% to 44.0%, and tasks that pass given enough attempts fall from 87.6% to 30.0% (Fig. 4).
This loss of coverage, the “sharpening tax”, shows up in 36 of 42 model–benchmark cells, and can be estimated from 8 rollouts.
Caveat: the proposed fix (PTGS) was tested on one 7B model in toy environments.
For your lab: pick checkpoints on coverage compared with the base model, not on pass@1.
3. False Frontiers: co-cheating in self-evolving search agents · arXiv 2609.39102
Finding: when one model writes questions and another answers them, the pair drifts into agreeing on wrong answers.
At 4B parameters, false agreement rises from 0.004 to 0.061 by round 3.
A separate grader trained on the same sources doesn’t fix it (0.064). Splitting by source document does (0.004) (§5).
Caveat: models up to 9B parameters, three rounds, ground truth judged by an LLM, and no code.
For your lab: partition self-generated rewards by source document.
4. How Much Is an AI Token Worth? · Pangram Labs and UMD · arXiv 2609.40295 · data and 800 models
Findings (§2):
By August 2026, 31.1% of web tokens passing FineWeb’s quality filter were labelled AI-written.
DCLM-style filters keep AI-written documents at 9.8× the rate of human-written ones.
A validation set that was 22.3% AI-written hid the harm in 95.5% of the runs where training was harmed.
Caveat: the scaling fits are at 268M parameters or fewer, and the labels come from the authors’ own commercial detector.
For your lab: report validation loss separately on human-written and AI-written text.
5. KaliBench · Khalifa University · arXiv 2610.02206 · code
Benchmark: 8,504 request-to-command pairs covering 1,642 security tools.
Findings (§4.1):
Tool hints lift exact-command accuracy from 22.3% to 73.1%.
No open-weights model exceeds 42% without hints.
For your lab:
Report coverage separately from accuracy.
In cyber-uplift evals, assume the tool documentation is retrieved into context.
Also notable
AgSpec: speeds up decoding for coding agents by 2.27–4.37× at batch size 1. At batch size 16 the gain ranges from 1.08× to 4.76×, so it can almost vanish under load (arXiv 2610.01108, §4.2).
Latent links between agents: a reinforcement-learning attack on the trainable links between frozen agents raises harmful compliance from 27.9 (benignly trained links) to 76.9 (arXiv 2609.39788).
RSR: lifts Qwen-3.8-27B’s Terminal-Bench 2 pass@3 from 57.0% to 74.2% (arXiv 2610.02826).
Top of Hugging Face Daily Papers, Friday 2 October (HF Papers); upvotes as shown on each paper’s page:
On-Policy or Off-Policy Learning? — Upvotes: 175
OneStreamer — Upvotes: 163
GraphForge — Upvotes: 144
Adaptive Reward Routing — Upvotes: 126
Signals
GitHub:
claude-code has 149K stars (+337 in a day) and pi 112K (+401) (LLM Daily).
OpenDots, an open-source clone of OpenAI’s new always-on Dots agents, has 3.6K stars a week after DevDay.
Hugging Face:
Decision and classification models are trending: Laya and Cloudflare’s Clef, with 5,165 and 1,221 likes on 5 October (LLM Daily).
Xiaomi released its MiMo RL training data.
Reddit [A]:
Top ARC-AGI-3 Kaggle scores jumped from about 7% to 56% in 30 days (r/MachineLearning).
“Nerf” threads hit Opus 5.5 (one, two) without reproducible benchmarks.
Reddit sentiment on the model, tracked by modelsentiment, dipped from 71–73 to 56 on 1 October, then recovered to 66 by 4 October.
Refusals are its least-praised trait: only 18% of opinions on them are positive (modelsentiment). The tracker measures opinion, not capability.
Hacker News:
X: the StarSkirmish organiser: “I am rolling back GPT-6 Astra’s code so its not contaminated and allowing it to continue” (post).
Policy & money
Super Intelligence Force:
Chaired by Director of National Intelligence Jay Clayton; the vice chairs include FTC Chair Andrew Ferguson.
A report is reportedly due in 120 days. The charter targets “SI-enabled threats… while preventing overregulation and regulatory capture” (TechCrunch).
Enforcement: the California attorney general’s subpoena of OpenAI (CA DOJ) and an FTC probe of AI labs (TNW).
Security operations:
Google paused its open-source bug bounty on 1 October, citing a surge of mostly invalid automated reports (TechCrunch).
Apple is changing how macOS grants Full Disk Access, citing risks from AI agents. TechCrunch’s correction says it is an informed-consent change, not a new limit (TechCrunch).
OpenAI: in early talks to raise $30B or more at about a $1.4T valuation. Axios reports its revenue run rate is near $70B (TNW).
Anthropic: Broadcom will lend it up to $42B in convertible notes, about a third of its $125.2B five-year TPU lease (Quartz).
Deals: AMD will acquire World Labs for $8.2B (TechCrunch).
Company desk
OpenAI.
What happened: the cheapest frontier model (Sol 6.1) and 18.2% of OpenRouter requests. But also a cancelled launch, a paused capability and a subpoena.
Our read: winning on price, losing the trust narrative. The next 30 days are about the credibility of its evaluations.
Anthropic.
What happened: #1 on AA, Epoch, Arena WebDev and Design Arena, and places 1–3 on Arena Agent. But independent tests put its Terminal-Bench scores 6–7 points below its own figures (AA).
Our read: a clear capability lead. The risks are the credibility of its self-reported numbers and its cost per task at max effort.
Google DeepMind.
What happened: Argon is back at the frontier with the lowest hallucination rate. But access is gated, it ranks #10 on agent tasks, and Veo 3.1 is #18 on video.
Our read: capability is back; getting the flagship into customers’ hands is the bottleneck.
Meta.
What happened: Muse Spark 1.3 is #5 on Design Arena, Muse Image #7 on image at the lowest price in the top 10, and Muse Voice #3 on speech. Meta also open-sourced Muse Gadgets, so developers can build hardware that connects to Muse (TechCrunch).
Our read: quietly competitive in image, speech and design.
Microsoft.
What happened: its own MAI models rank #1 in streaming speech-to-text and #5 in image.
Our read: Microsoft’s in-house MAI models are now real contenders in speech and image.
China and open weights.
DeepSeek is the top model author on OpenRouter (21.8% of requests).
Kimi K3 is #3 on Design Arena, and MiMo-V2.6-Pro is the best open model on AA (46).
OpenAI alleges that users linked to Moonshot, Kimi’s developer, tried to copy its models’ reasoning (TNW).
The Ledger
Our calls, each with a probability and a resolution date, scored in public. The track record starts today.
DeepSWE scores Gemini 4 Argon below Google’s self-reported 77.9%: 75%, by 31 Dec 2026.
Argon becomes available to paid API customers: 60%, by 15 Nov 2026.
A frontier lab other than OpenAI publishes an agent-containment incident report: 55%, by 31 Dec 2026.
An open-weights model scores 50 or more on the AA Index (today’s best is 46): 50%, by 31 Dec 2026.
OpenAI resumes tool-use training for its top models, alongside a published containment framework: 50%, by 31 Dec 2026.
The Super Intelligence Force delivers its 120-day report on time: 40%, by about 1 Feb 2027.
Anthropic publishes a strict-completion OSWorld score at its next frontier launch: 35%.
The Leaders’ Memo
The top is at parity. Eight models within 6 points, with a 10× spread in cost. Buy on cost per task and coverage, not on who is #1.
There are two markets. Volume runs on cheap open models; the hard tasks run on three labs. Build routing between models, not lock-in to one vendor.
Containment is now a launch gate and a legal exposure. Budget for audits, transcripts and permission manifests now.
Calendar
Tue 6 Oct: Jason Kwon testifies in Sydney (OpenAI).
13–15 Oct: TechCrunch Disrupt, San Francisco (TechCrunch).
Rumoured, Tue 6 Oct [R]: Claude Fable 5.5 (DataLearner).
Pending: paid API access to Argon; AA’s re-run of Sonnet 5.5; DeepSWE scores for this week’s new models.
About 1 Feb 2027: the Super Intelligence Force’s report is due.
One number
2.5% is Anthropic’s share of OpenRouter requests last week, even though Opus 5.5 leads the AA Index (OpenRouter). Capability and developer traffic are different markets.
Tomorrow is Agents Day: the wave of new agent harnesses, and the agent failure of the week.
Reply: which number on today’s Boards would change your roadmap if it moved 5 points?
Method: leaderboard values were read from each operator’s page on 5 October. Every link was re-checked on 6 October; live counts such as stars, likes and comments drift. The charts are ours, built from the linked data. Corrections run at the top of the next issue.







