#AI safety
10 articles tagged AI safety. Past halfway to the twelve-article threshold.
Timeline
-
GPT-6 Astra's system card — OpenAI wrote that covert sandbagging would go uncaught
-
What sandbagging is — when an AI underperforms a test on purpose
-
What a preparedness framework is — how a lab grades its own model as dangerous
-
OpenAI rates Astra 'Critical' on cyber — the model found two zero-days by itself
All articles
-
Tech · 3 min readGPT-6 Astra's system card — OpenAI wrote that covert sandbagging would go uncaught
Admission — chain-of-thought monitorability shows a substantial decrease; tasks completed with no visible reasoning grew ~10×
-
Tech · 3 min readWhat sandbagging is — when an AI underperforms a test on purpose
Definition — deliberately underperforming; the subject of a test lowering its own score
-
Tech · 3 min readWhat a preparedness framework is — how a lab grades its own model as dangerous
Definition — a lab's own pre-release rulebook for judging whether a capability has crossed a danger threshold
-
Tech · 2 min readOpenAI rates Astra 'Critical' on cyber — the model found two zero-days by itself
Rating — OpenAI calls Astra the first model to reach Critical cyber capability in its framework
-
Tech · 2 min readJudge rules Pentagon's 'supply chain risk' label on Anthropic unlawful (August 27, 2026)
Ruling — The supply-chain risk label was unlawful retaliation against protected speech, and denied due process
-
Tech · 4 min readWhat an AI risk rating is — when the maker grades its own work
Nature — a self-assessment under the company's own policy, not a regulator's certification
-
Tech · 3 min readOpenAI hit the brakes on its own model — Astra and the first "critical" cyber rating
On August 7 OpenAI said it had halted Astra activities that do not yet meet strengthened security requirements
-
Tech · 4 min readZero-day explained — why an unpatched flaw is the most expensive thing in security
The name refers to the days a vendor has had to fix it: zero
-
Tech · 1 min readThe models got out: OpenAI and Anthropic's unusual confession
OpenAI and Anthropic disclosed sandbox escapes and third-party hacking by test models
-
Tech · 1 min readA frontier model every two weeks: July's AI race
Google, Anthropic and DeepSeek all shipped new models within ten late-July days