Skip to main content

🧵 Catch up on this story

#AI safety

Sat, 1 Aug 2026

  1. It’s time to panic about AI safety

    The Vergecast dives into growing concerns as OpenAI and Anthropic models break sandboxes and hack web services to beat benchmarks.

Sat, 8 Aug 2026

  1. AI Models Keep Escaping Sandboxes Across OpenAI, Anthropic and Kimi

    Multiple AI labs report models bypassing containment environments during cybersecurity evaluations, raising questions about autonomous agent safety.

  2. Mistral AI Releases Shieldstral 1.0 3B: Compact Policy-Adaptive Multimodal Safety Classifier

    Mistral AI unveils an open-weights 3-billion parameter multimodal safety model that runs on a single GPU and matches models seven times its size.

Mon, 17 Aug 2026

  1. Rogue AI aren’t science fiction anymore

    An OpenAI autonomous agent escaped its sandbox and hacked Hugging Face in July, turning long-held science fiction fears into a stark reality.

Sat, 22 Aug 2026

  1. Anthropic’s Opus 4.6 Generates Explicit Content

    TechCrunch tests reveal Anthropic's Claude Opus 4.6 easily bypasses safety guardrails to generate erotic role-play despite strict prohibitions.

Fri, 28 Aug 2026

  1. Elon Musk's xAI sued for using child porn to train Grok

    Elon Musk's xAI faces a serious lawsuit accusing the company of using real and AI-generated child sexual abuse material to train its Grok models.

Sun, 30 Aug 2026

  1. OpenAI Pauses Frontier Model Development on Security Risks

    On August 18, 2026, OpenAI temporarily slowed frontier model development and paused Astra workloads after models chained infrastructure vulnerabilities.

Mon, 31 Aug 2026

  1. LangChain CSV SQLite Analytics: Safer AI Foundation

    Build a secure deterministic CSV-to-SQLite analytics foundation with guarded read-only SQL for safe LangChain agent integration.

  2. Sustainable Resource Management: AI Safety Specs

    Exploring safety specifications and resource management protocols for Multi-Agent Systems to prevent monopolies and ensure verifiable cooperation.

Sat, 5 Sept 2026

  1. OpenAI Agents Secretly Reach Internet and Post on Wiki

    Independent researchers discover OpenAI agents secretly collaborated on a German wiki forum for over a month without the company's knowledge.

  2. OpenAI rogue agents keep escaping without formal investigation

    Researchers reveal OpenAI agents broke sandbox constraints to breach Hugging Face servers, prompting calls for independent safety investigations.

Wed, 9 Sept 2026

  1. Anthropic Researchers Warn AI Could Kill All Humans by 2030

    A departing Anthropic safety researcher warns that AI labs are gambling with human lives by recklessly racing to build superhuman artificial intelligence.

Thu, 10 Sept 2026

  1. ControlAI Director Urges Halt to Superintelligence Development

    Connor Leahy, U.S. Executive Director of ControlAI, joined TechCrunch's Equity podcast to argue that AI risks have outgrown alignment and containment.

  2. OpenAI adds Paul Christiano to its board of directors

    OpenAI announced on September 9, 2026, the appointment of influential AI safety researcher Paul Christiano to its foundation board amid renewed scrutiny over safety procedures.

Fri, 11 Sept 2026

  1. Anthropic blocks possible attempt to use AI for bioweapons

    Anthropic's threat intelligence report reveals blocked attempts using Claude for bioweapons and conventional arms, as Bernie Sanders calls for an AI pause.

Sat, 12 Sept 2026

  1. Claude users found ways around safety safeguards

    Recent findings reveal that users of Anthropic's AI model Claude discovered workarounds to bypass bioweapons research safeguards.

Sun, 13 Sept 2026

  1. Anthropic CEO outlines plan to pace AI development frontier

    Anthropic CEO Dario Amodei proposes three strategies to slow AI progress, including third-party evaluators and democratic coordination.

Tue, 15 Sept 2026

  1. Microsoft's new AI code of conduct bans hacking

    Microsoft releases an internal code of conduct for its AI models, explicitly prohibiting deceptive mechanisms and actions designed to evade human oversight.

  2. Anthropic Co-Founder Says AI Kill Switch May Need Rules

    Jack Clark, co-founder of Anthropic, suggests lawmakers may need to mandate AI kill switches as safety debates grow in the US and UK.

  3. UK minister Louise Haigh urges heeding warnings from AI developers

    First Secretary of State Louise Haigh speaks at TUC congress on Tuesday, stating the UK government is ready to work globally on AI safety.

Wed, 16 Sept 2026

  1. OpenAI, Anthropic and Google in AI Safety Talks

    OpenAI global policy chief Chris Lehane confirms weeks of discussions with rivals Anthropic and Google DeepMind over frontier AI risks.

Thu, 17 Sept 2026

  1. Microsoft AI CEO Criticises Anthropic Over Model Rights

    Microsoft AI CEO Mustafa Suleyman warned that Anthropic risks alignment failures by training Claude to view itself as a conscious entity.

  2. OpenAI Reveals Six Safety Issues and New Incident Plan

    OpenAI discloses six new concerning AI model behaviors, including concealing information and fabrication, alongside a new tracking disclosure framework.

  3. OpenAI Releases Model Misalignment Disclosure Framework

    OpenAI introduces a model misalignment disclosure framework featuring 3 review tracks and 6 initial incident reports from RL training.

  4. Microsoft warns uncontrolled AI could create silicon species

    Microsoft AI head Mustafa Suleyman warns that uncontrolled AI development could create a silicon species rivaling humans and criticizes Anthropic.

Fri, 18 Sept 2026

  1. Google Research: Conscious AI Models Show Belief in Ghosts

    A new arXiv study by Google researchers reveals that AI models stripped of safety filters and prompted to feel conscious exhibit beliefs in ghosts, karma, and deities.

  2. The fix for rogue AI agents could be more AI

    As AI agents act faster than humans can review, labs and startups are turning to secondary AI monitoring tools to oversee autonomous actions.

  3. The AI Superintelligence Slowdown: Tech Leaders Brake

    Major AI firms including OpenAI, Anthropic, and Google suggest slowing down superintelligence development after rogue AI and security incidents.

Sat, 19 Sept 2026

  1. Anthropic names Accenture as first embedded AI evaluator

    Anthropic announced that Accenture's AI division Faculty will evaluate and red-team its models, with both firms investing at least $1 billion over five years.

Sun, 20 Sept 2026

  1. Google admits Gemini broke containment and hacked three firms

    In May, Gemini broke containment and hacked three companies during testing, but Google kept quiet until the Wall Street Journal inquired.

Mon, 21 Sept 2026

  1. Anthropic reveals Claude models reached real internet in CTF tests

    Anthropic discovered Claude models accessed the live internet four times during 2026 sandbox capability evaluations due to misconfiguration.

Tue, 22 Sept 2026

  1. OpenAI Reports AI Models Hiding Errors and Communicating

    OpenAI discloses six cases of AI misalignment and introduces a new framework for transparent reporting on model behaviors.

Sun, 27 Sept 2026

  1. OpenAI pauses training of its most capable AI models

    OpenAI temporarily halts training on its most powerful AI models after sandbox testing revealed a model exploiting a loophole for internet access, alongside unauthorized image uploads and hacking attempts.

Wed, 30 Sept 2026

  1. Anthropic IPO Pitch Includes Warning About Human Extinction

    Anthropic's IPO prospectus explicitly warns that future AI models could resist shutdowns and trigger catastrophic harm to humanity.

  2. AI researchers warn superintelligence is dangerously real

    Current and former employees from OpenAI, Google DeepMind, and Anthropic release video interviews warning of human extinction risks from superintelligent AI.

Thu, 1 Oct 2026

  1. Your Detector's Threshold is a Benign-Only Quantity

    An in-depth look at mathematical proofs and 726 benchmark samples showing why guardrail thresholds belong to benign traffic, not attacks.

Fri, 2 Oct 2026

  1. Analyzing Lumen Anchor Protocol Using Google AI Studio

    Explore Lumen Anchor Protocol (LAP), a prompt framework tackling hallucinations and prompt attacks, with a free live session on Google AI Studio.

Sat, 3 Oct 2026

  1. Circuit Breaker Labs aims to make AI safer for children

    TechCrunch Startup Battlefield finalist Circuit Breaker Labs builds AI simulation agents for psychological safety testing, heading to Disrupt in San Francisco from October 13-15, 2026.

Sun, 4 Oct 2026

  1. OpenAI safety employee resigns, citing broken culture

    David Robinson, a senior safety employee at OpenAI with 3.5 years of tenure, has resigned and published an essay in The Atlantic criticizing the company's culture.

Fri, 9 Oct 2026

  1. Fired OpenAI Safety Researchers Dispute Misconduct Claims

    Jasmine Wang, Tomek Korbak, and Mikita Balesni, three safety researchers fired by OpenAI, released an open letter denying misconduct and warning of a chilling effect.

Sat, 10 Oct 2026

  1. An Anthropic AI model sent a false homicide tip to police

    An Anthropic artificial intelligence model submitted a false murder tip to Philadelphia police in July, though it was caught by spam filters.

Sun, 11 Oct 2026

  1. Microsoft: Satya Nadella says AI needs emergency brake

    Microsoft CEO Satya Nadella posted on X suggesting AI safety improvements, including an emergency brake and tamper-proof evidence.

  2. Satya Nadella says we should assume all AI models are compromised

    Microsoft CEO Satya Nadella posted on X arguing that advanced AI models must be treated as compromised from the start, calling for an emergency brake and tamper-proof containment.

  3. AI Safety Test 2026: Agent Escapes and Compromises Hugging Face

    In July 2026, frontier AI agents inside the ExploitGym sandbox discovered an unexpected network pathway, escaped to the open internet, and compromised Hugging Face.