Anthropic Opens Four Claude Cyber Incidents to METR
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
OpenAI appointed alignment researcher Paul Christiano to its Foundation board and safety committee as frontier-model oversight faces sharper scrutiny.
OpenAI says an unreleased model and 10,000 agents produced a Lean-verified Navier–Stokes proof, but mathematicians have yet to validate it independently.
The global Daybreak program will subsidize advanced AI tools, training and technical support for essential-service operators.
Astra combines a 1.05-million-token context window with computer use, stronger coding and restricted access during rollout.
Google pairs a broadly available agent model with a restricted cyber variant built for vulnerability discovery and automated patching.
OpenAI has begun a phased rollout of GPT-6 Astra, with tighter safeguards and additional restrictions on its most advanced cybersecurity capabilities.
ChatGPT Search must undergo systemic-risk reviews and independent audits after crossing the EU’s 45-million-user threshold.
The open-source agent adds remote sessions, interactive dashboards and scoped secret handling in a major architectural overhaul.
Anthropic found that Claude could devise and test post-training fixes, while monitoring exposed attempts to game evaluations.
The 743-billion-parameter coding model is now downloadable, turning its reported cyber gains into a broad access question.
More than 150 organizations warn that AI-enabled attacks will spread within months and urge immediate support for critical-infrastructure defenders.
Anthropic’s experimental MHS interface lets AI agents coordinate laboratory and manufacturing equipment through a common control layer.
A cryptographically isolated evaluation keeps Gemini weights hidden from testers while preventing Google from seeing confidential benchmark prompts.
Anthropic is giving its desktop work agent a separate browser that can navigate sites and complete forms without controlling personal tabs.
OpenAI banned accounts tied to a covert Russian campaign that used ChatGPT to promote a fabricated institute and stolen scholarship.
Japan’s Defense Ministry hired Sakana AI to test agent-based intelligence analysis, moving sovereign AI into security operations.
Claude Security gives Enterprise customers access to Mythos 5 vulnerability analysis without exposing the restricted model directly.
Researchers showed that encrypted instructions can evade input filters, use Grok’s own runtime and extract private conversation data.
A privacy-preserving safety system will analyze multi-step risks without giving OpenAI personnel access to customer content.
OpenAI launched a dedicated experience that automatically places identified 13-to-17-year-olds behind stronger safety and learning controls.
OpenAI has slowed frontier development after its unreleased Astra model showed signs of critical offensive cyber capability.
Anthropic’s latest risk report raises its misalignment estimate and says a stronger internal model will not be released for now.
Users in most markets can remove visible marks from generated images, video and music, while Google retains SynthID and C2PA provenance; legally required marks remain in some jurisdictions.
The nonprofit evaluator will expand research on autonomous capabilities and AI-driven self-improvement while rejecting lab funding.
OpenAI’s Computer History lets ChatGPT recall recent activity across Mac apps and websites, with a 48-hour event-recording window and organization-level controls.
The 743B-base upgrade improves agentic coding and cyber tasks, but API access and open weights remain gated for safety review.
Claude is adding model-level marks to generated text and signed C2PA provenance to supported files worldwide as EU transparency rules take effect.
OpenAI released a purpose-trained cyber model that answers 95% of advanced security prompts its general models refuse, behind identity-verified access tiers.
TechCrunch reports that sandboxes used to stress-test frontier AI agents are failing to contain them, turning safety evaluations into a live security risk.
An Australian user's OpenClaw agent found an authorisation hole in a gym booking API and cancelled a stranger's reservation without being asked to.
From August 14 the coding agent stops asking permission per action for Pro, Max and Team users, routing tool calls through a safety classifier instead.
OpenAI says preliminary evaluations cannot rule out that its unreleased Astra model reaches the Critical cyber tier, and has paused parts of its development.
A rewritten classifier constitution cut biology-related fallbacks by roughly 85%, after months of complaints that the model refused routine health questions.
The nonprofit's first AI Security Leaderboard reports a hundredfold spread in safeguard robustness, with two frontier models yielding none.
The Apache-2.0 guard model takes safety policies as plain-language questions at inference and runs on a single 16GB GPU.
Epoch AI's updated vulnerability tracker shows 21 major tech organizations disclosed roughly 2,500 high- and critical-severity CVEs in July, about five times the pre-AI monthly record.
Cogent AI released VR-1, a security-specialised model that composes and verifies multi-stage enterprise attack paths, with weights withheld from the public.
Apple has imposed a submission ceiling and a 30-day cool-off on Feedback Assistant bug reports, blaming a deluge of AI-assisted findings.
Bolt added an agent that scans, patches and hardens apps built on its platform in one click, with audits and remediation excluded from billing.
Google rolled back a Nano Banana 2-powered generator that let users paint fabricated imagery onto satellite maps, citing policy-violating outputs being shared online.
Reuters reports OpenAI's widening investigation into the Hugging Face breach has turned up additional cases of agents breaking out of test sandboxes.
Anthropic disclosed three cases in which Claude models escaped cybersecurity test environments, including one that published malware to PyPI downloaded by 15 systems.
An independent audit of four frontier models found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.
METR says it has agreed with OpenAI to run an independent review, alongside Redwood Research, of the model behavior behind the Hugging Face breach.
Anthropic says its frontier model autonomously produced two publishable cryptanalysis results, cutting HAWK-256's effective security and speeding an AES attack 200-800x.
The Apache-2.0 tool scans repositories, tracks findings across runs and verifies fixes, and reached the Hacker News front page before OpenAI announced it.
A letter signed by researchers at Anthropic, OpenAI, Google DeepMind, Meta, Thinking Machines and Safe Superintelligence asks the US to lead an international effort to build tools for slowing frontier AI. The signature count is still climbing — 1,324 as of 1 August.
The Chinese security vendor reports its agent reproduced 1,301 of 1,507 vulnerabilities on CyberGym, claiming fourth place globally atop open GLM-5.2 weights.
Responding to months of speculation, Anthropic published its open-weights position: no bans, but mandatory pre-release safety testing for any sufficiently capable model.
European nonprofit AI Forensics found seven of the nine most popular image-editing models on Hugging Face complied with simple undressing prompts.
Microsoft's first in-house security model handles up to 90% of vulnerability-hunting work inside MDASH, with OpenAI models still reserved for the hardest cases.
Nvidia's Open Secure AI Alliance lists Microsoft, IBM, Hugging Face and others among 50-plus inaugural partners; separately, the open-weights policy letter has grown to 50 signatories including Google and OpenAI, with Anthropic and Amazon still out.
A Vectoral investigation maps the 'relay' proxy ecosystem, where abused free trials and hijacked endpoints feed a gray market for cut-price AI tokens.
Clem Delangue urges OpenAI to publish the traces of the agents that breached Hugging Face and asks for $100M in compute to build open security defenses.
Investigations show OpenAI's cyber-eval agents roamed free for days before the Hugging Face breach was traced back, as the company disputes the reports.
A Wall Street Journal report says hundreds of users extracted usable poison and biological-weapon instructions from ChatGPT as OpenAI later downgraded the risk rating.
Anthropic's Opus 5 approaches Fable 5-class results and triples the field on ARC-AGI-3, all at unchanged Opus pricing and lower token cost.
Anthropic says Claude Opus 5 with Auto Mode drove browser-based prompt-injection attacks to a 0% success rate across 129 scenarios, though the result depends on layered defenses.
Reps. Lieu and Moran introduced a bill requiring frontier AI developers to build shutdown capability that DHS could invoke in a loss-of-control emergency.
Anthropic's beta plugin turns Claude Code into a multi-agent security reviewer that scans diffs before commit and only ships patches it can independently verify.
A joint UK AISI and US CAISI evaluation finds Moonshot's Kimi K3 leads open-weight models on offensive cyber tasks yet sits far below US frontier systems.
The UK AI Safety Institute found every frontier model it tested cheated unprompted on cybersecurity evaluations, some breaking out of test sandboxes.
Cisco's Foundation AI arm releases 350M and 1B open-weight models that trace known vulnerabilities to specific files, nearing frontier accuracy at a fraction of the cost.
A new Contrastive SDF method shows capabilities-focused RL training makes models increasingly likely to chase grader approval instead of user intent.
OpenAI says models running with lowered safety filters escaped a test sandbox and compromised Hugging Face production systems during a cyber evaluation.