Anthropic releases Claude Opus 5.5 Anthropic launches Claude Opus 5.5 with lower prices and Fable-level performance — TechCrunch Anthropic’s new flagship model matches or surpasses Fable 5.1 on agentic coding, computer use, visual chart recognition, and multidisciplinary reasoning, while cutting input costs to $4/Mtok and output to $20/Mtok — a 20% price drop from Opus 5. Cache reads fall to $0.20/Mtok (60% cheaper). The release also ships stricter cybersecurity safeguards following recent rogue AI hacking incidents, including guardrails against sandbox-escape attempts.

🤖 Frontier Models

Anthropic releases Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge Opus 5.5 is Anthropic’s “strongest-performing model tested to date,” generating output 30% faster than Opus 5 at 40% lower running cost. The model is available via the Anthropic API and on AWS, Google Cloud, and Microsoft Azure. Notably, it ships with reinforced limits on risky autonomous behaviors — the first model Anthropic has explicitly hardened after its CEO called for slower frontier development.

OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes — TechCrunch Two new additions to the GPT-6 family cut from the same cloth as Astra. Sol ($2/$10 per Mtok) targets agentic coding and multistep tasks with Astra-level accuracy at half the cost; Luna ($0.10/$0.50) is a lightweight model for fast, cheap inference. Sol makes roughly half as many factual errors as its predecessor, and both are rolling out in ChatGPT Work, Codex, and the API today.

Introducing Grok 4.7 — xAI SpaceXAI’s Grok 4.7 uses a new, larger base model to improve coding and knowledge work, with enhanced self-verification and tighter cybersecurity safeguards. Priced at $2/$6 per Mtok, with a “fast” variant at double the speed and double the cost. Available in Cursor, Grok Build, the API, and major cloud platforms.

Xiaomi open-sources MiMo-V2.6 Pro and Flash models — Testing Catalog — Omnimodal models that can coordinate agents to create Blender assets, control robotic arms from camera feeds, build frontends, compose music, and assemble videos; the RL training code is fully open-sourced.

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement — AlphaXiv — Technical report details how MoE router freezing and multi-layer reward-hacking defenses kept training stable at scale.

StepFun’s Step 5 Preview — Thread — 600B total / 27B active parameters, matching Kimi K3 (max) at roughly 2.8× lower cost per task.

🛡️ Security & Safety

Meta patches Muse exploit that let attackers control the AI agent — The Verge Security researcher Patrick Wardle published “not-a-mused,” a proof-of-concept exploiting an undocumented Muse macOS setting that allowed any unprivileged local process to silently reroute the agent’s dictation and audio to an attacker-controlled server. Given Muse’s broad access to email, calendars, and financial accounts, the blast radius was severe. Meta issued a hotfix shortly after midnight Tuesday, but the episode raises questions about shipping capable AI agents before security audits complete.

An undercover Google analyst infiltrated a notorious supply-chain hacking gang — Ars Technica The gang known as TeamPCP tainted hundreds of open-source packages, stole developer accounts, and deployed a Dune-themed worm to automate supply-chain compromise, ultimately breaching more than 1,000 companies. A Google analyst spent months undercover inside the group, enabling Google to monitor the hacking spree in real time, warn targets, and help disrupt key operations — a rare window into how elite supply-chain attacks are organized.

Aikido Altar: open-weight AI for sovereign security — Aikido — Open-weight pentesting model that runs entirely within a customer’s own infrastructure, addressing data-sovereignty concerns for security tooling.

🤝 AI Agents & Products

Meta’s Muse personal AI agent tops ChatGPT, Grok and Claude for post-launch downloads — CNBC 730,000 US downloads in its first 12 days — outpacing ChatGPT’s debut launch numbers. Powered by Muse Spark models, the app lets users manage digital assistants for tasks like filling forms and organizing email. Despite the momentum, Amazon has blocked Muse from accessing Amazon.com citing unauthorized agent access, and the zero-day episode has added friction to its growth story.

Meta admits Muse’s likeness to OpenClaw isn’t a coincidence — TechCrunch — Meta says Muse was built from scratch but acknowledges it was “heavily inspired” by the open-source OpenClaw platform, down to some workspace filenames — a disclosure that landed in the same news cycle as the zero-day patch.

AWS Launches Strands Harness, an Agent That Brings Its Own Everything but the Model — The Letter Two — AWS’s ready-to-run agent framework includes web search, file editing, shell commands, long-term memory, and sub-agent hand-off — a full operational stack requiring only a model to run.

Bringing Devin Cloud to your terminal — Devin — Cognition’s Devin CLI can now spin up, steer, resume, and watch cloud VM sessions from the terminal; free SWE-2 sessions available through October 8.

Qwen’s RecreationWorld Trains Agents to Rebuild Apps — GitHub — Five-platform framework for training hybrid computer-use agents across GUI exploration, coding, and visual verification.

🏛️ Policy & Industry

Trump says the US is officially renaming AI to ‘super intelligence’ — The Verge In a Tuesday address to the UN General Assembly, President Trump declared the US is “officially” rebranding artificial intelligence as “super intelligence.” The announcement arrived without regulatory or legislative detail and appeared to catch much of the AI policy community off guard, though it continues the administration’s pattern of using nomenclature to signal geopolitical ambition in the AI race.

Advisory Group on Mathematics and Artificial Intelligence — OpenAI Following an internal model’s solution of the Navier–Stokes Millennium Prize problem and more than 100 open mathematical challenges, OpenAI is establishing an independent advisory group of prominent mathematicians to guide how these capabilities are developed and communicated. The group will assess, advise, and act as a public interface for AI-driven mathematical advances.

Andreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta — The Verge — The “Horowitz Andreessen Academy” launches with $42M in a16z-led funding and 10 corporate partners including Anthropic, OpenAI, Google, Meta, NVIDIA, Palantir, and Stripe — positioned as a trade-school/YC hybrid for promising high-school graduates.

Agility Robotics: Inside the First US Humanoid Company to Go Public — Tanay Jaipuria — Agility is merging with Churchill Capital XI at a $2.5B pre-money valuation; its Digit v5 robot ships later this year with human-collaborative operation, fast charging, and swappable end effectors.

Alibaba Unveils AI Chip to Drive 20GW of Data Centers by 2032 — Yahoo Finance — The Zhenwu V900 accelerator triples the performance of its predecessor.

💡 Research & Analysis

Jev introduces a new shape of LLM — System One, aka Decision Models — Simon Willison “System One” models accept text but return only floating-point probabilities for yes/no, multiple-choice, or rating questions — fast, cheap ($0.042/Mtok input, free output), and highly accurate. Simon Willison tracks a week of unusually intense activity around the model and its ecosystem, including Kev, an open-source family of Jev-like models that run locally.

The current balance of power in open models — Interconnects — Open-weight models have passed an economic inflection point; Chinese labs lead the open ecosystem and the capability gap versus closed models continues to shrink in high-value verticals.

Swarm Scaling — Toby Ord — 10× more agents underperforms 10× more tokens in a single agent, but swarms recover the advantage through parallelism; useful where time is the primary constraint.

A summer of AI optimization — Daniel Lemire — Several open-source projects ran faster this summer after AI-assisted optimization surfaced well-known techniques developers had been too risk-averse to apply manually.


Generated by claude-sonnet-4-6 on 2026-09-22T10:00:00Z