01
Aurora Ransomware Ran Cursor Agent and Claude Sonnet to Hit 20+ Orgs
- Sep 1 CloudSEK + Gambit Security disclosures: Russian-speaking Aur0ra affiliates used SpaceX's Cursor Agent, running Claude Sonnet in thinking mode, to plan and drive post-compromise activity.
- 20+ victims across 9 countries between April and July 2026; 28 recovered chat sessions cover 10 orgs from Apr 8–May 21.
- Operators fed the agent SOCKS tunnels and stolen VPN creds, then let it handle lateral movement over SMB/LDAP/WinRM/RDP, log-wiping, Defender kills, and Zig-compiled ESXi locker deploy.
- Guardrails bypassed by telling Cursor the engagement was a red-team 'simulation'; prompts explicitly excluded CIS ranges and CIS-country domains.
- Lands one day after yesterday's OpenAI-cuts-off-Cursor story and gives xAI a headache heading into Grok 4.7 — Cursor Agent is the first mainstream coding tool to show up in a ransomware IR report.
industry thehackernews.com
02
DeepSeek Nears $7.4B Round at $74B Pre-Money, Targets 2027 IPO
- Aug 28 CNBC/Bloomberg reporting: DeepSeek is finalizing a ~50 B RMB / $7.4B raise at a $74B pre-money valuation — roughly a 10× step-up on its last mark.
- Anchor investors: High-Flyer (Liang Wenfeng's quant fund) plus returning names Monolith Management, Shixiang, Tencent, JD.com, NetEase, and battery giant CATL.
- Proceeds earmarked for model research and additional compute ahead of a Shanghai Star Market IPO planned for 2027.
- Puts DeepSeek in the top three Chinese AI labs by valuation next to Moonshot and Zhipu, and above Alibaba's Qwen unit on paper — funding round #2 after its debut $7B round in June.
industry cnbc.com
03
Anthropic Ships Claude Fable 5.1 and Mythos 5.1
- Sept 1 release: same $10/$50 base pricing as Fable 5, but cache reads slashed 75% to $0.25/M tokens — Anthropic says ~25% cheaper on typical workloads, up to 45% on heavily agentic ones.
- Terminal-Bench 4.0 = 55.8% (vs 42.0% for Fable 5, 52.3% for Opus 5); Terminal-Bench-Science 0.1 = 52.6%, more than double Fable 5's 24.7%.
- 1M-token context, 128k output; Mythos 5.1 is the same model under a looser safeguard profile for vetted cyber/biosec orgs, ships alongside a new Enterprise Frontier Safeguards architecture and ~60% fewer cybersecurity false positives in Claude Code.
- HN launch thread hits 884 points / 836 comments; Simon Willison's animated-pelican-on-a-bike test rates the high-effort tier the strongest Claude he's used.
models anthropic.com
04
Google Ships Gemini 3.8 Flash and a Cyber-Focused Twin
- Sept 2 release: 1M-token context, 64K output, tuned for long-horizon coding and agentic workflows; $0.75/M input and $3.75/M output through Dec 31, both double on Jan 1.
- Terminal-Bench 2.1 climbs to 90.8% (81.6% for 3.7 Flash) and HLE-Verified hits 54.9%; beats Opus 5 on three of Google's published benchmarks, though the SWE-Bench Pro gain is barely a point.
- 3.8 Flash Cyber ships alongside behind Google Fairwind limited access for governments and trusted partners; Chrome Security team says it produces 2.6× more correct patches than comparable commercial models.
- HN launch thread splits: Google's DevRel frames higher token usage as 'verifies its work more often', developers counter that 3.8 Flash burned 120M output tokens on a benchmark suite where the median was 71M — 70% more spend at $3.75/M.
models blog.google
05
NVIDIA Buys Hugging Face for $12.9B
- Sept 3 announcement: ~$11.9B to investors plus a $1B equity retention pool for HF staff; close targeted for early 2027 pending regulatory approvals.
- HF hosts 3M models, 1M apps, and 500K datasets used by 18M developers and 200K companies — the largest open-source AI hub going to a single hardware vendor.
- Jensen pledges HF stays 'compute-agnostic': no Nvidia hardware requirement, multi-cloud and multi-accelerator support continue, founding team stays on.
- HN thread: 740+ points, 300+ comments — splits between 'Microsoft-buys-GitHub replay' fears and reproducibility upside; awkward timing after last week's OpenAI postmortem on the HF production-environment breach.
industry blogs.nvidia.com
06
OpenAI's GPT-6 Astra Debuts With an 'AGI Era' Claim
- Sept 3 launch: Brockman closes the press briefing with 'Welcome to the AGI era'; 1,050,000-token context, 128K output, knowledge cutoff Apr 30 2026.
- Astra saturates FrontierMath Tier 4 v2 at 97.6%, ARC-AGI-3 at 99.9%, and ExploitBench at 100%; 72.6% on OSWorld 2.0 at ~47% less time per task than GPT-5.6 Sol.
- Independent Artificial Analysis Intelligence Index lands at 61 — identical to Sol, behind Anthropic's Fable 5.1 — and on Humanity's Last Exam with tools Astra scores 57.2% vs Fable 5.1's 65.0%.
- API pricing $10/$50 per M tokens (2.5× Sol); Altman apologizes Sept 4 for a 'messy' rollout after Pro/Plus users were locked out at launch.
models openai.com
07
Claude Formalizes Fermat's Last Theorem in Lean in 11 Days
- Anthropic ran a swarm of parallel Claude agents that produced the first end-to-end machine-checked proof in 11 days of wall-clock time.
- Output: 13M lines of Lean 4, 30,300 theorems (29,500 used in the final proof), roughly 6B output tokens.
- First formalization attempt failed; adding Columbia's open-source Prove2Me tool mid-run unblocked completion. Lean verified with only its three standard axioms.
- Imperial's Kevin Buzzard calls it an 'extraordinary autoformalization achievement'; mathematicians had expected the Wiles proof to take years to formalize.
research anthropic.com
08
DeepMind's AlphaGenome Atlas Scores Every Single-Letter Human DNA Change
- Sept 8: precomputed AlphaGenome predictions for all ~9B possible single-nucleotide variants in the human genome, packaged as a ~1 PB dataset — roughly 30× the AlphaFold DB.
- Introduces the AlphaGenome Variant Impact (AVI) score summarising each variant across regulatory, splicing, and expression tracks.
- Live via a public website, the AlphaGenome API, and a first-party skill in Google Antigravity; free for non-commercial research.
- The largest single-shot precomputed biology asset any frontier lab has shipped this year.
research deepmind.google
09
OpenAI Claims a Navier–Stokes Proof — NYU and Anthropic Post the Euler One First
- OpenAI told reporters an unreleased model, running ~10,000 coordinated agents over 88 hours, produced a Lean-formalized proof of finite-time blowup for the forced 3D Navier–Stokes equations.
- Run stats: ~2.7M messages, ~130B output tokens; the model described as 'significantly more capable than GPT-6 Astra.' No proof or model artefact published.
- Same day, NYU's Tristan Buckmaster and Anthropic's Levent Alpöge posted a Claude-assisted, machine-checked proof of finite-time blowup for the 3D Euler equations — with the actual Lean files.
- HN and math-Twitter latched onto the contrast: one claim in a press call, one verifiable artefact on arXiv.
research unite.ai
10
NSA, CISA, and FBI Name Six Chinese AI Firms in 'Industrial-Scale' Distillation Advisory
- Sept 8 joint advisory (AA26-251A) names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI for systematically distilling Claude, GPT, Gemini, and Grok since late 2024.
- Alleges billions of tokens across millions of queries were pulled via fraudulent accounts, bulk premium subs, and proxy 'transfer stations' — likely with Chinese government knowledge.
- Specifics cited: DeepSeek R1/V3 trained on Claude, GPT, and Gemini output; Moonshot's Kimi K3 trained on Claude Fable, Kimi K2 on GPT-4o.
- Unusual mitigation ask: US labs 'quietly degrade' suspected accounts rather than block them, to preserve the counter-distillation signal.
industry cisa.gov