A million downloads in three days, and what the memory wall costs

Tuesday 25 August 2026 · the board, the read, and today's 28 signals

The board

▼ MDB 6.5% ▼ MU 5.8% ▼ SMCI 5.6% ▼ AMKR 4.7% ▼ GFS 4.4% ▼ NET 4.4%

  • Markets — Monday's session was red across the slice we track: MongoDB −6.5%, Micron −5.8%, Super Micro −5.6%, Amkor −4.7%, GlobalFoundries −4.4%, Cloudflare −4.4% and Datadog −4.2%. Nothing in it finished higher.
  • Open weights — Qwen3.8-27B is at 2,645,226 downloads in thirty days, with the community's uncensored builds now trending alongside the official ones.
  • Models — the value pick changed hands for the first time in weeks: GPT-5.6 Sol (max) takes it at $8/M blended, while Opus 5 keeps top intelligence at 63.1.

See the full dashboard →

The read

The open-weights number is the one to sit with. Qwen3.8-27B's rolling thirty-day counter reached 2,358,347 on Monday morning, up from 1.37 million on Friday — about a million in three days, and Mark's reaction was that this is what a watershed looks like. It is a download counter, not an adoption figure, and worth reading as pulls rather than users. Alongside it, the Ox Alpha mystery narrowed on two fronts: fingerprinting work by an outside developer points at GLM-5.3 on six of nine probes, and the headline 80% turns out to have been a DeepSWE subset, with the full run reported at 63%. The correction is more instructive than the claim. And on the paid side, cited figures put Anthropic's costliest tier at 11.4% of customer spend against 6% of token volume — a publisher we could not identify, so treat it accordingly.

Two Australian items belong together. Assistant Minister Andrew Charlton put Australian AI spending at $5–8bn a year and described the value leaving the country as a "distant sucking sound", most of it offshore. And the ABC assembled the deflation paradox with local numbers: per-token prices have collapsed — Epoch has GPT-4-Turbo-class capability going from US$15 per million tokens to 17.5 US cents — while the bills companies actually pay keep rising. Both halves are true at once, which is exactly why the measurement matters.

The money kept moving. Alibaba priced a HK$80bn placement of new shares, about US$10.2bn, with all net proceeds earmarked for its full-stack AI capabilities — equity rather than debt, which is a different bet on the cost of capital. And Hugging Face is reported to be fielding interest at US$13bn or more, with a bank engaged and no bidder named.

Silicon had a busy day, and the through-line was memory rather than compute. Nvidia has put Groq 3 LPX racks into full production for decode workloads, commercialising its US$20bn acquisition — and The Register is worth reading on what the 3,400 tokens-per-second figure actually measures: Gemma 4 31B dense at FP8, small enough to sit in about 64 of the 256 LPUs. Nvidia also set a demanding bar for bringing CUDA to RISC-V. High Bandwidth Flash offers capacity relief that software may not take up — simulation-only so far. Mark spent Monday evening working through the architecture and landed on the memory wall as the real constraint. And Xiaomi claims a three-chip stack for serving large models locally, on vendor specifications for an engineering prototype.

Governance moved from containment to something more like law. Steve Yegge argues agent governance should look like fences rather than sandboxes — mechanisms that turn you away if you are not supposed to be there, rather than walls that try to hold everything in. His second argument is the one with the organisational edge: agents will write down the rules nobody ever wrote down. Against that, two new attack surfaces: a paper names "agentic flooding" of government services as a live risk, what happens when agents act for the public against institutions that cannot scale; and an essay argues inference engines are under-examined as a target, where a model emits tokens chosen to exploit the parser reading them. On the building side, Microsoft shipped a skill for coding agents to optimise other agents, and a paper claims stability gains over GRPO from single-rollout asynchronous RL — which matters more as training targets shift from chat to long-horizon agentic work.

Three on expertise, which is becoming this publication's most persistent thread. Studies suggest AI coding assistance can block the expertise it requires — the tools need judgement to use well and remove the friction through which judgement forms. A Goldman Sachs partner warned on the bank's own podcast of cognitive atrophy in banking's apprenticeship model, which is a notable place for that argument to surface. And Chinese survey optimism sits alongside AI-attributed layoffs — fewer than 10% of respondents worried, in a country already reporting the job losses. Coexistence, not contradiction.

Three more. Zvi Mowshowitz argues American data-centre opposition has become close to unconditional — a verdict on the industry rather than an argument about siting, on polling showing 75% local opposition despite the economics. Taiwanese prosecutors have indicted nine people over allegedly forged documents concealing AI server exports to China, said to include an Nvidia senior manager. And Pew finds AI-authorship signals in about a third of pages published since ChatGPT, with two denominators that need keeping apart.

Two to close, both about dependency. A near-three-hour error window across Claude models is a small reminder of what it means to build on a runtime somebody else operates. And the most portable idea of the day: when the cost of searching collapses but the incentive to search never existed, what actually happens?

Read all 28 signals →