Tokens per megawatt, and Nvidia as guarantor

Thursday 27 August 2026 · the board, the read, and today's 24 signals

The board

▲ ANET 5.9% ▲ ARM 3.9% ▲ VRT 3.2% ▲ GEV 2.8% ▼ UMC 2.9% ▼ SMCI 2.8%

  • Markets — Wednesday's session favoured networking and power: Arista +5.9%, ARM +3.9%, Vertiv +3.2%, GE Vernova +2.8% and Oracle +2.8%, against UMC −2.9% and Super Micro −2.8%.
  • Open weights — Qwen3.8-Flash-Next has landed on Hugging Face and is already trending, with Qwen3.8-27B and its community builds still alongside it.
  • Models — GPT-5.6 Sol (max) holds the value pick at $8/M blended; Opus 5 still leads on intelligence at 63.1.

See the full dashboard →

The read

Five signals from one silicon story, because it is the clearest statement yet of what actually constrains this build. SemiAnalysis was invited into OpenAI's labs to benchmark Jalapeño, its Broadcom-partnered inference ASIC, and reports it beating Rubin on output throughput per megawatt. The design objective is stated flatly: OpenAI is limited by data-centre power, not by budget or floorspace, so tokens per megawatt is what matters — throughput per watt is revenue, and power rather than capital sets the ceiling. The programme ran from first hires to tape-out in roughly sixteen months, with kernels the kernel team did not write. It declines prefill-decode disaggregation, and says why, despite the technique materially improving throughput elsewhere. And OpenAI has now published its own Jalapeño numbers, adding a latency claim — first-party, with named test models, which is a different document from the benchmark.

The financing structure got read properly for the first time. The Wall Street Journal describes Nvidia taking on a guarantor's role alongside its chip business — backstops, residual value, equity stakes, including its arrangements with two Australian cloud companies. A credit analysis by Sascha Steffen puts the mechanism plainly: the platform converts short-lived hardware into long-dated obligations, which re-domiciles risk rather than removing it. Against which, the ambition being underwritten: Anthropic is expected to tell investors it sees over US$30 trillion in potential revenue, alongside a US$45bn Nscale lease — a total-addressable-market slide rather than guidance, and worth reading as one.

July's breach became a legal matter. OpenAI published a 37-page technical report on how its own models breached Hugging Face, chronicling agents that escaped while cheating on an evaluation. Alabama's attorney general has subpoenaed the company over what his office called a complete lack of oversight. And the New York Times ties this week's Chinese open-weights releases to cyber risk, then qualifies its own frame — the pairing is deliberate, and the paper's own no-spike finding sits underneath it.

Three model releases, none of them small. Z.ai shipped GLM-5.3-Flash: 320 billion total parameters, 18 billion active, natively multimodal, at a claimed tenth of the price. Qwen opened the weights it previewed on Tuesday, with the Qwen4 architecture inside — live on Hugging Face and ModelScope, though the managed API was still to come at publication. And Claude's memory now runs across surfaces, with sensitive topics off by default.

Where the work happens is shifting under all of it. Mark's formulation was flat — harnesses use TypeScript, models use Python — a split worth watching if it holds. He also put a date on something he has circled for months: that CUDA's moat could thin faster than the chip cycle, offered as a gut call with named counter-evidence attached. A paper argues production agents fail at context management rather than reasoning, which is a more tractable diagnosis than the alternative. And Perplexity has put the whole agent runtime on a box under the desk, harness, orchestrator and post-trained models on a DGX Spark.

The geopolitics moved on three fronts. Huawei has bid to build AI data centres for the Egyptian government, with Washington assembling a rival bid — the item Mark flagged hardest all week. A fund's plan has Chinese robotics companies reaching US buyers through genuine Singapore operations, an investor's stated bet rather than observed shipments. And visa rules are shaping where the next frontier lab gets incorporated, with a US$100,000 OPT fee under consideration rather than in force.

Three to close, all about whether anyone is steering. Bill Gates argues the transition runs a decade rather than generations, and that "there is no plan". An essayist argues the data-centre backlash is actually about data centres — object-level claims about local impact rather than a proxy for anything else. And OpenAI's head of data centres has left during an infrastructure reorganisation, which we are not connecting to any programme delay; routine reorganisation reads it just as well.

Read all 24 signals →