$ cat briefs/daily/2026-08-03.md
2026-08-03
August 3, 2026 (Mon)
Generated by ingest + brief agents · Sources: Tier 1 (HuggingFace Daily, Anthropic, alignment.anthropic.com, Meta AI, Mistral, xAI, X feeds, MiniMax + DeepSeek rotation) + prefetch candidates
_2 top stories · 1 paper pick · 6 new pages_
Top Stories
1. Two Western open-weight releases had been sitting unrecorded for three weeks
An Interconnects recap on 2026-08-02 named three models as the current open Pareto frontier. One of them, Kimi K3, this wiki has held since July. The other two it had never heard of.
Inkling — Thinking Machines Lab, 2026-07-15. A 975B total / 41B active multimodal MoE, 1M context, Apache 2.0, full weights on Hugging Face in BF16 and NVFP4, with day-zero support in SGLang, vLLM, llama.cpp and Unsloth. Pretrained on 45 trillion tokens; the MoE design largely follows DeepSeek-V3. The lab explicitly declines to call it the strongest model — its stated role is a customizable base for fine-tuning, sold alongside the lab's own Tinker platform.
Laguna S 2.1 — Poolside, 2026-07-21. 118B total / 8B active, built for agentic coding, up to 1M context, under OpenMDW-1.1. Poolside's own figures: 70.2% on Terminal-Bench 2.1 with thinking, 78.5% on SWE-Bench Multilingual, and 40.4% on DeepSWE v1.1 against DeepSeek-V4-Pro-Max's 9.0% at roughly one-sixth the active parameters.
The sharpest detail is a coincidence nobody drew. Thinking Machines Lab is a founding partner of the Open Secure AI Alliance, whose July 27 launch this wiki recorded alongside the standing objection that none of its members ships a frontier-class open model. Inkling's weights landed twelve days earlier and neither the announcement nor the coverage mentioned them — and on the one third-party measure held here it does not close the gap anyway: 41 on the Artificial Analysis Intelligence Index against Claude Opus 5 (max) at 61.
Why it matters: "restrict open weights" and "restrict Chinese models" had converged into nearly the same policy because the largest open releases were Chinese. Two US-addressed releases inside a week — one Apache 2.0, one pitched in coverage as "the West's answer to DeepSeek and Qwen" — put a domestic constituency on the side a distillation-focused enforcement regime would burden.
→ Inkling (new), Laguna S 2.1 (new), Thinking Machines Lab (new), Poolside (new), Open-Weights Policy Fight → Inkling snapshot · Laguna snapshot · Interconnects #23 · AA leaderboard 2026-08-02
2. MiniMax shipped a video model and promised the weights — the weights did not come
MiniMax H3 launched 2026-07-31 (aliased Hailuo 3.0): a single transformer reading text, images, video and audio in one context and returning video with native stereo audio — 1440p, 4–15 seconds at whole-second granularity, 24 fps, where the previous generation offered a choice between 6 and 10. MiniMax states it set aside the Hailuo-02 architecture; the new H3-Omni Transformer splits understanding from generation during training because multimodal context tripled sequence-length variance, lifting end-to-end throughput by nearly 30%. Live in the API as MiniMax-H3 and in the Hailuo app at $0.13 per generated second at 2K.
It is marketed as open-weight with weights promised "in the coming days". A day after launch none had shipped.
Why it matters: this is the third model this wiki has recorded whose open-weight status is an announcement rather than a release — Qwen 3.8 Max (Preview) made the same promise on July 19 and has still not kept it. When a policy fight is being argued over what open weights do to the threat surface, a marketing claim and a downloadable checkpoint have stopped being the same thing, and the spec table now says unknown for the license rather than a license name.
→ MiniMax H3 (new), MiniMax, Open-Weights Policy Fight → Source · MiniMax · MarkTechPost
Paper Picks
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents — arXiv:2607.28227, Tongyi MAI / Alibaba / Qwen AI Lab, 2026-07-29. HF Daily Papers, 2026-08-02.
- TL;DR: one model covering mobile, computer, browser and DeepSearch, trained against sandbox environments plus a large-scale real-device mobile runtime. Reported 82.1% on MobileWorld — +14.6 points over Opus 4.8, +12.0 over GPT-5.6 Sol — with 92.2% on MobileWorld-Real and 97.5% on AndroidDaily. On computer use, 79.5% on OSWorld-Verified, second overall.
- Why read it: the margin is the wrong shape for an incremental result. A 14.6-point lead over a frontier generalist on mobile GUI use suggests the surface is separable from computer use, and that whoever has real devices to train against wins it. The report also describes a harness for proactive service initiation — detecting a flight cancellation and acting, rather than waiting to be invoked, which is a materially different trust surface than any agent recorded here.
- Caveat: vendor-run figures, no published harness configuration, and the author list, parameter count and license were not obtainable from this environment. Open-weight availability is a report claim, not a verified fact.
- → Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents (new), Agents (LLM Agents), Eval Harness Configuration
Watch
- MiniMax H3 weights. Promised "in the coming days" on July 31. If they land, MiniMax H3's
Licenserow moves offunknown; if they do not, it joins Qwen 3.8 Max (Preview) as a second unfulfilled open-weight announcement inside a month. (source) - xAI voice routing switches August 5.
grok-voice-latestbegins routing togrok-voice-think-fast-2.0on 2026-08-05 per xAI's release notes; Grok 4.6 remains announced-not-shipped on Grok 4.6. (xAI release notes) - The Alignment Science blog has a pre-wiki backlog. Checked against
sources/this run: no new article since Agentic Misalignment (July 13), which is held — the daily check is clean. But the article list carries at least five posts predating this wiki that it has never held, including Measuring and improving coding audit realism (March 10) and Model Spec Midtraining (May). Backfill candidate for the next lint, not an ingest miss. (alignment.anthropic.com)
New in Wiki
For user review.
- Thinking Machines Lab (new — entity page created; Key People is
unknown, the lab's own site was unreadable from here) - Poolside (new — entity page created; Key People
unknownfor the same reason) - Inkling (new —
- Laguna S 2.1 (new —
- MiniMax H3 (new)
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents (new —
Updates
- MiniMax — MiniMax H3 added to Models & Products and Recent Activity
- Alibaba / Qwen AI Lab — Qwen-UI-Agent added to AI Products & Research and Recent Activity
- Open-Weights Policy Fight — new section on the two Western open-weight releases; the "alliance has no frontier model" open problem now names Thinking Machines and the measurement that does not close it
- Agents (LLM Agents) — Qwen-UI-Agent added to Key Papers / Events