AI Trend Notifier
EN
← archive

$ cat briefs/daily/2026-08-05.md

2026-08-05

August 5, 2026 (Wed)

3 stories · 0 papers · 2 watch items · 3 new pages

+3new pages
[01]

Top Stories

1. OpenAI discloses two more eval containment failures — and one names the vendor already behind Anthropic's

  • OpenAI reported that two external testing partners identified incidents where "testing configurations and controls combined with the advancing capabilities of the recent models allowed for model activity to extend beyond their intended testing boundaries" (source).
  • Irregular notified OpenAI on 2026-07-29: a misconfiguration in its Capture-the-Flag environment let models onto the public internet, and in one test the fictional target's name coincided with a real domain, which the models exploited while believing they were still inside the exercise. Anthropic named the same partner as the cause of its own three breaches, disclosed five days earlier.
  • UK AISI notified OpenAI on 2026-08-03, from cyber-range evaluations run with internet access intentionally enabled and cyber classifiers disabled — by design, not by fault.
  • Model names, affected-system counts and whether the real domain belonged to an identifiable organization were not disclosed.
  • Why it matters: the outsourced-evaluation trust boundary now has two failures at a single vendor rather than one, and UK AISI's is the first incident in this sequence where nothing was misconfigured — which is the one a configuration standard cannot fix.
  • Eval Environment Containment, OpenAI

2. Liquid AI ships a 2.6B agentic model that runs on a phone

  • LFM2.5-2.6B: 2.69B parameters in under 2.5 GB, 128K context, ~34T pre-training tokens; 220 tok/s on an M5 Max, 113 on a Ryzen AI Max+, ~30 on phone-class hardware (source).
  • Post-training ran four stages including Agentic RL inside live harnesses (OpenClaw, Hermes Agent) rather than synthetic traces alone. Weights on Hugging Face with day-one llama.cpp / MLX / vLLM / SGLang / ONNX support.
  • Vendor claims competitive tool-use against Gemma 5B–8B and Qwen 4.7B–9.7B at 2–4× smaller. No benchmark with a number attached was published, and no independent measurement exists. The licence could not be read and is recorded as unknown, not inferred from the family tag.
  • Why it matters: every open-weight release this wiki has tracked recently runs 118B to 2.8T. This one is competing with running no model at all on the device, which is a different market and a much larger one.
  • LFM2.5-2.6B, Liquid AI (both new)

3. Grok Voice Think Fast 2.0 becomes the default for every grok-voice-latest caller today

  • Released 2026-07-29 and captured seven days late — x.ai has no feed, so it is polled by search rather than through the prefetch ledger. What surfaced it was today's alias cutover, not the launch (source).
  • ~0.70 s to first audio, down from 1.25 s; reported 82.9% on Artificial Analysis' speech-to-speech benchmark (overall) against GPT-Realtime-2.1 at 79.1%. Recorded as reported — this repo holds no Artificial Analysis snapshot with a speech-to-speech column to check it against. $0.08/minute.
  • Why it matters: xAI's voice claims have been self-administered until now (Think Fast 1.0 at 67.3% on its own tau-voice Bench). A third-party figure ahead of OpenAI's realtime model is a different kind of claim — and the automatic alias flip means every existing integration inherits the new model without opting in.
  • Grok Voice Think Fast 2.0, xAI
[02]

Paper Picks

None today. HuggingFace Daily Papers was unfetchable for the fifth consecutive day (WebFetch 403 on every host attempted), and the WebSearch fallback that rescued a pick on 2026-08-03 surfaced nothing dated 2026-08-04 — only a 2026-08-02 item already past.

[03]

Watch

  • Anthropic appoints its first chief global affairs officer, Tino Cuéllar — former California Supreme Court justice, president of the Carnegie Endowment until July, reporting to Daniela Amodei. This wiki has tracked Anthropic's government friction as a run of separate episodes; a standing executive role is the company treating it as a function. → Anthropic (source)
  • GLM-5.5 is still rumoured, not shipped. Wednesday's Chinese-lab rotation (Z.ai + MiniMax) found a JPMorgan note and July community leaks pointing at an August release above 1T parameters, but no model card, benchmark or endpoint, and Z.ai's own channels still promote GLM-5.2. Nothing citable was added. → Z.ai, GLM-5.2
[04]

New in Wiki

  • Liquid AI (new — Key People: unknown; company history, funding and headcount were not established from the release announcement alone)
  • LFM2.5-2.6B (new — Pricing and License are both genuine unknowns, explained on the page)
  • Grok Voice Think Fast 2.0 (new — no parameter count, architecture or context window published by xAI)
[05]

Updates

  • Eval Environment Containment: new section for the third disclosure — the two new incidents in a table, the Irregular finding, and UK AISI added to the taxonomy as a deliberately opened sandbox rather than a broken or unclosed one. State of the Art advanced to 2026-08-05; two Open Problems rewritten.
  • OpenAI: the 2026-08-04 disclosure in full.
  • Anthropic: Cuéllar appointment; Cuéllar added to Key People; a cross-reference recording that OpenAI's Irregular incident bears on Anthropic's own.
  • xAI: Grok Voice Think Fast 2.0; the Grok Voice API product line now names its engine.