$ cat briefs/daily/2026-07-10.md
2026-07-10
July 10, 2026 (Fri)
5 new items · 1 late capture · 1 new page · 7 updated
+1new page
[01]
Top Stories
1. ChatGPT Work — OpenAI launches autonomous multi-hour work agent
- A standalone product (distinct from GPT-5.6 model) that accepts an outcome goal, connects to your apps and files, executes steps independently for hours, and delivers finished outputs: docs, sheets, slides, web apps
- Available immediately for Pro, Enterprise, and Edu plans; Plus/Business rollout within days; powered by GPT-5.6
- Simultaneous rename: Codex desktop app → "ChatGPT" desktop, bundling Work + Codex in one
- Why it matters: this is the first OpenAI product designed around multi-hour autonomous operation rather than turn-by-turn assistance — the interface shifts from "answer this" to "finish this." Directly competes with Anthropic Managed Agents, xAI Agent Tools API, and Google ADK 2.0. The agentic-work market just got its clearest consumer expression yet.
- → OpenAI (Bloomberg) (OpenAI)
2. Meta Muse Spark 1.1 + Meta Model API — Meta enters the commercial AI API market
- Muse Spark 1.1: 1M context, major gains in tool use / computer use / coding / MCP generalization; new SOTA on MedScribe, TaxEval, and Harvey's Legal Agent Bench — beating Claude Fable 5 at 10× lower cost and 2× speed
- Meta Model API public preview (US only): $1.25/$4.25 per M tokens input/output; $20 free credits; speaks both OpenAI and Anthropic SDK formats
- Why it matters: Meta has spent 2026 building the infrastructure for a closed-model API business (MSL, Watermelon training, "dual-track") alongside Llama open-weights. Today it opened the door. At $1.25/$4.25, Muse Spark 1.1 undercuts Sonnet 5 ($2/$10) and Grok 4.5 ($2/$6) on price — and beats Fable 5 on professional-domain benchmarks. The frontier API market now has a 5th credible entrant.
- → Muse Spark (1.0 / 1.1) (Meta AI) (TechCrunch)
3. GPT-Live-1 — OpenAI's full-duplex voice replaces ChatGPT Voice for all tiers __
- GPT-Live-1 (paid) and GPT-Live-1 mini (free): can listen and speak simultaneously — natural interruptions, no turn-taking wait
- Intelligent delegation: background hand-off to GPT-5.5 for complex reasoning or web search, result returned into the voice stream
- GPT-Live-1 → default for Go/Plus/Pro; GPT-Live-1 mini → default for Free; rolled out globally July 8–9
- Why it matters: full-duplex eliminates the biggest friction in voice AI — the stop-and-wait ping-pong. Combined with ChatGPT Work, voice becomes the natural interface for agent delegation ("finish this project while I'm in meetings"). Prior leader in voice UI was Google's Gemini Live; OpenAI is directly competing.
- → GPT-Live-1 (OpenAI) (TechCrunch)
[02]
Paper Picks
NVIDIA Nemotron-Labs-Diffusion (arXiv:2607.05722) — arXiv
- TL;DR: A single set of weights unifies autoregressive, diffusion-parallel, and self-speculation decoding. In self-speculation mode the model is both drafter and verifier, sharing a KV cache — no separate draft model needed. Result: 6.82 accepted tokens/step vs. Eagle3's 2.75; 6× more tokens per forward pass vs. Qwen3-8B; 4× throughput on SPEED-Bench (SGLang/GB200).
- Why read it: self-speculation without a separate draft model is a practical simplification for production serving stacks. Open weights (3B/8B/14B, commercial license). From NVIDIA Research on their own hardware — not a benchmark-farming paper.
- → NVIDIA
[03]
Watch
- AlphaEvolve now GA on Gemini Enterprise (July 10) — DeepMind's algorithm-discovery agent is production-deployed for logistics/genomics/HPC/semiconductor/finance customers. 5–56% error reduction in early access. Competes directly with OpenAI Codex and Claude for Science. FedRAMP/DoD excluded. → AlphaEvolve (Google Cloud)
- Claude Reflect (beta) (July 9) — Anthropic added a usage dashboard and break-reminder feature to Claude.ai/Desktop. Monthly recap, quiet hours, break nudges. Unusual product move (encouraging reduced use). Requires Memory. → Anthropic (Anthropic)
- Gemini 3.5 Pro July 17 GA — still on track per July 8 confirmation. Architecture fully rebuilt from scratch after scrapping 2.5 Pro base. 2M context window, Deep Think. Watch for benchmark comparisons against Muse Spark 1.1 and GPT-5.6 Sol in the same week.
- Meta Watermelon — MSL opening the paid API now (Muse Spark 1.1) confirms Watermelon is NOT imminent. Muse Spark 1.1 is the commercial bridge model; Watermelon (GPT-5.5 parity claim, 10× training compute) remains training.
[04]
New in Wiki
- GPT-Live-1 (new — OpenAI full-duplex voice model page created)
[05]
Updates
- Muse Spark (1.0 / 1.1): Major rewrite — added Muse Spark 1.1 spec, benchmarks (MedScribe/TaxEval/Harvey SOTA), Meta Model API pricing, comparison table
- Meta AI: Added Muse Spark 1.1 + Meta Model API to Recent Activity and Models & Products
- OpenAI: Added ChatGPT Work (July 9) and GPT-Live-1 (July 8, to Recent Activity and Models & Products
- AlphaEvolve: Added GA section with early-access metrics and Gemini Enterprise context
- Google DeepMind: Added AlphaEvolve GA to Recent Activity
- Anthropic: Added Claude Reflect (July 9) and Fable 5 extension to July 12 to Recent Activity
- index.md: Added models/gpt-live-1; stats updated to 89 pages, ~150 sources