the brief

Flagship models and agent stacks moved fast: OpenAI’s GPT‑5.6 hit GA and became Microsoft 365 Copilot’s preferred model, while Meta’s Muse Spark 1.1 finally shipped an API focused on agentic tool use. Devtools kept pace with Claude Code and Next.js updates, Cloudflare pushed for ML‑DSA in post‑quantum signatures, and new studies surfaced on RAG strategies and jailbreak robustness.

the poursit · sip · 14 items

pulse

(09)
  • simonw/blog· AnalysisJul 9, 07:46 PM

    OpenAI ships GPT-5.6 family

    GPT‑5.6 is GA with Luna, Terra, and Sol tiers priced per million tokens (Luna $1/$6, Terra $2.5/$15, Sol $5/$30), aiming at stronger tool use and long‑context work.

    The new GPT-5.6 family: Luna, Terra, Sol — <p>OpenAI's latest flagship model <a href="https://openai.com/index/gpt-5-6/">hit general availability this morning</a>, and comes in three sizes: Luna, Terra, and Sol (from smallest to largest).</p> <p>The new models are priced per 1M input/output tokens as Luna $1/$6, Terra $2.50/$15, Sol $5/$30. For comparison, the Claude Opus series are $5/$25 and the Claude Fable 5 is $10/$50, but price-per-million tokens doesn't tell us much now that the number...

    signal 9hype 1model_releaseopenaipricinglaunchsource ↗
  • simonw/blog· AnalysisJul 9, 04:24 PM

    Meta unveils Muse Spark 1.1 API

    Meta’s Spark line gets its first API with reported gains in agentic tool calling and computer use, plus extensive evals in the released datasheet and blog.

    Introducing Muse Spark 1.1 — <p><strong><a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Introducing Muse Spark 1.1</a></strong></p> Following <a href="https://simonwillison.net/2026/Apr/8/muse-spark/">Muse Spark in April</a>, here's Muse Spark 1.1 - the first Spark model to offer an API. Meta claim significant improvements in agentic tool calling and computer use.</p> <p>There are a lot more details are in the <a href="https://ai.meta.com/static-resource/muse-spark-1...

    signal 9hype 1model_releaseapiagentic_tool_uselaunchsource ↗
  • simonw/blog· AnalysisJul 10, 01:05 AM

    ChatGPT Work cloud versus desktop clarified

    OpenAI docs note desktop Work runs locally with permitted apps/files and is not synced with cloud Work at launch—important for data boundaries and workflow planning.

    Quoting OpenAI — <blockquote cite="https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex"><p>[...] Work on web and mobile runs in the cloud. Work in the desktop app can also use local files and desktop apps with your permission. At launch, cloud Work conversations do not appear in desktop Work; desktop Work threads and local files remain on that computer.</p></blockquote> <p class="cite">&mdash; <a href="https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex">OpenAI...

    signal 6hype 1openaichatgpt_workdesktop_applaunchsource ↗
  • simonw/blog· AnalysisJul 9, 04:12 PM

    llm-meta-ai adds Muse Spark support

    Simon Willison’s llm-meta-ai 0.1 lets developers run prompts against Meta’s muse‑spark‑1.1 from the llm CLI, easing quick experiments and eval loops.

    llm-meta-ai 0.1 — <p><strong>Release:</strong> <a href="https://github.com/simonw/llm-meta-ai/releases/tag/0.1">llm-meta-ai 0.1</a></p> <p>Let's LLM run prompts against the new <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">muse-spark-1.1</a> model.</p> <p>Tags: <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/meta">meta</a></p>

    signal 7hype 1plugin_releasecli_toolmeta_modellaunchsource ↗
  • anthropics/claude-code· First-partyJul 10, 01:45 AM

    Claude Code v2.1.206 ships updates

    Release adds /cd directory suggestions, a /doctor to trim CLAUDE.md, safer auto‑push for /commit-push-pr, gateway login improvements, and smoother worktree handling.

    v2.1.206 — What's changed Added directory path suggestions to /cd, matching /add-dir behavior Added a /doctor check that proposes trimming checked-in CLAUDE.md files by cutting content Claude could derive from the codebase /commit-push-pr now auto-allows git push to the repo's configured push remote (remote.pushDefault, or the sole remote when only one is configured) in addition to origin Gateway: /login now supports Anthropic-operated public gateway endpoints EnterWorktree now asks for confi...

    signal 9hype 1release_notesclaude_codedev_ergonomicslaunchsource ↗
  • vercel/next.js· First-partyJul 10, 12:07 AM

    Next.js canary adds service worker support

    v16.3.0‑canary.82 compiles service workers via Turbopack, outputs them under /_next/static, and bumps React to the 20260708 snapshot alongside rendering benchmarks and agentic checks.

    v16.3.0-canary.82 — Misc Changes [turbopack] Compile service workers registered from pages router pages: #95583 [turbopack] Output service workers to /_next/static/: #95554 Reduce the size of OperationVc from 8 bytes to 4: #95614 Upgrade React from df4bd1b4-20260708 to 5123b063-20260708: #95642 Add attribute rendering benchmark: #95621 docs(opengraph-image): load assets at module scope to keep route static under Cache Components: #95246 Convert agent-041's blocking-data check to the agentic L...

    signal 8hype 1release_notesnext_jsturbopacktechnicalsource ↗
  • openai/blog· First-partyJul 9, 01:00 PM

    Microsoft 365 Copilot moves to GPT-5.6

    Microsoft 365 Copilot now prefers GPT‑5.6 across Word, Excel, PowerPoint, Chat, and Cowork, promising faster, higher‑quality outputs and stronger reasoning/tool use.

    GPT-5.6 is now the preferred model in Microsoft 365 Copilot — Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.

    signal 7hype 2model_updatemicrosoft_365copilotlaunchsource ↗
  • techmeme· AggregatorJul 9, 11:03 PM

    OpenAI AGI head shifts to adviser

    Fidji Simo will not return to a full‑time OpenAI role due to health, remaining as an adviser—a notable leadership change as the company readies an IPO.

    Fidji Simo says she has decided to leave her full-time role at OpenAI and transition to being a part-time adviser after her medical condition worsened (Wall Street Journal) — Wall Street Journal: Fidji Simo says she has decided to leave her full-time role at OpenAI and transition to being a part-time adviser after her medical condition worsened — The head of artificial general intelligence won't return from extended medical leave, a major leadership change as the AI giant prepares to go public

    signal 7hype 2openaileadership_changeorg_newsculturalsource ↗
  • techmeme· AggregatorJul 9, 09:55 PM

    Mercor buys Deeptune for agent RL

    Mercor acquired Deeptune, which builds reinforcement‑learning environments for AI agents, signaling consolidation and demand for richer evals and training grounds for autonomous systems.

    Mercor acquires Deeptune, which builds reinforcement learning environments for AI agents, three months after CEO Brendan Foody backed Deeptune's $43M Series A (Lily Mae Lazarus/Fortune) — Lily Mae Lazarus / Fortune: Mercor acquires Deeptune, which builds reinforcement learning environments for AI agents, three months after CEO Brendan Foody backed Deeptune's $43M Series A — Brendan Foody started Mercor when he was 19 years old. Now he's 23, worth billions on paper, and running one of the fast...

    signal 5hype 2acquisitionagentsreinforcement_learninglaunchsource ↗

findings

(04)
  • cloudflare/blog· First-partyJul 9, 02:00 PM

    Cloudflare urges ML-DSA now for PQ signatures

    Cloudflare reviews nine emerging post‑quantum signature candidates and argues teams should deploy ML‑DSA today rather than wait, balancing security with operational readiness.

    Why we cannot wait for better post-quantum signature algorithms — NIST is advancing nine new post-quantum signature algorithms as potential candidates for future standardization. We take a closer look at all of them, and argue that while they are in the works and show great potential, we should use ML-DSA for now — the best one currently available.

    signal 7hype 1post_quantumcryptographynisttechnicalsource ↗
  • hn/frontpage· AggregatorJul 10, 01:28 AM

    Cloudflare's vulnerability harness blueprint

    A practical guide to building a reusable vulnerability harness improves reproducible exploit testing and regression checks across complex stacks—useful for maturing app‑sec pipelines.

    Build your own vulnerability harness — Article URL: https://blog.cloudflare.com/build-your-own-vulnerability-harness/ Comments URL: https://news.ycombinator.com/item?id=48854681 Points: 12 # Comments: 5

    signal 7hype 1securitytutorialtestingtechnicalsource ↗
  • tmlr-pub.bsky.social· BlueskyJul 9, 08:20 PM

    Iterative RAG can beat ideal evidence

    A TMLR study on scientific multi‑hop QA finds iterative retrieval/search sometimes outperforms setups with preselected ‘ideal’ evidence, challenging assumptions about RAG evaluation design.

    When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answ... Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee et al. Action editor: Zhangyang Wang https://openreview.net/forum?id=pa5TnBdyDP #retrieval #iterative #knowledge

    signal 6hype 1paperragretrievaltechnicalsource ↗
  • tmlr-pub.bsky.social· BlueskyJul 9, 12:19 PM

    Character noise amplifies LLM jailbreaks

    TMLR paper shows simple character‑level perturbations can significantly boost jailbreak success by exploiting tokenization quirks, underscoring fragile guardrails and the need for robust defenses.

    Random Character-Level Perturbations Amplify LLM Jailbreak Attacks Shuyi Yu, Zhe Cao, Kohei Tsuji et al. Action editor: Hanwang Zhang https://openreview.net/forum?id=BXsOIppKEI #tokenization #jailbreak #adversarially

    signal 5hype 1paperjailbreakadversarial_attackstechnicalsource ↗

voices

(01)
  • pragmatic/engineer· AnalysisJul 9, 05:20 PM

    Cursor shares eye-opening coding stats

    Cursor reports power users generate 10× more code than median, most AI spend is on input tokens, and nearly half of AI code changes merge without manual review.

    The Pulse: Interesting AI coding stats from Cursor — Power users generate 10x as many lines of code vs the median, most of the AI spend is coming from input tokens not output ones, and almost half of AI changes are accepted without manual review by devs (!!)

    signal 7hype 2ai_codingcursorproductivity_metricsculturalsource ↗