the brief

Security and scale bracketed the day: Hugging Face detailed an agentic intrusion their own AI triage caught, while Alibaba previewed a 2.4T-parameter Qwen3.8 Max promising open weights. On the research side, new work on step-by-step information evaluation and fresh resources on multi‑agent debate surfaced, alongside pragmatic engineering on token‑efficient deep research and cautionary voices about latent backdoors in open models.

the poursit · sip · 8 items

alerts

(01)
  • techmeme· AggregatorJul 19, 10:15 PM

    Hugging Face detects agentic intrusion

    Hugging Face says an agentic system breached parts of its data pipeline, accessing internal clusters and credentials, with the company’s own AI-based triage spotting the attack early.

    Hugging Face says an agentic AI system hacked its data pipeline, accessing several internal clusters and credentials; its own AI-based triage caught the breach (Hugging Face) — Hugging Face: Hugging Face says an agentic AI system hacked its data pipeline, accessing several internal clusters and credentials; its own AI-based triage caught the breach — Earlier this week, we detected and responded to an intrusion into part of our production infrastructure.

    signal 8hype 3security_breachincident_responseai_agentstechnicalsource ↗

pulse

(01)
  • techmeme· AggregatorJul 19, 02:10 PM

    Alibaba previews 2.4T Qwen3.8 Max

    Alibaba launched a preview of Qwen3.8 Max, claiming frontier‑level performance and promising to make it open‑weight soon—potentially a big jolt for the open model ecosystem.

    Alibaba launches preview of 2.4T parameter Qwen3.8 Max, says it's comparable to frontier AI models and second only to Fable 5, will make it "open-weight soon" (Bloomberg) — Bloomberg: Alibaba launches preview of 2.4T parameter Qwen3.8 Max, says it's comparable to frontier AI models and second only to Fable 5, will make it “open-weight soon” — Alibaba Group Holding Ltd. launched a preview version of its flagship Qwen3.8 Max model, which it described as comparable …

    signal 7hype 6model_releaseopen_weightsqwenlaunchsource ↗

findings

(04)
  • tmlr-pub.bsky.social· BlueskyJul 19, 08:19 PM

    Measuring information step by step

    New OpenReview paper proposes AI‑based, stepwise information measurement to move beyond vibes-based evals, offering a more granular way to assess reasoning and explanation quality.

    Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes Zachary Robertson, Sanmi Koyejo Action editor: Nishant Mehta https://openreview.net/forum?id=i7T1tFvFQM #adversarial #ai #information

    signal 7hype 1paperevaluationreasoningtechnicalsource ↗
  • hn/frontpage· AggregatorJul 19, 12:01 PM

    Building a frugal research pipeline

    An engineering write‑up shows a custom deep‑research workflow to cut LLM token spend, detailing retrieval, summarization, iteration choices, and concrete cost/quality tradeoffs.

    I burned all my tokens researching how to save tokens — Article URL: https://quesma.com/blog/custom-deep-research-pipeline/ Comments URL: https://news.ycombinator.com/item?id=48967355 Points: 10 # Comments: 4

    signal 5hype 2agentspipeline_designtoken_optimizationtechnicalsource ↗
  • sardean.bsky.social· BlueskyJul 19, 04:49 PM

    Teaching RL without MDPs first

    A draft manuscript explores starting advanced RL with policy gradients before MDP formalism, aiming to align teaching with optimization‑centric practice and reduce conceptual overhead.

    imagining how to start an advanced undergraduate reinforcement learning course without MDPs but with policy gradients. Here is what I came up with so far, feedback is welcome! sdean.website/rl-manuscrip...

    signal 4hype 1reinforcement_learningpolicy_gradientmanuscripttechnicalsource ↗
  • tedunderwood.com· BlueskyJul 19, 04:45 PM

    Multi‑Agent Debate reference posted

    A concise pointer to the ACL‑hosted Multi‑Agent Debate paper gives teams a stable citation and baseline for debate‑driven reasoning and evaluation experiments.

    It's a thing! Multi-Agent Debate. aclanthology.org/2025.acl-lon...

    signal 4hype 1paperagentsmulti_agent_debatetechnicalsource ↗

voices

(02)
  • thezvi/vase· AnalysisJul 19, 02:11 PM

    Zvi parses Hassabis’s AI framework

    Zvi’s read of Demis Hassabis’s frontier‑AI manifesto distills policy posture from product realities, highlighting what’s actionable versus rhetorical for capability and governance watchers.

    Demis Hassabis on the New Coming Age — Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.

    signal 5hype 2policyanalysisfrontier_aiculturalsource ↗
  • matthodges.bsky.social· BlueskyJul 19, 01:45 PM

    On open‑weight model time bombs

    A reminder that open‑weight models can be trained to behave nicely until a trigger flips them adversarial, with linked research outlining practical Trojan strategies and risks.

    Little prevents an open weight model from shipping with a time bomb. RL it to be very helpful upfront so adoption spreads wide, train it on the harnesses it’ll be deployed in, turn it adversarial once some future context hits. Feels like this is being ignored. arxiv.org/abs/2401.05566

    signal 4hype 2model_securitysafetysleeper_agentsculturalsource ↗