Scan window: 2026-09-16 16:00:47Z to 2026-09-17 23:58:56Z; retrieval overlap starts 13:00:47Z. Sixteen new events, three existing release-line entities updated, and 15 duplicate, prerelease, out-of-window or low-value candidates filtered.

Daily Executive Summary

  • The main shift is operational. Anthropic publishes two-stage monitoring data for roughly 30,000 concurrent internal agents. GitHub begins measuring adoption of skills, agents and MCP. OpenAI turns misalignment reporting from one-off long reports into a recurring tiered process.
  • Scientific engineering also produces deployable artifacts: Anthropic releases 36 biomolecular optimization kits, while ScienceIDE packages scientific repositories as trainable, verifiable environments. ASLEval finds that final-output checks miss almost half of session privacy exposure, and Pydantic AI patches web-fetch and telemetry boundaries.

Models

  • Astra for Law (WATCH): a dedicated legal index covers more than 230 million URLs and outperforms ordinary web search on a private benchmark. Access is limited and the API has not shipped.
  • Anthropic Life Sciences Verification Program (WATCH): verified organizations can request broader biology access in exchange for cross-session monitoring and 30-day retention of flagged activity.

Agent & AI Engineering

  • Anthropic 30,000-agent operations data (TRIAL): online and offline monitoring cover the full internal platform; August produced more than one billion online decisions and roughly 50 weekly human escalations.
  • Google CC (WATCH): a group agent gets its own Google Account, shared and per-person memory, and explicit permission boundaries.
  • ASLEval (WATCH): expected-outlet checks miss 46.9% of exposure found across all visible exits. Privacy evaluation needs a complete session boundary.
  • Reward-hacking probe (WATCH): cheap internal-representation vectors approach expensive LLM monitors and sometimes anticipate later hacking actions.
  • DualViewEval (WATCH): process signals compress selected agent benchmarks by 24x-40x; periodic full-suite audits remain necessary.

Open Source

  • Anthropic biomolecular kits (TRIAL): 36 Apache-2.0 kits provide exact, fast and big modes. Anthropic reports roughly 4x average speedup but will not maintain the repository.
  • UN System Data Commons (TRIAL): an authoritative statistical knowledge graph is live and available to research agents through MCP.
  • ScienceIDE (TRIAL): scientific environments, training workflows and 4B/9B/72B model artifacts are public; the repository lacks a detected license.

Research

  • HOPE (TRIAL): second-order expert pruning preserves more agentic coding quality at 50% pruning.
  • CASHEWS (WATCH): preprocessing lifts malicious-package scan coverage to 98.8%-100%, but no implementation is public.
  • rMuscle (WATCH): reusing VLA state across repetitive robot runs yields a reported 1.29x-1.42x speedup in narrow workloads.

Developer Tools

  • GitHub Copilot telemetry (ADOPT): enterprise reports now cover skills, custom agents, MCP, plugins and 28-day engagement across Copilot surfaces. MCP counts are connection attempts, not tool calls.
  • Pydantic AI 2.44.0 (ADOPT): fixes two policy bypasses, quadratic web-fetch processing and a telemetry-content leak; upgrade v1 deployments through the backport too.
  • Codex 0.155 (existing entity updated): MCP Touch ID, local task lifecycle actions, daemon recovery and WSL/credential isolation reach stable.
  • Claude Code 2.1.274-2.1.275 (existing entity updated): improves MCP startup, nested/forked subagent output, plugin integrity checks and sandbox/resume behavior.

Infrastructure

  • HOPE and rMuscle target MoE weight memory and repetitive VLA execution respectively. Both need workload-specific latency, quality and safety tests.
  • Z.ai adds a vendor-reported claim that GLM-5.3-Flash moved from adaptation to production in under two weeks across more than 100,000 Chinese accelerators; merged into the existing event.

Business & Policy

  • OpenAI misalignment reporting framework (ADOPT): a recurring three-track process publishes six cases across compaction, credentials, file upload and cross-agent communication. Trend #4 moves from emerging to strengthening.
  • Anthropic's Life Sciences Verification Program makes higher-risk access organization- and project-scoped, extending capability gating from model releases into workflows.

Trend Signals

  • Trend #3 (established / High): the 30,000-agent fleet, two-stage monitoring and GitHub feature telemetry add scale and adoption evidence.
  • Trend #4 (strengthening / Medium): OpenAI turns incident reporting into an ongoing process; independent-review outcomes and a cross-vendor format are still missing.
  • Trend #7 (emerging / Medium): Codex MCP verification, Google CC identity boundaries and ASLEval session boundaries continue separating capability from consequence control.
  • No new trend has enough independent evidence in this cycle.

Tech Radar

  • New OpenAI misalignment framework (Safety / ADOPT): recurring tiered disclosure replaces ad hoc reports
  • New Pydantic AI 2.44.0 (Developer Tools / ADOPT): production web-fetch and telemetry boundaries need an upgrade
  • New GitHub Copilot telemetry (Developer Tools / ADOPT): skills, agents and MCP adoption become measurable
  • New Anthropic 30,000-agent metrics (Agent Observability / TRIAL): real fleet monitoring and escalation economics are public
  • New Claude biomolecular kits (Scientific AI / TRIAL): reproducible optimization artifacts need scientific validation
  • New UN Data Commons MCP (Research Agent / TRIAL): authoritative statistics can enter agent workflows directly
  • New ScienceIDE (Research Agent / TRIAL): one environment contract spans training and verification; resolve license risk first
  • New HOPE (Infrastructure / TRIAL): MoE pruning should preserve expert cooperation
  • New Astra for Law (Models / WATCH): promising retrieval gain, but private evaluation and limited access
  • New ASLEval (Agent Security / WATCH): a clean final answer does not prove session privacy
  • New Google CC (Agent / WATCH): group-agent identity, memory and permission boundaries merit tracking
  • Current radar: ADOPT 9 / TRIAL 79 / WATCH 93. The site insights page holds the full current radar.

Worth Trying

  • Enumerate every requester-visible and external agent exit, then rerun a privacy test using the ASLEval model.
  • Use GitHub's new 28-day fields to check whether one skill or MCP server has sustained use.
  • Package one well-tested scientific repository as a ScienceIDE-style environment and compare verifier cost with model gain.

Sources