This run covers scan window [2026-09-13T01:02:19Z, 2026-09-13T16:00:31Z], with retrieval overlap back to 2026-09-12T22:02:19Z. It adds one event, updates two existing events, and skips one duplicate release. Events are filed by official UTC date; this daily is filed by Asia/Shanghai calendar day.

Daily Executive Summary

  • LiteLLM RC adds Skills discovery and MCP permissions (ev-20260913-01, WATCH): 1.102.0-rc.1 serves Agent Skills through a well-known index and can configure Codex and Claude Code with a gateway key. Per-user MCP permissions now cover both tools/list and tools/call, alongside customer-managed KMS, routing telemetry and guardrail failure-path fixes. The direction is coherent, but this is still an RC; production upgrades should wait for stable.
  • DeepSeek V4.1-Flash's vLLM support reaches merged code (ev-20260910-01 update): DeepSelect measures 191 microseconds versus 616 for the prior path in one GB200 batch-256 / 1M-KV kernel test. Engram tables also gain node-local DP sharding, asynchronous CPU-offload prefetch and shared memory. Both live only on vLLM main and are not an end-to-end performance promise for a tagged release.
  • Qwen3.8-Flash-Next-NVFP4 gains an upstream single-box path (ev-20260826-04 update): SGLang main can serve it on one DGX Spark / GB10 and publishes a 200-question GSM8K and throughput check. The result is TP=1 and not a controlled before/after comparison, so the recommendation remains WATCH.

Models

  • DeepSeek V4.1-Flash (TRIAL): vLLM #56464 and #56512 add the sparse-indexer TopK and Engram DP-sharding/offload paths. Framework integration is becoming concrete, but it has not reached a tagged release.
  • Qwen3.8-Flash-Next (WATCH): SGLang #39126 brings NVIDIA's NVFP4 derivative to one DGX Spark. This is a reproducible workstation path, not full production validation.

Agent & AI Engineering

  • LiteLLM 1.102.0-rc.1 (WATCH): Skills discovery, MCP authorization, gateway keys, customer-managed KMS and routing observability now sit in one control plane. A proof of concept should test permission consistency across tools/list and tools/call, spend accounting and fail-closed guardrail behavior.
  • Gemini CLI's 0.61 nightly republished the same source hash at 01:29Z. The official diff contains only package versions and sandbox-image tags, so it is skipped as a duplicate.

Open Source

  • The valuable movement in this window is on framework main branches: vLLM merged two DeepSeek V4.1-Flash serving paths, while SGLang merged a single-box path for the NVFP4 Qwen Flash-Next derivative.
  • Codex had only unreleased main commits; llama.cpp had rolling builds and isolated fixes. Neither becomes a radar event.

Research

  • The arXiv watermark remains 2609.11923. No v1/v2 timestamp falls in the strict window, and the Sunday digest had not appeared by the cutoff.

Developer Tools

  • Codex stable remains 0.154.0, Claude Code remains 2.1.270, and Gemini CLI stable remains 0.59.0. LiteLLM's RC can configure Codex/Claude Code gateways, but it does not change those client release lines.

Infrastructure

  • DeepSelect's 191-microsecond result covers only a TopK kernel on GB200. Engram prefetch improves reported TTFT by at most 0.9%. Do not extrapolate either into an end-to-end throughput claim during capacity planning.
  • The Qwen NVFP4 path checks 200 GSM8K questions on GB10: 97.5%/97.0% with/without MTP and 92.87/89.53 tok/s. The samples use different concurrency settings, so they do not isolate MTP's benefit.

Business & Policy

  • No policy event in this window materially changes model access, API cost, chip supply or the open-source landscape. An Anthropic sitemap refresh appeared after the scan cutoff and is deferred for content-level verification next run.

Trend Signals

No evidence-supported new trend emerged in this cycle. Trend #1 gains one serving-ecosystem evidence group: the DeepSeek and Qwen open-weight lines now have merged vLLM and SGLang code, but third-party Terminal-Bench / SWE reproduction on the weights remains absent, so it stays strengthening / Medium. Trend #3 gets an adjacent signal on skills/orchestration standardization from the LiteLLM RC; it does not count as formal evidence before stable. Every other trend keeps its status and confidence.

Tech Radar

  • New LiteLLM 1.102.0-rc.1 (ai-engineering / WATCH): Agent Skills discovery and MCP authorization reach one gateway; wait for stable before a production upgrade review.
  • DeepSeek V4.1-Flash (foundation-model / TRIAL): vLLM merged DeepSelect and Engram production paths, but they are not in a tagged release yet.
  • Qwen3.8-Flash-Next (foundation-model / WATCH): SGLang merged a single-DGX-Spark NVFP4 path; validation remains narrow.

Current radar: ADOPT 6 / TRIAL 59 / WATCH 79, 144 tracked entries total. See the site's Insights page for the complete current radar.

Worth Trying

  • Install the LiteLLM RC in isolation, verify that each user's MCP tool list matches actual call permissions, and reconcile spend before versus after the upgrade.
  • On GB200, reproduce the DeepSelect kernel result separately, then measure end-to-end TTFT and throughput before using it for capacity planning.
  • On a DGX Spark, reproduce SGLang #39126's Qwen NVFP4 path and expand the check to full GSM8K plus coding-agent workloads.

Watch Items

  • Whether LiteLLM 1.102 stable preserves these Agent Skills, MCP-permission and KMS behaviors with upgrade guidance.
  • Whether the next vLLM tag includes #56464/#56512 and the next SGLang tag includes #39126.
  • Trend #1's final gate: a third-party Terminal-Bench / SWE reproduction on public GLM-5.3 or DeepSeek V4.1-Flash weights.
  • Continue arXiv above 2609.11923; verify Anthropic's post-cutoff sitemap movement against actual publication timestamps.

Sources