This run covers scan window [2026-09-13T01:02:19Z, 2026-09-13T16:00:31Z], with retrieval overlap back to 2026-09-12T22:02:19Z. It adds one event, updates two existing events, and skips one duplicate release. Events are filed by official UTC date; this daily is filed by Asia/Shanghai calendar day.
Daily Executive Summary
- LiteLLM RC adds Skills discovery and MCP permissions (ev-20260913-01, WATCH): 1.102.0-rc.1 serves Agent Skills through a well-known index and can configure Codex and Claude Code with a gateway key. Per-user MCP permissions now cover both
tools/listandtools/call, alongside customer-managed KMS, routing telemetry and guardrail failure-path fixes. The direction is coherent, but this is still an RC; production upgrades should wait for stable. - DeepSeek V4.1-Flash's vLLM support reaches merged code (ev-20260910-01 update): DeepSelect measures 191 microseconds versus 616 for the prior path in one GB200 batch-256 / 1M-KV kernel test. Engram tables also gain node-local DP sharding, asynchronous CPU-offload prefetch and shared memory. Both live only on vLLM main and are not an end-to-end performance promise for a tagged release.
- Qwen3.8-Flash-Next-NVFP4 gains an upstream single-box path (ev-20260826-04 update): SGLang main can serve it on one DGX Spark / GB10 and publishes a 200-question GSM8K and throughput check. The result is TP=1 and not a controlled before/after comparison, so the recommendation remains WATCH.
Models
- DeepSeek V4.1-Flash (TRIAL): vLLM #56464 and #56512 add the sparse-indexer TopK and Engram DP-sharding/offload paths. Framework integration is becoming concrete, but it has not reached a tagged release.
- Qwen3.8-Flash-Next (WATCH): SGLang #39126 brings NVIDIA's NVFP4 derivative to one DGX Spark. This is a reproducible workstation path, not full production validation.
Agent & AI Engineering
- LiteLLM 1.102.0-rc.1 (WATCH): Skills discovery, MCP authorization, gateway keys, customer-managed KMS and routing observability now sit in one control plane. A proof of concept should test permission consistency across
tools/listandtools/call, spend accounting and fail-closed guardrail behavior. - Gemini CLI's 0.61 nightly republished the same source hash at 01:29Z. The official diff contains only package versions and sandbox-image tags, so it is skipped as a duplicate.
Open Source
- The valuable movement in this window is on framework main branches: vLLM merged two DeepSeek V4.1-Flash serving paths, while SGLang merged a single-box path for the NVFP4 Qwen Flash-Next derivative.
- Codex had only unreleased main commits; llama.cpp had rolling builds and isolated fixes. Neither becomes a radar event.
Research
- The arXiv watermark remains 2609.11923. No v1/v2 timestamp falls in the strict window, and the Sunday digest had not appeared by the cutoff.
Developer Tools
- Codex stable remains 0.154.0, Claude Code remains 2.1.270, and Gemini CLI stable remains 0.59.0. LiteLLM's RC can configure Codex/Claude Code gateways, but it does not change those client release lines.
Infrastructure
- DeepSelect's 191-microsecond result covers only a TopK kernel on GB200. Engram prefetch improves reported TTFT by at most 0.9%. Do not extrapolate either into an end-to-end throughput claim during capacity planning.
- The Qwen NVFP4 path checks 200 GSM8K questions on GB10: 97.5%/97.0% with/without MTP and 92.87/89.53 tok/s. The samples use different concurrency settings, so they do not isolate MTP's benefit.
Business & Policy
- No policy event in this window materially changes model access, API cost, chip supply or the open-source landscape. An Anthropic sitemap refresh appeared after the scan cutoff and is deferred for content-level verification next run.
Trend Signals
No evidence-supported new trend emerged in this cycle. Trend #1 gains one serving-ecosystem evidence group: the DeepSeek and Qwen open-weight lines now have merged vLLM and SGLang code, but third-party Terminal-Bench / SWE reproduction on the weights remains absent, so it stays strengthening / Medium. Trend #3 gets an adjacent signal on skills/orchestration standardization from the LiteLLM RC; it does not count as formal evidence before stable. Every other trend keeps its status and confidence.
Tech Radar
- New LiteLLM 1.102.0-rc.1 (ai-engineering / WATCH): Agent Skills discovery and MCP authorization reach one gateway; wait for stable before a production upgrade review.
- DeepSeek V4.1-Flash (foundation-model / TRIAL): vLLM merged DeepSelect and Engram production paths, but they are not in a tagged release yet.
- Qwen3.8-Flash-Next (foundation-model / WATCH): SGLang merged a single-DGX-Spark NVFP4 path; validation remains narrow.
Current radar: ADOPT 6 / TRIAL 59 / WATCH 79, 144 tracked entries total. See the site's Insights page for the complete current radar.
Worth Trying
- Install the LiteLLM RC in isolation, verify that each user's MCP tool list matches actual call permissions, and reconcile spend before versus after the upgrade.
- On GB200, reproduce the DeepSelect kernel result separately, then measure end-to-end TTFT and throughput before using it for capacity planning.
- On a DGX Spark, reproduce SGLang #39126's Qwen NVFP4 path and expand the check to full GSM8K plus coding-agent workloads.
Watch Items
- Whether LiteLLM 1.102 stable preserves these Agent Skills, MCP-permission and KMS behaviors with upgrade guidance.
- Whether the next vLLM tag includes #56464/#56512 and the next SGLang tag includes #39126.
- Trend #1's final gate: a third-party Terminal-Bench / SWE reproduction on public GLM-5.3 or DeepSeek V4.1-Flash weights.
- Continue arXiv above 2609.11923; verify Anthropic's post-cutoff sitemap movement against actual publication timestamps.
Sources
- https://github.com/BerriAI/litellm/releases/tag/v1.102.0-rc.1 and https://pypi.org/project/litellm/1.102.0rc1/
- https://github.com/vllm-project/vllm/pull/56464 and https://github.com/vllm-project/vllm/pull/56512
- https://github.com/sgl-project/sglang/pull/39126
- https://github.com/google-gemini/gemini-cli/releases/tag/v0.61.0-nightly.20260913.g9c1b0a610 and https://github.com/google-gemini/gemini-cli/compare/v0.61.0-nightly.20260912.g9c1b0a610...v0.61.0-nightly.20260913.g9c1b0a610
- https://openai.com/news/rss.xml, https://github.blog/changelog/feed/, https://blog.google/technology/ai/rss/, and https://arxiv.org/list/cs.AI/recent