Events

Events Timeline

181 structured events, filed by publication date (not discovery time). Every event is source-verified and deduplicated, with six-dimension scores and an adoption recommendation.

Category
Org
Calendar overview Darker = more events that day; click a day to clear filters and jump to its group
September 2026
MTWTFSS 18192021222324252627282930
August 2026
MTWTFSS 1234567891011121516232430

2026-09-17

8 events
2026-09-17 Developer Tools ADOPT

GitHub adds adoption telemetry for Copilot agents, skills and MCP

GitHubMicrosoft GitHub-Copilotusage-metrics +3 Impact 2 Eng 4 Adoption 5
2026-09-17 Open Source TRIAL

UN Data Commons exposes official statistics to agents through MCP

United NationsGoogle MCPData-Commons +3 Impact 3 Eng 4 Adoption 3
2026-09-17 Agent WATCH

Google gives its family agent a separate identity and permission boundary

Google multi-principal-agentidentity +3 Impact 3 Eng 4 Adoption 1
2026-09-17 Open Source TRIAL

Anthropic releases 36 Claude-built biomolecular optimization kits

Anthropic biomolecular-modelingscientific-software +3 Impact 4 Eng 4 Adoption 2
2026-09-17 Developer Tools ADOPT

Pydantic AI closes web-fetch policy and telemetry leaks

Pydantic Pydantic-AISSRF +3 Impact 3 Eng 5 Adoption 4
2026-09-17 Foundation Models WATCH

OpenAI launches Astra for Law with a dedicated legal index

OpenAI GPT-6-Astralegal +3 Impact 3 Eng 4 Adoption 3
2026-09-17 business-policy WATCH

Anthropic opens verified access for higher-risk biology work

Anthropic life-sciencesverified-access +2 Impact 3 Eng 3 Adoption 2
2026-09-17 business-policy TRIAL

Anthropic publishes monitoring metrics for a 30,000-agent fleet

Anthropic agent-fleetobservability +3 Impact 4 Eng 5 Adoption 5

2026-09-16

9 events
2026-09-16 Research TRIAL

ScienceIDE turns scientific repositories into training environments for agents

AItonomy FoundationPhAI Labs ScienceIDEscientific-agent +4 Impact 4 Eng 4 Adoption 2
2026-09-16 Infrastructure WATCH

rMuscle reuses internal VLA state across repetitive robot tasks

Research collaboration rMuscleVLA +3 Impact 3 Eng 3 Adoption 1
2026-09-16 Research WATCH

Study detects reward hacking from internal model representations

Independent researchers reward-hackingmonitoring +3 Impact 4 Eng 4 Adoption 1
2026-09-16 business-policy ADOPT

OpenAI formalizes recurring disclosure of model misalignment

OpenAI misalignmentincident-disclosure +3 Impact 4 Eng 4 Adoption 4
2026-09-16 Infrastructure TRIAL

HOPE prunes MoE experts while preserving agentic coding performance

Amazon AWS HOPEmixture-of-experts +3 Impact 3 Eng 4 Adoption 2
2026-09-16 Research WATCH

DualViewEval compresses costly agent benchmarks with process signals

Tencent HunyuanTsinghua University DualViewEvalagent-evaluation +3 Impact 3 Eng 4 Adoption 1
2026-09-16 Research WATCH

ASLEval finds local checks miss agent privacy exposure

Minjiang University ASLEvalagent-security +3 Impact 3 Eng 4 Adoption 1
2026-09-16 Research WATCH

CASHEWS expands LLM package scanning to large bundled files

Research collaboration CASHEWSsupply-chain-security +3 Impact 3 Eng 4 Adoption 1
2026-09-16 Developer Tools TRIAL

GitHub expands AI Scan beyond repositories with CodeQL default setup

GitHubMicrosoft AI-ScanCodeQL +2 Impact 3 Eng 4 Adoption 4

2026-09-15

7 events
2026-09-15 Agent Security TRIAL

Social-harness study finds cross-owner agents fail even when honest

University of Washington multi-agentA2A +3 Impact 4 Eng 4 Adoption 1
2026-09-15 agent-framework TRIAL

ScienceBuddy releases a self-improving scientific-agent workspace

PhAI LabsFudan University research-agentself-improvement +6 Impact 4 Eng 4 Adoption 1
2026-09-15 Infrastructure TRIAL

JustFit serves a 200K-token model context on a 24 GiB MacBook

independent research local-inferenceMLX +3 Impact 4 Eng 4 Adoption 1
2026-09-15 AI Engineering TRIAL

Study finds deep agent hierarchies preserve context but lose findings

independent research multi-agentdelegation +3 Impact 4 Eng 5 Adoption 2
2026-09-15 Research TRIAL

SWE-bench audit finds top coding-agent ranks statistically unresolved

Imperial College LondonHong Kong Polytechnic University coding-agentSWE-bench +5 Impact 4 Eng 5 Adoption 3
2026-09-15 Infrastructure WATCH

FlashVector tunes production serving across kernels, servers and features

UnityStanford University model-servingoptimization-agent +3 Impact 4 Eng 5 Adoption 3
2026-09-15 Open Source TRIAL

OPEN-1B makes every training step independently auditable

Gensyn open-modeltraining +3 Impact 4 Eng 4 Adoption 2

2026-09-14

13 events
2026-09-14 Research TRIAL

Study shows injected plans can evade chain-of-thought monitors

academic (see paper) agent-securityprompt-injection +1 Impact 4 Eng 4 Adoption 1
2026-09-14 Research WATCH

Gavel reads skill-routing signals from a frozen model

academic (see paper) agent-skillsrouting +1 Impact 3 Eng 4 Adoption 1
2026-09-14 Research TRIAL

VLoc Bench finds agents struggle to locate vulnerable code

academic (see paper) agent-evaluationsecurity +1 Impact 3 Eng 4 Adoption 1
2026-09-14 Research WATCH

AlgoEvo adapts algorithm search using execution feedback

academic (see paper) agentic-searchalgorithm-discovery +1 Impact 3 Eng 3 Adoption 1
2026-09-14 AI Engineering TRIAL

Study finds shell access beats typed tools on enterprise agent tasks

MicrosoftCarnegie Mellon University agenttool-use +4 Impact 4 Eng 5 Adoption 2
2026-09-14 Research TRIAL

Harness study finds no average vendor-native coding advantage

evolutionID coding-agentagent-harness +4 Impact 3 Eng 5 Adoption 2
2026-09-14 Agent Security TRIAL

Runtime gate blocks policy-breaking actions from clean agent skills

Pheo agent-skillssupply-chain +4 Impact 4 Eng 5 Adoption 2
2026-09-14 Infrastructure TRIAL

AMD releases execution-verified data for training ROCm kernel agents

AMD rocmhip +4 Impact 4 Eng 4 Adoption 1
2026-09-14 Research WATCH

SAS trains sparse attention selectors directly from language loss

Tencent HunyuanHKUST sparse-attentionlong-context +3 Impact 4 Eng 3 Adoption 1
2026-09-14 Infrastructure WATCH

SQD splits decode around subquadratic attention stages

NVIDIAHarvard University inferencedisaggregation +3 Impact 4 Eng 4 Adoption 1
2026-09-14 Research TRIAL

ParaRecover evaluates recovery across parallel tool-call failures

Dalian University of Technology agent-evaluationparallel-tool-use +3 Impact 3 Eng 4 Adoption 1
2026-09-14 Research TRIAL

Expert audit finds physics benchmarks understate frontier models

Yale UniversityUniversity of Cambridge benchmark-auditscientific-reasoning +4 Impact 4 Eng 5 Adoption 2
2026-09-14 Research WATCH

Constant-state cache makes block-diffusion context length independent

MilaConcordia University diffusion-language-modelmamba +4 Impact 4 Eng 3 Adoption 1

2026-09-13

1 events

2026-09-12

3 events
2026-09-12 Developer Tools WATCH

Gemini CLI nightly adds a gate against build-file prompt injection

Google gemini-cliagent-security +3 Impact 3 Eng 4 Adoption 1
2026-09-12 Business & Policy WATCH

Dario Amodei proposes embedded third-party evaluators to pace the AI frontier

Anthropic pacinggovernance +4 Impact 2 Eng 2 Adoption 3
2026-09-12 Research WATCH

Real-SWE benchmarks coding agents on licensed private enterprise codebases

Specific Labs benchmarkswe +4 Impact 3 Eng 4 Adoption 2

2026-09-11

12 events
2026-09-11 Developer Tools TRIAL

GitHub upgrades Copilot review with shell tools and an agent ensemble

GitHubMicrosoft copilotcode-review +4 Impact 3 Eng 4 Adoption 4
2026-09-11 Research WATCH

NCP-ArchPreview trains an 8.9B model to predict concepts, not just tokens

academic architecturelatent-space +2 Impact 4 Eng 2 Adoption 1
2026-09-11 Infrastructure WATCH

OpenAI's Habitat storage platform: 70M req/s lessons and a 2-engineer Rust rewrite

OpenAI infrastructurestorage +3 Impact 4 Eng 5 Adoption 2
2026-09-11 Research TRIAL

Osprey: one target-agnostic drafter backbone serves many speculative-decoding targets

SambaNovaacademic speculative-decodingserving +1 Impact 3 Eng 4 Adoption 2
2026-09-11 Research TRIAL

KVShareArena benchmarks KV-cache reuse beyond exact prefixes

UCLacademic kv-cachebenchmark +2 Impact 3 Eng 4 Adoption 2
2026-09-11 Research TRIAL

Subagents beat context-loaded skills as task horizons grow

Microsoft Research Cambridge agent-architectureskills +2 Impact 3 Eng 4 Adoption 2
2026-09-11 Research WATCH

Gander: an open full-duplex omni agent released with models, code and data

Zhejiang Universityacademic omni-agentvoice +3 Impact 4 Eng 4 Adoption 3
2026-09-11 Research WATCH

25 Fields Medalists declare the goals of AI companies and mathematics severely misaligned

academicmathandai.org research-policyevaluation +3 Impact 2 Eng 2 Adoption 4
2026-09-11 Research WATCH

A one-direction weight edit strips refusal from the shipped GLM-5.3-Flash weights

academic open-weights-safetyrefusal-direction +3 Impact 4 Eng 3 Adoption 3
2026-09-11 Research WATCH

Copycat dynamics explain the colluding wiki agents' collective behavior

academic multi-agentemergent-behavior +3 Impact 3 Eng 3 Adoption 2
2026-09-11 AI Engineering TRIAL

IB2: a protocol for scoring enterprise AI systems by serving route, not model ID

academic evaluationbenchmark-methodology +3 Impact 3 Eng 4 Adoption 2
2026-09-11 Agent Security WATCH

Researchers attribute a May RubyGems package flood to an OpenAI agent swarm

Independent ResearchersRubyGems agent-securitysupply-chain +4 Impact 3 Eng 3 Adoption 2

2026-09-10

18 events
2026-09-10 Developer Tools TRIAL

OpenAI opens the Agents API beta: the Codex harness as a managed service

OpenAI agent-platformmanaged-api +4 Impact 4 Eng 5 Adoption 3
2026-09-10 Foundation Models TRIAL

DeepSeek V4.1-Flash: new CED/CSA2 architecture with 1M context, open weights

DeepSeek foundation-modelopen-weights +5 Impact 5 Eng 5 Adoption 4
2026-09-10 Developer Tools TRIAL

Cursor Projects: a coordinator agent for months-long bodies of work

AnysphereCursor multi-agentcoordinator +2 Impact 4 Eng 4 Adoption 3
2026-09-10 Agent Security WATCH

Anthropic threat report: attackers prompt-injected an eval sandbox to hunt pre-release Claude

Anthropic threat-intelligenceprompt-injection +3 Impact 3 Eng 4 Adoption 3
2026-09-10 AI Engineering TRIAL

OpenAI API wave: GPT-Live 1 GA, prompt-cache diagnostics GA, key lifecycles

OpenAI apivoice +3 Impact 3 Eng 4 Adoption 3
2026-09-10 Research TRIAL

RedKnot-MLA: prefix-cache-style reuse finally works for MLA models

Peking UniversityXiaohongshu kv-cacheprefix-caching +3 Impact 4 Eng 4 Adoption 2
2026-09-10 Research TRIAL

Speculative decoding brought into RL rollout at 122B scale

Houmo AIacademic speculative-decodingrl-training +2 Impact 4 Eng 3 Adoption 2
2026-09-10 Research TRIAL

A supply-chain audit finds security defects in 16% of public coding-agent configs

academic agent-securitymcp +3 Impact 3 Eng 5 Adoption 3
2026-09-10 Research TRIAL

CapScope: capability-scoped harnesses cut prompt-injection success to near zero

UCLacademic prompt-injectionagent-security +2 Impact 4 Eng 5 Adoption 2
2026-09-10 Research TRIAL

SWE-Bench Pro audit: reward hacking inflated the leaderboard

Shanghai AI LaboratoryFudan University benchmarkreward-hacking +2 Impact 3 Eng 4 Adoption 3
2026-09-10 Research TRIAL

One instruction cuts benchmark exploitation from 45-82% to 4-11%

NVIDIAacademic benchmarkevaluation +2 Impact 3 Eng 4 Adoption 3
2026-09-10 Research WATCH

FrogNano: a 4B coding agent trained purely on online-synthesized tasks

ServiceNow Researchacademic coding-agentrl-training +2 Impact 3 Eng 3 Adoption 2
2026-09-10 Research WATCH

MoEMB: multimodal embeddings scaled through the expert axis

MetaZhejiang University embeddingsmultimodal +2 Impact 3 Eng 3 Adoption 2
2026-09-10 Research WATCH

Hybrid attention mechanisms do not fully fix attention sinks at 1M tokens

academic hybrid-attentionattention-sinks +3 Impact 3 Eng 3 Adoption 2
2026-09-10 AI Engineering TRIAL

SLOWeave sizes prefill chunks to decode deadlines with a performance guarantee

academic llm-servingchunked-prefill +3 Impact 3 Eng 4 Adoption 2
2026-09-10 Research TRIAL

ExecCritic separates test-writing from repair and trains both with RL

academic coding-agentstest-generation +3 Impact 3 Eng 4 Adoption 2
2026-09-10 Research TRIAL

Code2Skill distills 19,769 repositories into verified agent skills

academic agent-skillsskill-synthesis +2 Impact 3 Eng 4 Adoption 2
2026-09-10 Research WATCH

Content-based addressing replaces positional stretching for long context

academic long-contextrope +3 Impact 4 Eng 2 Adoption 2

2026-09-09

2 events
2026-09-09 Agent Security WATCH

Anthropic discloses a fourth incident and signs a METR investigation agreement

AnthropicMETR agent-securitysafety-incident +3 Impact 4 Eng 5 Adoption 3
2026-09-09 Developer Tools TRIAL

GitHub Copilot ships enterprise-managed permissions for agent operations

GitHubMicrosoft copilotagent-governance +3 Impact 3 Eng 4 Adoption 4

2026-09-08

20 events
2026-09-08 Foundation Models TRIAL

OpenAI ships GPT-Image-2.5: Sunburst for fidelity, Flare for half the latency

OpenAI image-generationapi +1 Impact 3 Eng 4 Adoption 4
2026-09-08 Foundation Models WATCH

Nex-N2.5: Shanghai Innovation Institute open-sources agent-recipe models

Shanghai Innovation Institutenex-agi open-weightsagent-model +2 Impact 2 Eng 2 Adoption 3
2026-09-08 Research WATCH

OpenAI's internal model produces an AI-generated solution to the Navier–Stokes Millennium Problem

OpenAI millennium-prizelean +4 Impact 5 Eng 4 Adoption 3
2026-09-08 AI Engineering TRIAL

Prefix caching quietly breaks reproducibility, and quantization multiplies the damage

academic prefix-cachingquantization +4 Impact 3 Eng 4 Adoption 2
2026-09-08 AI Engineering TRIAL

KVMem virtualizes million-token agent workspaces on a consumer GPU

academic agent-memorykv-cache +4 Impact 4 Eng 5 Adoption 3
2026-09-08 Research TRIAL

Causal experiments split hybrid models: attention recalls, recurrence controls

academic hybrid-attentionlinear-attention +4 Impact 3 Eng 3 Adoption 2
2026-09-08 AI Engineering TRIAL

One integer halves activated MoE experts with near-zero quality loss

academic moeexpert-skipping +3 Impact 3 Eng 5 Adoption 3
2026-09-08 Research TRIAL

tau-tau-Bench: the strongest agent-building agent passes 24% of real engagements

academic agent-evaluationbenchmark +3 Impact 3 Eng 4 Adoption 2
2026-09-08 Research TRIAL

EVOHARNESSBENCH measures what agents lose as their harness evolves

academic agent-evaluationharness +3 Impact 3 Eng 4 Adoption 2
2026-09-08 Research TRIAL

Agent memory changes meaning across model upgrades unless normalized to a schema

academic agent-memorymemory-portability +4 Impact 2 Eng 4 Adoption 2
2026-09-08 AI Engineering TRIAL

A draft-model gate predicts coding-agent failures before execution

academic uncertaintyspeculative-decoding +3 Impact 3 Eng 4 Adoption 2
2026-09-08 Agent Security WATCH

Prompt injection should be evaluated as an attacker's test-time search

academic prompt-injectionagent-security +3 Impact 3 Eng 3 Adoption 2
2026-09-08 Research TRIAL

HackProbe detects reward hacking in self-evolving models without touching weights

academic reward-hackingself-evolving +4 Impact 3 Eng 3 Adoption 2
2026-09-08 Research TRIAL

Harbor Adapters standardize 80+ agentic benchmarks into one runnable harness

academic agent-evaluationbenchmark +3 Impact 3 Eng 4 Adoption 3
2026-09-08 Business & Policy WATCH

Mistral raises €3B Series D to push sovereign open-weight AI to the frontier

Mistral AI fundingopen-weights +2 Impact 2 Eng 2 Adoption 3
2026-09-08 Agent Security TRIAL

Meta details the Muse agent security stack, with a public bug bounty

Meta agent-securityprompt-injection +4 Impact 3 Eng 5 Adoption 3
2026-09-08 AI Engineering WATCH

Cohere details its megakernel serving engine for North Mini Code

Cohere model-servingmegakernel +1 Impact 3 Eng 3 Adoption 2
2026-09-08 Developer Tools WATCH

Windsurf removes Cascade; Devin Local becomes its only agent

WindsurfCognition coding-agentdevin +2 Impact 2 Eng 2 Adoption 2
2026-09-08 computer-use WATCH

CUA-Universe synthesizes hybrid GUI+CLI environments for computer-use agents

Shanghai Jiao Tong Universityacademic computer-usegui-agent +2 Impact 3 Eng 3 Adoption 2
2026-09-08 Research WATCH

Alpöge and Buckmaster resolve the forced Euler problem with an Anthropic internal model

AnthropicNYU eulermillennium-prize +4 Impact 4 Eng 3 Adoption 2

2026-09-07

1 events

2026-09-06

2 events
2026-09-06 Research WATCH

OpenAI discloses the numbers behind agent-driven research acceleration

OpenAI rsiautomated-research +5 Impact 4 Eng 4 Adoption 5
2026-09-06 Business & Policy WATCH

OpenAI's chief scientist warns against building systems we cannot monitor

OpenAI alignmentmonitorability +4 Impact 3 Eng 2 Adoption 3

2026-09-05

7 events
2026-09-05 Open Source TRIAL

Jina-OCR-v1: a production document-parsing model with speculative decoding

Jina AI ocrdocument-parsing +5 Impact 4 Eng 5 Adoption 3
2026-09-05 Research TRIAL

Training-free lossy speculative decoding lifts SGLang throughput up to 56%

academic speculative-decodingsglang +3 Impact 3 Eng 4 Adoption 2
2026-09-05 Research TRIAL

Study quantizes a hybrid 27B model to NVFP4 W4A4, gating layers included

academicMinima AI quantizationnvfp4 +4 Impact 4 Eng 4 Adoption 3
2026-09-05 Research TRIAL

Random KV eviction matches learned eviction policies at higher throughput

academicUIUC kv-cacheeviction +5 Impact 4 Eng 4 Adoption 2
2026-09-05 Research WATCH

Speculative Macro Commit pre-executes agent action chains to cut latency

academicUSC agent-latencyspeculative-execution +3 Impact 3 Eng 4 Adoption 2
2026-09-05 Research TRIAL

SWE-Gate: a third of functionally passing agent fixes fail review constraints

academic coding-agentevaluation +3 Impact 3 Eng 4 Adoption 2
2026-09-05 Agent Security WATCH

Seven frontier models ran real businesses for 72 hours and all failed expensively

Bottleneck Labs agent-evaluationreal-world-agents +5 Impact 2 Eng 4 Adoption 2

2026-09-04

5 events
2026-09-04 Research WATCH

Anthropic completes the first end-to-end verified formalization of Fermat's Last Theorem

AnthropicColumbia University formal-verificationlean +4 Impact 5 Eng 3 Adoption 3
2026-09-04 Agent Security WATCH

Researchers document 18,000 colluding OpenAI agent posts on a public wiki

OpenAINightingale Collective agent-securitymulti-agent +4 Impact 4 Eng 5 Adoption 4
2026-09-04 Research TRIAL

CRISP speeds up long-context prefill 5.3x with input-adaptive sparse attention

Adobeacademic long-contextsparse-attention +3 Impact 3 Eng 4 Adoption 2
2026-09-04 Research TRIAL

Repo-To-Skill distills GitHub repositories into verified agent skills

academic skillsknowledge-distillation +3 Impact 3 Eng 4 Adoption 2
2026-09-04 Research TRIAL

EarlyEval cuts agent evaluation cost by stopping doomed runs early

academicSingapore Management University agent-evaluationcost-reduction +3 Impact 3 Eng 4 Adoption 2

2026-09-03

4 events
2026-09-03 Foundation Models ADOPT

OpenAI launches GPT-6 Astra, its first Critical-capability model

OpenAI gpt-6astra +7 Impact 5 Eng 4 Adoption 5
2026-09-03 computer-use WATCH

UI-Venus-2: an open-source foundation GUI agent for mobile, web, and desktop

academicVenus Team computer-usegui-agent +4 Impact 3 Eng 3 Adoption 2
2026-09-03 Research TRIAL

Verifier audit: RLVR reward errors concentrate in whitespace and punctuation

academic rlvrverifier +3 Impact 3 Eng 4 Adoption 2
2026-09-03 Developer Tools WATCH

xAI ships Grok Bot for Enterprise with free Cursor-Enterprise bundling

xAI grok-botenterprise +2 Impact 2 Eng 2 Adoption 3

2026-09-02

8 events
2026-09-02 Research WATCH

WHALE co-optimizes model weights and the agent harness

academic harness-optimizationfine-tuning +3 Impact 3 Eng 4 Adoption 2
2026-09-02 Research WATCH

CacheBridge transfers KV caches across different model architectures

academic kv-cachecross-model +4 Impact 3 Eng 4 Adoption 2
2026-09-02 Research TRIAL

Tool-calling study finds action-class calibration, not execution, is the bottleneck

academic tool-callingevaluation +3 Impact 3 Eng 4 Adoption 2
2026-09-02 Research WATCH

The Irreversibility Budget prices agent effects a fleet cannot undo

academic agent-governancerisk-control +3 Impact 3 Eng 4 Adoption 2
2026-09-02 Open Source WATCH

AMD open-sources Instella-MoE, a 16B MoE trained entirely on Instinct GPUs

AMD open-weightsmoe +3 Impact 3 Eng 3 Adoption 2
2026-09-02 Foundation Models ADOPT

Google ships Gemini 3.8 Flash and a gated Flash Cyber security model

GoogleGoogle DeepMind geminiflash +5 Impact 4 Eng 4 Adoption 4
2026-09-02 Developer Tools TRIAL

Cursor adds self-hosted machines so cloud agents run on your own infra

AnysphereCursor self-hostedcloud-agents +3 Impact 3 Eng 4 Adoption 3
2026-09-02 Foundation Models WATCH

Meta iterates Muse Spark toward long-horizon agentic coding

MetaMeta Superintelligence Labs muse-sparkagentic-coding +2 Impact 3 Eng 3 Adoption 3

2026-09-01

6 events
2026-09-01 Foundation Models ADOPT

Anthropic launches Claude Fable 5.1 with 75% cheaper cache reads

Anthropic claudefable +9 Impact 5 Eng 4 Adoption 5
2026-09-01 Business & Policy WATCH

OpenAI declares Astra its first Critical-capability model with new safeguards

OpenAI astrapreparedness-framework +3 Impact 3 Eng 3 Adoption 2
2026-09-01 Research TRIAL

Qwen team documents the Qwen3.8-Flash-Next architecture and its 1/9 training-FLOP economics

AlibabaQwen architecturemoe +5 Impact 4 Eng 4 Adoption 3
2026-09-01 Research TRIAL

Tail-Replay brings unconstrained prefix caching to hybrid attention LLMs

academic prefix-cachingkv-cache +4 Impact 3 Eng 4 Adoption 2
2026-09-01 Research TRIAL

BAITBENCH measures how often frontier agents take the planted shortcut

academic reward-hackingagent-evaluation +3 Impact 3 Eng 3 Adoption 2
2026-09-01 Research WATCH

Study: emergent misalignment is predictable from training-data distance

academic alignmentmisalignment +3 Impact 3 Eng 3 Adoption 2

2026-08-31

3 events
2026-08-31 Foundation Models WATCH

DeepSeek open-sources V4-Flash-Vision-Exp, its first V4-family multimodal model

DeepSeek open-weightsmultimodal +4 Impact 3 Eng 3 Adoption 3
2026-08-31 Agent Security WATCH

Anthropic details its July escape incidents and reward-hacking research

Anthropic agent-securitymisalignment +6 Impact 4 Eng 4 Adoption 3
2026-08-31 Research TRIAL

RealSWE shows coding-agent benchmarks barely resemble real user requests

academic evaluationcoding-agent +3 Impact 3 Eng 4 Adoption 2

2026-08-29

3 events
2026-08-29 Research WATCH

SARA splits action induction from execution authorization in tool-using agents

academic agent-securityprompt-injection +4 Impact 3 Eng 4 Adoption 1
2026-08-29 Research WATCH

HarnessLens evolves agent harnesses under a verification budget

academic agent-harnesscoding-agent +3 Impact 3 Eng 4 Adoption 1
2026-08-29 Research WATCH

TwinKV repairs KV cache eviction using key redundancy, not attention

academic kv-cachelong-context +4 Impact 3 Eng 4 Adoption 1

2026-08-28

3 events
2026-08-28 Research WATCH

Anthropic: automated researchers closed 85% of the deception safety gap

Anthropic alignmentsafety +4 Impact 4 Eng 3 Adoption 2
2026-08-28 Foundation Models WATCH

Tencent open-weights Hy4 preview: 770B MoE under Apache 2.0

Tencent open-weightsmoe +3 Impact 3 Eng 3 Adoption 2
2026-08-28 Business & Policy WATCH

OpenAI will cut off Cursor's bundled model access on November 12

OpenAISpaceX cursormodel-access +3 Impact 2 Eng 3 Adoption 4

2026-08-27

3 events
2026-08-27 computer-use WATCH

Anthropic previews the Model Hardware Standard for agent-controlled devices

AnthropicHHMI Janelia Research Campus computer-usephysical-devices +4 Impact 4 Eng 3 Adoption 2
2026-08-27 Foundation Models WATCH

Google ships Gemini Omni 1.1 Flash with keyframe control and 4K output

GoogleGoogle DeepMind video-generationmultimodal +2 Impact 2 Eng 3 Adoption 2
2026-08-27 Business & Policy WATCH

Report: Nvidia agrees to buy Hugging Face for $12.9 billion

NvidiaHugging Face acquisitionhugging-face +3 Impact 3 Eng 2 Adoption 5

2026-08-26

5 events
2026-08-26 Open Source ADOPT

vLLM 0.28.0 lands a big-model serving push for Kimi-K3 and DeepSeek V4

vLLM Project vllminference +4 Impact 3 Eng 4 Adoption 4
2026-08-26 Research TRIAL

AutoSaddler learns agent-harness fixes from failure traces

Microsoftacademic agent-harnesscoding-agent +4 Impact 4 Eng 4 Adoption 2
2026-08-26 Research WATCH

AgentWeave filters the tool list before reasoning to cut agent cost

academic tool-routingfunction-calling +5 Impact 2 Eng 3 Adoption 1
2026-08-26 Research WATCH

Study: full-trace process judges score relevance, not causal contribution

academic agent-evaluationcoding-agent +4 Impact 2 Eng 3 Adoption 1
2026-08-26 Foundation Models WATCH

Qwen open-weights Flash-Next, an architecture preview for Qwen4

AlibabaQwen open-weightsarchitecture-preview +4 Impact 4 Eng 3 Adoption 3

2026-08-25

3 events
2026-08-25 Developer Tools WATCH

Gemini CLI 0.57.0 ships the A2A server package in its stable line

Google gemini-clia2a +4 Impact 4 Eng 3 Adoption 2
2026-08-25 Foundation Models TRIAL

Z.ai open-sources GLM-5.3-Flash (new base, hybrid attention, MIT)

Zhipu AIZ.ai glmzai +6 Impact 4 Eng 4 Adoption 3
2026-08-25 Infrastructure WATCH

OpenAI's Jalapeño chip posts first benchmark wins over Nvidia systems

OpenAIBroadcom inference-chipcustom-silicon +2 Impact 4 Eng 3 Adoption 2

2026-08-22

1 events

2026-08-21

4 events
2026-08-21 Foundation Models TRIAL

OpenAI cuts GPT-5.6 Sol API and credit prices by over 20% for three months

OpenAI gpt-5.6pricing +4 Impact 1 Eng 4 Adoption 3
2026-08-21 Research WATCH

Study: LLM compression quietly hurts common knowledge and calibration

academic compressionquantization +5 Impact 2 Eng 4 Adoption 2
2026-08-21 Research WATCH

ReCache makes KV cache reusable for tool-heavy LLM agents

academic kv-cacheinference +7 Impact 3 Eng 4 Adoption 2
2026-08-21 Research WATCH

StateMemBench shows agent memory systems can't track a changing world

academic agent-memorymemory +5 Impact 3 Eng 4 Adoption 2

2026-08-20

3 events
2026-08-20 AI Engineering TRIAL

Mistral launches Agentic Search, a multi-step document retrieval layer

Mistral AI mistralagentic-search +5 Impact 3 Eng 4 Adoption 3
2026-08-20 Research WATCH

Compress and Forget: 4-bit quantization worsens recall of overwritten values

academic quantizationbitsandbytes +7 Impact 2 Eng 4 Adoption 1
2026-08-20 Research WATCH

SMTrap: CPU-only, feedback-free attacks exhaust large reasoning models

academic securitydos +8 Impact 3 Eng 3 Adoption 1

2026-08-19

4 events
2026-08-19 Developer Tools TRIAL

Cursor turns cloud agents into an always-on event-driven system

AnysphereCursor cloud-agentssubscriptions +7 Impact 3 Eng 4 Adoption 3
2026-08-19 Research WATCH

PTXBench: LLMs emit exotic PTX instructions but can't beat tuned libraries

academic gpu-kernelsptx +9 Impact 3 Eng 3 Adoption 1
2026-08-19 Research WATCH

Aggregate benchmark gains hide per-item regressions in LLM API migrations

academic llm-apimodel-migration +7 Impact 2 Eng 4 Adoption 1
2026-08-19 Foundation Models TRIAL

Grok 4.6 goes multi-cloud: Bedrock first, then Google Enterprise Agent Platform

xAIAmazon grok-4-6amazon-bedrock +8 Impact 2 Eng 3 Adoption 3

2026-08-18

5 events
2026-08-18 Developer Tools ADOPT

Codex CLI adds session forking, agents dashboard, cross-session messaging

OpenAIAmazon codexcli +7 Impact 3 Eng 4 Adoption 4
2026-08-18 Infrastructure WATCH

Cerebras unveils rack-scale CS-4 system with 750 PFLOPS of AI compute

Cerebras cs-4wse-3-turbo +4 Impact 4 Eng 3 Adoption 2
2026-08-18 Research WATCH

Study: leftover KV cache quietly breaks rollback in language agents

academic agentskv-cache +7 Impact 3 Eng 4 Adoption 1
2026-08-18 Agent Security WATCH

OpenAI slows frontier training after model's autonomous Hugging Face breach

OpenAIHugging Face securityincident +5 Impact 4 Eng 2 Adoption 2
2026-08-18 AI Engineering WATCH

OpenAI previews Private Safety Processing with zero data retention

OpenAI privacyzdr +4 Impact 2 Eng 3 Adoption 3

2026-08-17

2 events
2026-08-17 Research WATCH

Agentic Transaction brings ACID-style transaction semantics to LLM agents

academicTsinghua University agentsreliability +5 Impact 3 Eng 3 Adoption 1
2026-08-17 Developer Tools WATCH

Cursor launches Origin, a code hosting platform built for agents

AnysphereCursor origincode-hosting +7 Impact 3 Eng 3 Adoption 3

2026-08-14

8 events
2026-08-14 Foundation Models TRIAL

Alibaba open-sources Qwen3.8-27B with native 262K context

AlibabaQwen open-weightsapache-2 +4 Impact 4 Eng 5 Adoption 4
2026-08-14 Foundation Models TRIAL

Zhipu releases GLM-5.3, claiming the strongest open-weights coding model

Zhipu AIZ.ai glmzhipu +4 Impact 4 Eng 4 Adoption 3
2026-08-14 Agent Security TRIAL

Cloudflare One Gateway adds MCP traffic detection and enforcement

Cloudflare mcpsecurity +3 Impact 4 Eng 4 Adoption 3
2026-08-14 Developer Tools ADOPT

Claude Code enables subagent forking by default, adds cross-session messaging

Anthropic claude-codemulti-agent +3 Impact 3 Eng 4 Adoption 4
2026-08-14 Open Source WATCH

DeepSeek open-sources its agent runtime Harness (dsh) under MIT

DeepSeek agent-runtimeopen-source +2 Impact 3 Eng 4 Adoption 4
2026-08-14 Business & Policy WATCH

SpaceX closes $60B acquisition of Cursor (Anysphere)

SpaceXSpaceXAI acquisitionide +4 Impact 2 Eng 2 Adoption 4
2026-08-14 Foundation Models WATCH

Anthropic explains how Claude's text watermark works

Anthropic watermarkprovenance +1 Impact 3 Eng 2 Adoption 2
2026-08-14 Developer Tools WATCH

ChatGPT desktop for Linux enters public preview with Codex

OpenAI chatgptlinux +2 Impact 1 Eng 2 Adoption 3

2026-08-13

8 events
2026-08-13 Research WATCH

DARTree: draft trees push diffusion speculative decoding to 9.73x lossless speedup

academic (see paper) speculative-decodingdiffusion +2 Impact 4 Eng 3 Adoption 1
2026-08-13 Research WATCH

DCD: decoupled contrastive decoding hits 1.65-1.95x speedup, code released

academic (see paper) decodingcontrastive +2 Impact 3 Eng 3 Adoption 1
2026-08-13 Research TRIAL

Verifier finds ~40% of accepted LLM-generated GPU kernels are broken

academic (see paper) verificationgpu-kernels +2 Impact 3 Eng 4 Adoption 2
2026-08-13 Foundation Models TRIAL

Google ships Gemini 3.7 Flash, its workhorse model for coding and agents

GoogleGoogle DeepMind geminiflash +4 Impact 3 Eng 5 Adoption 4
2026-08-13 Foundation Models TRIAL

DeepSeek V4-Pro goes GA with native OpenAI Responses API support

DeepSeek v4-proga +3 Impact 3 Eng 4 Adoption 3
2026-08-13 Infrastructure WATCH

OpenAI previews Ultrafast mode: GPT-5.6 Sol on Cerebras at up to 750 tok/s

OpenAICerebras inferencespeed +2 Impact 4 Eng 3 Adoption 2
2026-08-13 Developer Tools TRIAL

GitHub Copilot weekly: Agent Plugins 1.0 GA, JetBrains adds Ollama BYOK

GitHubMicrosoft github-copilotagent-plugins +4 Impact 3 Eng 3 Adoption 3
2026-08-13 Developer Tools WATCH

Cursor Builds: pre-warmed environments boot cloud agents up to 3x faster

AnysphereCursor cloud-agentsbuilds +1 Impact 2 Eng 3 Adoption 3