Summary

Tencent released and open-sourced Hy4 preview, its next-generation LLM: a 770B-parameter MoE with 49B active per token and a context window beyond 1M tokens, shipped as standard and FP8 weights under Apache 2.0 (per HF tags; the announcement itself names no license). Target uses are software engineering (planning, debugging, front-end quality), office work (financial and data analysis, cross-document collaboration), game development (playable prototypes from a single prompt) and scientific research. Public benchmark evidence is thin: one internal Tencent blind test with 163 experts over 203 engineering tasks scored Hy4 preview 2.99/4.00, slightly ahead of GLM-5.3 (2.92) and Kimi K3 (2.94). Tencent says the model assisted its own training pipeline and self-optimized its inference stack for a 31.8% end-to-end throughput gain. The full Hy4 is not released yet — a preview-first rollout with the next batch 'soon'. API access via Tencent Cloud TokenHub and OpenRouter at $0.834/M input, $2.501/M output, $0.042/M cached.

Why it matters
This is the third Chinese lab in one week to put major open weights on Hugging Face (Z.ai 8/25, Qwen 8/24-26, Tencent 8/27-28), and at 770B Apache 2.0 it is the closest thing yet to a same-tier, different-organization release for the open-weight frontier — direct evidence for trend #1. But treat capability claims as unverified: no public benchmark scores, one vendor-run blind test, and a preview badge. Serving teams get a vLLM recipe day-0; expect a real read on the model only when public harnesses (SWE-bench, Terminal-Bench) publish runs.
Technical details
Architecture 770B total / 49B active MoE; context beyond 1M tokens
License Apache 2.0 per Hugging Face tags (announcement text names no license)
Artifacts tencent/Hy4-preview + tencent/Hy4-preview-FP8 on HF (repos created 2026-08-27T08:52Z); ModelScope mirror; vLLM recipe day-0; downloads ~1.4k main + ~1.3k FP8, 283 likes @2026-08-30
Eval internal blind test only: 163 experts, 203 engineering tasks; Hy4 preview 2.99/4.00 vs GLM-5.3 2.92, Kimi K3 2.94; no public benchmark scores (no SWE-bench/Terminal-Bench)
Self Improvement model assisted its own training pipeline; self-optimized inference system +31.8% end-to-end throughput (vendor claim)
API Tencent Cloud TokenHub + OpenRouter; $0.834/M input, $2.501/M output, $0.042/M cached; free two weeks on WorkBuddy/CodeBuddy; Hy3 free extended to 2026-09-30
Rollout preview-first; full Hy4 weights not yet released; next Hy4-series batch 'expected soon'
Updates
2026-08-31 Adoption & follow-up check @2026-08-31T00:00Z: downloads 2,123 main (+~700 vs the 8/30 sample) + 319 likes; FP8 1,469 — modest velocity compared with the GLM-5.3-Flash cohort. No full Hy4 release yet; the announcement still only says the next batch is expected 'soon' (community speculates ~2.5 months to GA from the Hy3 precedent). Verification note: aggregators (datalearner etc.) list public benchmark scores (Terminal Bench 2.1 85.4, SWE-Bench Pro 65.7) that do NOT appear in the official announcement or the HF model card — unverified, not counted; the official page still shows only the internal blind test. Community GGUF compression to ~200GB claimed at ~98% performance retention (unverified). Recommendation stays WATCH.
2026-09-02 adoption check @2026-09-02T00:0xZ: 3,516 downloads (+66% vs 2,123 @8/31) + 383 likes — continued but still modest vs the GLM-5.3-Flash cohort; full Hy4 not shipped ('soon' unchanged), no official public benchmarks. Trend #1 criterion (b) candidate status unchanged. Stays WATCH
2026-09-05 Adoption check @2026-09-05: 5,684 downloads (+62% vs 9/2) + 430 likes; FP8 2,666. Full Hy4 still not shipped (tencent org's latest repos remain Hy4-preview/Hy4-preview-FP8 from 8/27; flagship line still Hy3 from July). Third-party quantization ecosystem growing despite the preview: AngelSlim/Hy4-preview-GGUF 109,240 downloads; mlx-community 4bit; inferencerlabs MLX Q4i (9/2); anemll FlashMoE-STQ1_0 (9/3). Stays WATCH
Tags
open-weightsmoelong-contextapache-2codingtencent