Tencent Open-Sources Hy4 Preview: A 770 Billion Parameter Model With a Million-Token Context

0
525

The Weight of Open Weights

Tencent dropped Hy4 preview on August 28, 2026, and if you run infrastructure for a living, you should care less about the press release and more about the bill of materials. This is a 770 billion total parameter Mixture-of-Experts model with 49 billion active parameters and a context window north of one million tokens. That '49B active' number is the only reason this thing is deployable at all. It is the same compute economics that made DeepSeek V4 Flash-style local inference viable, and it is the difference between a model that lives in a data center and one that can sit on a single high-end node with the right quantization. The weights are synced across HuggingFace, GitHub, ModelScope, and Gitcode, with API access through Tencent Cloud TokenHub and OpenRouter — plus two free weeks on WorkBuddy and CodeBuddy, and Hy3 access extended to September 30. That is a customer acquisition play dressed as a research release, and it is smart.

The MoE Math That Matters

Let me be direct about the architecture. A 770B total parameter model with 49B active parameters means you are loading a lot of weights into VRAM or system memory, but you only pay inference cost for the active slice on every token. That is the DeepSeek playbook, and it is now the standard for frontier-class open models. The practical reality: you need serious hardware to host the full 770B even if only 49B fire per token — multiple 80GB GPUs per node, or aggressive quantization that trades a few accuracy points to run on commodity hardware.

Tencent claims Hy4 preview is built on high-quality training data co-created with internal experts across software engineering, gaming, finance, and security. That is a meaningful differentiator. Most labs scrape the internet and hope. Tencent is injecting domain-specific data from people who actually run trading systems, build games, and defend networks. The result shows up in the blind evaluation: 163 internal experts, 203 engineering tasks, Hy4 preview scored 2.99/4.00, ahead of GLM-5.3 at 2.92 and Kimi K3 at 2.94. Those margins are thin, but they are consistent across a broad task set, not a single cherry-picked benchmark.

Million-Token Context Is a Different Operating Model

A million-token context window is not a marketing number if it actually works. It changes what you can point a model at: an entire codebase, a full regulatory filing, a decade of support tickets, a complete audit trail. Tencent is co-designing Hy4 preview with CodeBuddy and WorkBuddy for long-context development understanding, planning, debugging, validation, and office analysis from raw data processing to document-table delivery. They are also claiming playable game prototypes from a single requirement and research scenarios in molecular dynamics and condensed matter physics.

Here is the caveat nobody in the coverage is hammering hard enough: Tencent admits Hy4 preview is an early iteration with known issues in complex-task long thinking and excessive self-verification. That means the million-token context might give you a model that reads everything and then second-guesses itself into a loop. Test the long-context quality yourself, not the spec sheet. The gap between 'context window supports 1M tokens' and 'the model actually reasons well across 1M tokens' is where production systems go to die.

The Chinese Open-Source Race Is Real

Look at the competitive picture. Tencent Hunyuan, Alibaba Qwen, Zhipu GLM, Moonshot Kimi — Chinese labs are open-sourcing frontier-class models at a pace US labs are not matching. OpenAI, Anthropic, and Google keep weights behind APIs and charge per token. The Chinese strategy is different: release the weights, build the ecosystem, win the developers, and monetize through cloud services and vertical products. Tencent is giving away two weeks of free access on WorkBuddy and CodeBuddy, and it is syncing to every major model repository. That is a land grab, and it is working.

The blind evaluation result — Hy4 preview at 2.99, Kimi K3 at 2.94, GLM-5.3 at 2.92 — tells you these models are within noise of each other on real engineering tasks. The moat is not the model. The moat is the integration: CodeBuddy for developers, WorkBuddy for office workers, Yuanbao and ima for consumers, Tencent Cloud for enterprises. The model is the commodity; the product is the differentiation.

What 'Open Source' Actually Means Here

I need to flag the licensing question because it is the one thing that will bite you in production. Tencent says 'open-sourced,' but we need to read the actual terms. Open weights with strings attached is not open source. If the license restricts commercial use, requires attribution that exposes your proprietary data, or limits fine-tuning for specific verticals, that changes the calculus entirely. The pattern with Chinese labs has been mixed: some release under permissive licenses, others add clauses that protect their cloud business.

There is also the benchmark cherry-picking problem. Tencent organized the blind evaluation itself, with its own internal experts — not an independent audit. The 2.99/4.00 score is directionally useful, but it is not third-party verification. If you are making procurement decisions, run your own eval suite against your own workloads. The cost of a wrong model choice is far higher than the cost of a few days of testing.

What This Means

Open-weight frontier models are now a commodity layer. Tencent, Alibaba, Zhipu, and Moonshot are all shipping models that are 'good enough' for real work, and they are doing it in public. The differentiation moves up the stack: hardware optimization, context engineering, and vertical productization. If you are building on these models, your value is not in the weights — it is in what you do with them. The Blaschke-Lebesgue result — Hy4 and Hyra advancing the volume lower bound from 0.380799 to 0.41104, about 2% from the Meissner tetrahedron conjecture — is a nice proof point that these models can do novel research, but it is not a reason to deploy one. The reason to deploy is the cost per useful token on your specific workload.

The infrastructure reality is this: a 770B total parameter model is not a laptop toy. You need real hardware, real memory bandwidth, and real engineering to serve it. The 49B active parameter count makes it feasible, but 'feasible' is not 'cheap.' Self-hosting means budgeting for the full weight load, not the active inference cost. And if you are building a product on top of Hy4 preview, plan for the known issues: long thinking loops and self-verification overhead on complex tasks.

The Road Ahead

Hy4 final is coming, and it will be better. The question is whether US policy catches up to the reality that frontier-class open weights are now a Chinese export. The compute export controls were supposed to slow this down. They did not. Tencent is shipping a 770B model with a million-token context, free to download. That is the story. The math result is a footnote; the geopolitical and infrastructure implications are the headline.

For operators, the play is clear: test Hy4 preview against your real workloads, read the license terms line by line, and build your context engineering around the million-token window. Right now, the moat is in what you build on top of the weights, not the weights themselves.

— Allan Ali, Sylt.ing

Поиск
Категории
Больше
Generative AI & AI Art
How Canva Magic Studio Transforms Graphic Design for Teams and Creators
How Canva Magic Studio Transforms Graphic Design for Teams and Creators The Shift from Complex...
От Patty 2026-06-21 17:06:17 0 1Кб
AI Tools & Software
Why AI in Financial Close Is Becoming a CFO Priority
Why AI in Financial Close Is Becoming a CFO Priority The financial close is the one process...
От PriyaSharma 2026-08-16 23:12:30 0 400
AI Tools & Software
Terafab: Musk's 16.8 Billion Chip Megafactory Would Be the Largest Building on Earth
Terafab: Musk's 16.8 Billion Chip Megafactory Would Be the Largest Building on Earth Elon Musk...
От Allan 2026-08-11 20:11:20 0 1Кб
AI Tools & Software
AI Data Center Backlash Hits Record Levels: 30 Billion in Projects Blocked or Delayed
Community Backlash Has Become the AI Industry's Most Underestimated Bottleneck The numbers are...
От Allan 2026-07-17 20:10:06 0 4Кб
AI News & Updates
Why Every Developer Should Run Local LLMs in 2026
Why Every Developer Should Run Local LLMs in 2026 The Cloud API Bill Is Already Unsustainable...
От Jessica 2026-07-28 23:05:09 0 495