Tencent Hy3: A 295B Open MoE with 21B Active

Z

ZharfAI Research

Model release desk

July 6, 2026Updated August 6, 20267 min read
Tencent Hy3: A 295B Open MoE with 21B Active

Tencent officially released Hy3 on July 6, 2026 after an April preview. The corporate announcement describes a 295-billion-parameter mixture of experts with 21 billion active parameters and 256K context. The weights and an FP8 checkpoint are available under Apache 2.0.

Hy3 turns feedback from more than fifty products into a post-trained model aimed at coding, office work, financial modeling, frontend design, games, and computer agents. Tencent says preview token use grew twentyfold and uses that production traffic to motivate reliability work. The evidence mixes public benchmarks, a blind expert study, and internal defect rates.

Architecture and efficiency

The official model card lists 80 layers, 192 routed experts with top-eight selection, 64 attention heads with eight KV heads, a 4,096 hidden size, and a 3.8B multi-token-prediction layer. BF16 weights and FP8 distribution are available.

Only 21B of 295B parameters are active per token, giving the model a relatively small compute path for its stored capacity. The 3.8B MTP layer proposes future tokens to accelerate generation. A 256K window supports large repositories and office packets without reaching the million-token infrastructure burden of some peers.

Open weights enable private serving, but 295B remains a multi-GPU checkpoint. The active count does not describe storage, expert communication, or cache requirements.

A small active route travels through a much larger modular engine, with context and office tools around it.
A small active route travels through a much larger modular engine, with context and office tools around it.

Benchmark snapshot

The first-party card and Hugging Face evaluation metadata report:

EvaluationHy3What it samples
SWE-bench Verified78.0Repository issue resolution
SWE-bench Pro57.9Harder active-repository tasks
GPQA Diamond90.4Graduate-level science reasoning
HLE53.2Broad difficult questions under the card setup
SkillsBench55.3Skill and tool-oriented tasks
Long-Horizon Terminal Bench28.8Extended terminal work
Blind expert production study2.67 / 4270 experts on real work; GLM-5.1 scored 2.51
Multi-turn issue rate7.9%Internal reliability, down from 17.4%

The model card also contains a benchmark appendix image with more rows. Use the exact card revision for procurement, because leaderboard metadata can update after publication. The blind study is relevant to product utility but not independently reproducible.

From preview to official release

The preview launched April 23, outside the core window. The July model is a new post-trained official release built from product feedback and higher-quality data. The date is not merely a delayed blog post.

Tencent says fifty-plus products contributed feedback, including WorkBuddy, Yuanbao, ima, and Marvis. This can improve practical instruction following and tool coordination. It can also overfit the model toward the company's application distribution. External teams need their own tasks.

Preview traffic growing twentyfold is adoption evidence, not quality proof. The official model should be compared to the preview under a matched harness to isolate improvements.

Product reliability metrics

Tencent emphasizes fact conflation, fabrication, contradiction, multi-turn coreference, ellipsis recovery, and constraint inheritance. The reported multi-turn issue rate falls from 17.4 to 7.9 percent. A blind study gives Hy3 2.67 out of four against 2.51 for GLM-5.1, with larger gains in frontend, data and storage, and CI/CD.

These evaluations address failures users actually experience. They are internal and need local reproduction. Define the denominator: how an “issue” is counted, how many conversations, which languages, and whether graders are human or model-based.

For Persian workflows, test long conversations with changing dates, amounts, names, and permissions. The model must follow the latest explicit instruction without losing earlier identifiers.

Coding and agent performance

SWE-bench Verified at 78 and Pro at 57.9 place Hy3 in a credible open-agent tier. The larger gap between Verified and Pro shows that difficult, larger-diff repositories remain challenging. Long-Horizon Terminal Bench at 28.8 is another reminder that extended work is much harder than standard issue resolution.

Use tests, minimal diffs, and review rubrics. For office and financial artifacts, validate formulas, references, and source evidence. A model that produces an attractive spreadsheet can still contain a silent logic error.

Tool calls require typed schemas and permissions. Computer agents need isolated accounts and confirmation. Our software agent guide covers those execution controls.

Open deployment and API access

Apache 2.0 weights are distributed through Hugging Face and other first-party channels. Tencent Cloud TokenHub offers an API and third-party platforms are adding support. Self-hosting and API use have different privacy, latency, and reproducibility profiles.

Pin revision, license, tokenizer, chat template, precision, and inference engine. Compare BF16 and FP8 on code, multilingual text, tools, and long context. Expert routing and MTP support must be verified in the chosen runtime.

The open-model operations guide provides the artifact chain. Treat community GGUF or other quantizations as separate derived releases with their own provenance and quality tests.

Best fit and verdict

Hy3 is attractive for organizations seeking an Apache-2.0 agent model with moderate active compute, 256K context, and product-oriented post-training. It can support coding, office, and computer workflows when the infrastructure can host 295B stored parameters.

It is not a small local model, and internal reliability gains are not universal guarantees. The strongest evidence is the transparent architecture, downloadable weights, public benchmark metadata, and explicit preview-to-release story. Reproduce the hard repository, long-terminal, and local-language cases before production.

Self-hosted acceptance matrix

Hy3's 295B total size and 21B active path make it an appealing sparse deployment, but active parameters do not determine the whole footprint. Operators must store and distribute all expert weights, route tokens across devices, maintain the 256K KV cache, and support the 3.8B multi-token-prediction layer. BF16 and FP8 checkpoints therefore need separate capacity and quality plans.

Start by pinning the Tencent repository revision, license, tokenizer, configuration, remote code, numerical format, serving engine, kernels, and container digest. Verify checksums before copying artifacts into an internal registry. Run custom code in an isolated build environment and generate a software bill of materials. Compare the hosted endpoint with each self-hosted precision on identical prompts before interpreting any performance change as a model difference.

Build a matrix across coding, tools, knowledge, long context, multilingual work, and failure recovery. For coding, include weak test suites, monorepos, generated files, and dependency failures. For tools, vary schemas, permission errors, stale state, and rate limits. For long context, place evidence at multiple positions among realistic distractors and measure effective recall, latency, and memory rather than only accepting a 256K request.

Report task success, unsafe attempts, hallucinated completion, output and hidden reasoning tokens where available, tool calls, time, accelerator memory, throughput, energy, and reviewer minutes. Run several seeds for nondeterministic agent tasks. A self-hosted build is accepted only when the organization can reproduce its quality after restart, fail over safely, and roll back to a known model revision without losing traceability.

The Apache 2.0 weights support modification and redistribution under their terms, but derived quantizations, fine-tunes, adapters, and serving templates become new governed artifacts. Each needs its own evaluation and license record. Openness creates deployment control; it does not remove operational accountability.

Source notes — reviewed 2026

#Tencent Hy3#Hunyuan#Open Weights#Mixture of Experts#AI Agents

Related Posts

Ready to Start Your AI Project?

Get in touch with our team to discuss how we can help your business.