
Hy4 preview is a new-generation Mixture-of-Experts (MoE) flagship model developed by the Tencent Hy Team. The model comprises 770B total parameters, of which 49B are activated per token. Model supports 1M tokens context length.
Native FP8 weights are published as well.
Benchmarks:

Known issues with the preview model: spends longer than necessary reasoning through complex tasks, and a tendency to over-verify its own work.
Open-weights frontier models will continue to drop until RAM prices improve 😅
You must log in or # to comment.

