今日已更新 166 条资讯 | 累计 40611 条内容
关于我们

Proposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]

/u/Specific-Tax-6700 2026年09月07日 02:41 0 次阅读 来源:Reddit r/MachineLearning

I ported MoE expert expansion to llama.cpp 🚀 Run MoE models with MORE routed experts than the native top-K (8->x), adaptive threshold, 99→50% influence decay, layer range. Runtime-only, all backends. Tested on Qwen 3.6 35B A4B+ https://github.com/vagrillo/llama.cpp/blob/moe-expansion/docs/moe-expansion.md submitted by /u/Specific-Tax-6700 [link] [留言]

本文内容来源于互联网,版权归原作者所有
查看原文