mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-30 17:11:19 +02:00
0377426cef
The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the second time passing n_ff_chexp. gguf_set_val_u32 removes-then-appends, so the second call clobbers the first: the saved shared_feed_forward_length ends up as n_ff_chexp (0 for every arch except GroveMoE), and expert_chunk_feed_forward_length is never written at all. So a save->load roundtrip of any MoE model with a shared expert loses n_ff_shexp. On reload the arch falls back to n_ff for the shexp tensor shape, that no longer matches the saved tensor, and the model FAILS to load. Hits qwen2moe, qwen3-next, granite-moe, hunyuan-moe, ernie4.5, bailingmoe2, nemotron-h, and the other shared-expert MoEs. Fix: the second call writes LLM_KV_EXPERT_CHUNK_FEED_FORWARD_LENGTH. test-llama-archs: set expert_shared_feed_forward_length to a value distinct from n_ff in the MoE setup so the roundtrip exercises it. Without the fix the reload fails on a shexp tensor-shape mismatch; with it, every arch roundtrips clean.