Agens Volundr 32B Preview
ストックにはログインが必要です
Open 32B model: only 18 of its 72 layers keep a KV cache
Artificial Intelligence
Developer Tools
Open Source
Agens Volundr is a dense 32B open-weights model for agents and long working sessions. Only 18 of its 72 layers keep a KV cache; the other 54 are linear attention with a fixed-size state, so long context stays affordable on hardware you own. Decode holds at 23.9 tok/s even at 128K context (BF16, two 48 GB GPUs, single user). 262K window, Apache-2.0, BF16 and INT4 (one 48 GB GPU). Runs on our open sglang build; GGUF planned.
投票数: 0