Agens Volundr 32B Preview logo

Agens Volundr 32B Preview

Open 32B model: only 18 of its 72 layers keep a KV cache

Artificial Intelligence Developer Tools Open Source

Agens Volundr is a dense 32B open-weights model for agents and long working sessions. Only 18 of its 72 layers keep a KV cache; the other 54 are linear attention with a fixed-size state, so long context stays affordable on hardware you own. Decode holds at 23.9 tok/s even at 128K context (BF16, two 48 GB GPUs, single user). 262K window, Apache-2.0, BF16 and INT4 (one 48 GB GPU). Runs on our open sglang build; GGUF planned.

投票数: 0
← 投稿一覧に戻る