Step-3.7-Flash on AMD: ROCm corrupts long context past ~94k, and thinking needs a hard token budget
Signal
72
Hype
15
In three linesStep-3.7-Flash on AMD/ROCm corrupts long context beyond ~94k tokens. Reasoning mode is on by default and consumes 2000+ tokens without a budget, causing empty responses. Workaround: cap context at 90k, set thinking_budget_tokens to 256 via llama.cpp, ignore enable_thinking:false and reasoning_effort.Read source
Your take?
Summary generated by Claude — human-verified