burnt-synapse

← back to blog

#llama.cpp

every post tagged llama.cpp

22 August 2026

qwen3.8 27b on $500 of scrap

qwen3.8 27b — q8 weights, f16 kv cache — doing 200 t/s of prompt processing and 30 t/s of generation on one box of $500 scrap. no gpu in sight.

22 August 2026

one million tokens, ten hours, three cves

opencode, pointed at a local qwen, left alone for about ten hours. it came back with roughly a million output tokens, three promising cves, and proof-of-concept scripts that actually worked. clean run — no restarts, no…