
August 19, 2026
π€β‘ Ask Mauro: Why I Went Smaller, Not Bigger
I got tired of people browsing my resume, so I made myself queryable. Ask Mauro runs on a Qwen3 0.6B small language model on a single 8 vCPU / 24 GB server β and going smaller instead of bigger took the average response time from roughly 68 seconds to roughly 17. Here is the architecture reasoning behind that call.









