hn • r/hackernews
Comment on: Homebench – Benchmark local LLMs for speed, memory, and quality
Author here. I had a pile of models pulled locally and no way to answer "which of these is actually good, and what does it cost me in speed?" llama-bench gives you tok/s and nothing about output quality; lm-evaluation-harness gives you quality but isn't built around Ollama or LM Studio, which is what most people are actually running at home. So this measures both in one pass and puts them in one table.The part that turned out to be harder than expected was deciding what the numbers mean:tok/s sounds trivial until you pick a denominator. Prompt processing? Model load? I exclude both, use Ollama