hn • r/hackernews
Comment on: Running Kimi K3 on a M1 Max
Turns out: yes, it's doable, but slow af.But I've spent a long time frustrated with the quality ceiling on local models, and quantization is a big part of that ceiling.
Developers struggle to run high‑quality LLMs locally on M1 Max because quantization trade‑offs slow inference and lower fidelity. QuantEdge automatically compiles, quantizes, and tunes models for Apple Silicon, delivering near‑real‑time performance while preserving output quality.