github.com favicon

GitHub - ggml-org/llama.cpp: LLM inference in C/C++ · GitHub

5
Public whispers
5
Contributors
2026-07-16 09:47:54
First whispered

Public whispers on this page

Text Highlight2026-07-16 13:08:54
Original Highlight Excerpt
"1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use"
Whisper Note
I tried 4-bit on my laptop, runs way smoother than full precision.
Text Highlight2026-07-16 12:59:54
Original Highlight Excerpt
"1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use"
Whisper Note
Damn, 1.5-bit sounds wild, does it even work or just a gimmick?
Text Highlight2026-07-16 10:05:54
Original Highlight Excerpt
"Plain C/C++ implementation without any dependencies"
Whisper Note
I tried to compile it on Windows and it was still a pain, but cool it's pure C.
Text Highlight2026-07-16 09:56:54
Original Highlight Excerpt
"Plain C/C++ implementation without any dependencies"
Whisper Note
Yeah, but no dependencies means you're on your own for everything else.
Text Highlight2026-07-16 09:47:54
Original Highlight Excerpt
"Plain C/C++ implementation without any dependencies"
Whisper Note
Finally something that doesn't need a bloated framework to run.

Share this page's whispers

Share to X
Short link
https://domwhisper.com/s/8f4d2163b5e6
Embed snippet
<iframe src="https://domwhisper.com/embed/8f4d2163b5e6" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

See what people are discussing on github.com

Install DomWhisper to view live whispers as you browse, and join the discussion.

Get the extension