The Kaitchup – AI on a Budget
Subscribe
Sign in
Home
Notes
AI Notebooks
The Kaitchup's Book
Weekly Kaitchup
Tutorials
Archive
About
Latest
Top
Discussions
ThinkingCap-Qwen3.6-27B Review: 2x Fewer Tokens, Same Accuracy?
A faster, more stable Qwen3.6 for local AI inference
Aug 5
•
Benjamin Marie
10
4
July 2026
DeepSeek-V4-Flash-0731 and Inkling Small: Smaller, but Better?
The Weekly Kaitchup #153
Jul 31
•
Benjamin Marie
7
Bonsai 27B Review: Can a 3.9 GB 1-Bit Model Match Qwen3.6 27B?
An in-depth look at Bonsai 27B’s accuracy, token efficiency, reasoning stability, and production trade-offs.
Jul 28
•
Benjamin Marie
10
5
1
Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1
The Weekly Kaitchup #152
Jul 25
•
Benjamin Marie
9
2
Qwen3.8: What Hardware Will You Need to Run Alibaba’s 2.4T Model?
Estimating the memory, storage, and GPU requirements for BF16, NVFP4, Q4, and TQ1 versions of Qwen3.8.
Jul 22
•
Benjamin Marie
8
1
Inkling, Gemma 4 Updates, and 1-Bit Qwen3.6
The Weekly Kaitchup #151
Jul 18
•
Benjamin Marie
16
Qwen3.6-27B KV Cache Quantization in vLLM: Accuracy, Memory, and Speed
A smaller KV cache enables longer sequences and higher concurrency with virtually no loss in accuracy.
Jul 16
•
Benjamin Marie
16
1
1
Efficient and Reasoning AI at the ACL 2026
The Weekly Kaitchup #150
Jul 11
•
Benjamin Marie
7
2
LFM2.5 230M and 350M: How Accurate Are the GGUF Versions?
230M or 350M GGUFs?
Jul 8
•
Benjamin Marie
10
2
DSpark and NVIDIA's Qwen3.6 NVFP4 Models
The Weekly Kaitchup #149
Jul 4
•
Benjamin Marie
16
2
June 2026
MiniMax M3 GGUF Quantization: From 852 GB to ~150 GB Without Breaking Accuracy
Benchmarks, token efficiency, and tensor-level analysis of low-bit M3 GGUFs.
Jun 30
•
Benjamin Marie
12
4
This Week in Open Models: Tiny LFM2.5, Ornith-1.0, and GLM-5.2 REAP
The Weekly Kaitchup #148
Jun 27
•
Benjamin Marie
8
4
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts