The Kaitchup – AI on a Budget
Subscribe
Sign in
Home
Notes
AI Notebooks
The Kaitchup's Book
Weekly Kaitchup
Tutorials
Archive
About
Tutorials
Latest
Top
Discussions
Laguna S 2.1: How Agent Harnesses and Inference Budgets Shape Coding Performance
Testing long-horizon coding performance on DeepSWE and Terminal-Bench 2.1
Aug 13
•
Benjamin Marie
13
5
Make Your Own Optimized GGUFs with AutoRound
Build optimized GGUF models for llama.cpp and LM Studio using AutoScheme, custom bit-widths, and layer protection.
Jun 10
•
Benjamin Marie
3
DFlash vs MTP: Qwen3.6 Speculative Decoding Benchmarks with vLLM and llama.cpp
Up to 4x faster inference -- Benchmarking the speed on various task on coding, chat, and math tasks, with optimal hyperparameters
Jun 2
•
Benjamin Marie
10
1
1
Reasoning Budgets vs. Structured CoT: Controlling LLM Thinking Tokens
Evaluations of BNF grammars and reasoning budgets with Qwen3.6 27B
May 25
•
Benjamin Marie
10
6
2
Train and Run DFlash Speculative Decoding
A simple method to make your local model much faster
May 18
•
Benjamin Marie
11
1
How to Reduce LLM Inference Cost and Improve Accuracy with Pass@k and Majority Voting
Is thinking disabled + multiple retries better and still more efficient than thinking enabled?
Apr 27
•
Benjamin Marie
16
3
1
The KV-Cache of Small MoEs: Qwen3, Qwen3.5/3.6, GLM 4.7 Flash, and Nemotron 3 Nano Compared
A memory-first look at four efficient open LLM architectures.
Mar 18
•
Benjamin Marie
27
2
Qwen3.5 Quantization: Similar Accuracy, More Thinking — Best Models and Recipes
INT4, NVFP4, and FP8 evaluations — Thinking off and on
Mar 12
•
Benjamin Marie
34
7
1
How to Deploy Your LLM in the Cloud
The simple recipe to choose your GPU and anticipate costs
Feb 23
•
Benjamin Marie
8
GLM-5 Memory Requirements Explained: MLA + DeepSeek Sparse Attention (DSA)
How GLM-5 fits 200K context without terabytes of KV cache, and what GPUs you need.
Feb 16
•
Benjamin Marie
5
2
Serving ExLlamaV3 Models with tabbyAPI: Accuracy, Speed, and Recommendations
With comparisons against AutoRound and GGUF models served with vLLM
Jan 19
•
Benjamin Marie
7
4-bit GLM-4.7 (358B) on a Single NVIDIA B300 with vLLM: AWQ vs NVFP4 vs INT4
Just give it enough tokens to think
Jan 12
•
Benjamin Marie
8
3
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts