| Quantizing models without losing the edge cases | quantization, gguf | 9 min |
| Serving a 70B mixture-of-experts model on a single workstation | moe, serving | 11 min |
| Why we publish an eval harness with every release | evaluation, reproducibility | 7 min |
| Token economics: batching, the KV cache and the p99 cliff | performance, serving | 10 min |
| Fine-tuning small models on a budget | fine-tuning, lora | 8 min |
| How uploaded weights get verified | supply-chain, verification | 6 min |
| Speculative decoding in production | performance, decoding | 7 min |
| Choosing a license for open weights | licensing, policy | 8 min |