TIL

26 things learned
August 2026
Aug 27 Quantization

AWQ found that protecting roughly 1% of salient weights (identified through activation statistics) can dramatically reduce quantization error. A tiny fraction of weights can disproportionately control model quality.

Aug 19 IoT

Many inexpensive sensors labeled eCO2 contain no CO2-sensing element. Chips such as the SGP30 and CCS811 infer an equivalent CO2 value from volatile organic compounds; actual CO2 measurement requires technology such as NDIR or photoacoustic sensing.

Aug 14 vLLM

Repetition penalties see tokens, not meaning. In vLLM, repetition_penalty also examines the prompt, so repeatedly necessary JSON tokens (quotes, braces, commas, and key names) can be penalized alongside the repetition you actually wanted to stop.

Aug 8 VLMs

For Qwen2.5-VL, one visual token represents roughly a 28 × 28-pixel region. An uncapped 1080p image therefore becomes about 2,600 visual tokens before any text is processed, making image resolution a direct inference-cost control.

Aug 2 Semantic Cache

Embedding similarity can erase the word that matters most. Queries such as bread and gluten-free bread may remain dangerously close in vector space, so a semantic cache can confidently return the wrong answer unless entities, attributes, and negation also match.

June 2026
Jun 29 vLLM

vLLM caches only complete KV blocks, and every block’s hash includes its parent’s hash. With 16-token blocks, a 53-token prompt can reuse at most 48 tokens, and changing one early token invalidates every cached block after it. Static text first, dynamic text last.

Jun 25 Quantization

SmoothQuant moves the quantization problem instead of eliminating it: it mathematically transfers difficulty from activation outliers into the weights using inverse channel-wise scaling, while leaving the full-precision computation equivalent.

Jun 22 vLLM

A Python version change helped create a real LLM-serving vulnerability. Python 3.12 made hash(None) predictable, making crafted collisions in vLLM’s prefix cache more feasible and potentially reusing KV blocks from different content. This became CVE-2025-25183.

Jun 11 Inference

Triton’s response cache hashes the model name, version, and exact input tensors, so it only hits on identical requests. It also cannot cache decoupled streaming models (the execution pattern commonly used by LLM backends), making it the wrong layer for semantic LLM caching.

Jun 8 LLMs

Speculative decoding can preserve the target model’s exact output distribution and still make serving slower. It shines when decoding is memory-bandwidth-bound at low concurrency; once batches saturate compute, rejected draft tokens become extra work.

Jun 7 Fine-tuning

Hugging Face Trainer may retain every evaluation logit on the accelerator. For Qwen’s 151,936-token vocabulary, a single 8 × 2,048 batch produces about 4.6 GiB of bf16 logits, before counting the model or later batches.

Jun 5 LLM Architecture

DeepSeek-V4’s causal encoder-decoder architecture activates roughly 8B parameters per token while reading the prompt but 16B while generating. Prefill and decode can effectively use different-sized models inside the same architecture.

May 2026
May 28 Evals

LLM judges can prefer outputs from their own model family partly because those outputs have lower perplexity, and therefore look more familiar, to the judge. Reversing A/B order attacks position bias, but it does nothing to remove family bias.

May 25 Claude Code

Typing ultrathink anywhere in a Claude Code prompt requests deeper reasoning for that single turn without changing your session effort level. Unlike older builds, ‘think’, ‘think hard’, and ‘think more’ are no longer special keywords; they’re passed through as ordinary prompt text.

May 25 Claude Code

/model opusplan runs Opus during plan mode then auto-switches to Sonnet for execution (Opus reasoning, Sonnet efficiency). Subtlety: the automatic 1M-token context upgrade applies only to plain opus; opusplan’s plan phase stays on the standard 200K window.

May 25 LLMs

Speculative decoding uses a small ‘draft’ model to propose several tokens that the large model verifies in one parallel forward pass. It’s mathematically lossless: the output distribution is identical to running the big model alone, and you just get the speedup for free on accepted tokens.

May 25 LLMs

Even at temperature=0, LLM outputs aren’t bitwise-deterministic on GPU. Floating-point addition isn’t associative, so parallel reductions and changing batch sizes can flip the argmax whenever two top logits are near-tied.

May 25 LLMs

Transformers dump a surprising amount of attention onto the very first token (an ‘attention sink’ that acts as a no-op dumping ground). StreamingLLM exploits this by always keeping the first few tokens in the KV cache, which stabilizes generation for effectively infinite-length streaming.

May 25 MoE

In a Mixture-of-Experts model a router fires only a few experts per token, so total vs. active parameter counts diverge hugely: a model can advertise hundreds of billions of params while only a small fraction actually compute on any given token. That’s why MoE models punch above their inference cost.

May 25 LLMs

KV cache grows linearly with sequence length (it’s attention compute that’s quadratic). Per token it costs 2 × layers × KV heads × head_dim × bytes, so GQA shrinks it by the ratio of query heads to KV heads. Llama 3 8B (32 query heads, 8 KV heads) needs 128 KB per token instead of 512 KB.

May 25 LLMs

Temperature zero does not make GPU inference bitwise deterministic. Parallel reductions can add floating-point values in different orders, and because addition is not associative at finite precision, a near-tied argmax can flip.

May 25 LLMs

Transformers often dump attention onto the first token even when it carries no useful meaning. StreamingLLM preserves these attention-sink tokens alongside a sliding window and can keep generation stable across millions of streamed tokens.

May 24 Python

On CPython 3.11, d.get(k) beats ‘k in d’ + d[k] when the key exists (one lookup instead of two), but on a miss the plain ‘in’ check is slightly faster since it skips the method call. try/except KeyError is fastest on hits and ~4x slower on misses, so pick based on your hit rate.

May 22 CUDA

torch.compile() with mode=’reduce-overhead’ uses CUDA graphs to replay kernels without Python overhead. Biggest win on small models with repeated forward passes.

May 20 RAG

Hybrid search (BM25 + dense retrieval) consistently outperforms either alone in production. The BM25 leg catches exact keyword matches that embeddings miss.

May 18 ML

Gradient checkpointing trades ~30% more compute for a ~60-70% reduction in activation memory. Usually worth it when batch size is the bottleneck.

No entries match.