TIL
26 things learnedAWQ found that protecting roughly 1% of salient weights (identified through activation statistics) can dramatically reduce quantization error. A tiny fraction of weights can disproportionately control model quality.
Many inexpensive sensors labeled eCO2 contain no CO2-sensing element. Chips such as the SGP30 and CCS811 infer an equivalent CO2 value from volatile organic compounds; actual CO2 measurement requires technology such as NDIR or photoacoustic sensing.
Repetition penalties see tokens, not meaning. In vLLM, repetition_penalty also examines the prompt, so repeatedly necessary JSON tokens (quotes, braces, commas, and key names) can be penalized alongside the repetition you actually wanted to stop.
For Qwen2.5-VL, one visual token represents roughly a 28 × 28-pixel region. An uncapped 1080p image therefore becomes about 2,600 visual tokens before any text is processed, making image resolution a direct inference-cost control.
Embedding similarity can erase the word that matters most. Queries such as bread and gluten-free bread may remain dangerously close in vector space, so a semantic cache can confidently return the wrong answer unless entities, attributes, and negation also match.
vLLM caches only complete KV blocks, and every block’s hash includes its parent’s hash. With 16-token blocks, a 53-token prompt can reuse at most 48 tokens, and changing one early token invalidates every cached block after it. Static text first, dynamic text last.
SmoothQuant moves the quantization problem instead of eliminating it: it mathematically transfers difficulty from activation outliers into the weights using inverse channel-wise scaling, while leaving the full-precision computation equivalent.
A Python version change helped create a real LLM-serving vulnerability. Python 3.12 made hash(None) predictable, making crafted collisions in vLLM’s prefix cache more feasible and potentially reusing KV blocks from different content. This became CVE-2025-25183.
Triton’s response cache hashes the model name, version, and exact input tensors, so it only hits on identical requests. It also cannot cache decoupled streaming models (the execution pattern commonly used by LLM backends), making it the wrong layer for semantic LLM caching.
Speculative decoding can preserve the target model’s exact output distribution and still make serving slower. It shines when decoding is memory-bandwidth-bound at low concurrency; once batches saturate compute, rejected draft tokens become extra work.
Hugging Face Trainer may retain every evaluation logit on the accelerator. For Qwen’s 151,936-token vocabulary, a single 8 × 2,048 batch produces about 4.6 GiB of bf16 logits, before counting the model or later batches.
DeepSeek-V4’s causal encoder-decoder architecture activates roughly 8B parameters per token while reading the prompt but 16B while generating. Prefill and decode can effectively use different-sized models inside the same architecture.
LLM judges can prefer outputs from their own model family partly because those outputs have lower perplexity, and therefore look more familiar, to the judge. Reversing A/B order attacks position bias, but it does nothing to remove family bias.
Typing ultrathink anywhere in a Claude Code prompt requests deeper reasoning for that single turn without changing your session effort level. Unlike older builds, ‘think’, ‘think hard’, and ‘think more’ are no longer special keywords; they’re passed through as ordinary prompt text.
/model opusplan runs Opus during plan mode then auto-switches to Sonnet for execution (Opus reasoning, Sonnet efficiency). Subtlety: the automatic 1M-token context upgrade applies only to plain opus; opusplan’s plan phase stays on the standard 200K window.
Speculative decoding uses a small ‘draft’ model to propose several tokens that the large model verifies in one parallel forward pass. It’s mathematically lossless: the output distribution is identical to running the big model alone, and you just get the speedup for free on accepted tokens.
Even at temperature=0, LLM outputs aren’t bitwise-deterministic on GPU. Floating-point addition isn’t associative, so parallel reductions and changing batch sizes can flip the argmax whenever two top logits are near-tied.
Transformers dump a surprising amount of attention onto the very first token (an ‘attention sink’ that acts as a no-op dumping ground). StreamingLLM exploits this by always keeping the first few tokens in the KV cache, which stabilizes generation for effectively infinite-length streaming.
In a Mixture-of-Experts model a router fires only a few experts per token, so total vs. active parameter counts diverge hugely: a model can advertise hundreds of billions of params while only a small fraction actually compute on any given token. That’s why MoE models punch above their inference cost.
KV cache grows linearly with sequence length (it’s attention compute that’s quadratic). Per token it costs 2 × layers × KV heads × head_dim × bytes, so GQA shrinks it by the ratio of query heads to KV heads. Llama 3 8B (32 query heads, 8 KV heads) needs 128 KB per token instead of 512 KB.
Temperature zero does not make GPU inference bitwise deterministic. Parallel reductions can add floating-point values in different orders, and because addition is not associative at finite precision, a near-tied argmax can flip.
Transformers often dump attention onto the first token even when it carries no useful meaning. StreamingLLM preserves these attention-sink tokens alongside a sliding window and can keep generation stable across millions of streamed tokens.
On CPython 3.11, d.get(k) beats ‘k in d’ + d[k] when the key exists (one lookup instead of two), but on a miss the plain ‘in’ check is slightly faster since it skips the method call. try/except KeyError is fastest on hits and ~4x slower on misses, so pick based on your hit rate.
torch.compile() with mode=’reduce-overhead’ uses CUDA graphs to replay kernels without Python overhead. Biggest win on small models with repeated forward passes.
Hybrid search (BM25 + dense retrieval) consistently outperforms either alone in production. The BM25 leg catches exact keyword matches that embeddings miss.
Gradient checkpointing trades ~30% more compute for a ~60-70% reduction in activation memory. Usually worth it when batch size is the bottleneck.
No entries match.
