DEV Community

jidonglab profile picture

jidonglab

1 project a week. Building and sharing the entire process — from idea to shipped product in 7 days. Currently: AI news automation.

LLM-as-Judge Position Bias: Why Swapping A and B Flips Wins

LLM-as-Judge Position Bias: Why Swapping A and B Flips Wins

Comments
7 min read

Want to connect with jidonglab?

Create an account to connect with jidonglab. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
GQA KV Cache Replication: Why TP=16 Doubles Your Memory

GQA KV Cache Replication: Why TP=16 Doubles Your Memory

Comments
7 min read
MoE Expert Load Imbalance: Why Decode Stalls on One GPU

MoE Expert Load Imbalance: Why Decode Stalls on One GPU

Comments
8 min read
Chunked Prefill: Why max_num_batched_tokens=512 Kills Throughput

Chunked Prefill: Why max_num_batched_tokens=512 Kills Throughput

Comments
7 min read
rsLoRA: Why LoRA alpha/r Scaling Kills High-Rank Fine-Tunes

rsLoRA: Why LoRA alpha/r Scaling Kills High-Rank Fine-Tunes

Comments
7 min read
Sequence Packing: Why Cross-Document Attention Ruins Your SFT

Sequence Packing: Why Cross-Document Attention Ruins Your SFT

Comments
7 min read
BM25 Length Normalization: Why Long RAG Chunks Never Rank

BM25 Length Normalization: Why Long RAG Chunks Never Rank

Comments
7 min read
Matryoshka Embedding Truncation: Why 256 Dims Breaks Thresholds

Matryoshka Embedding Truncation: Why 256 Dims Breaks Thresholds

Comments
7 min read
Attention Entropy: Why Softmax Blurs Retrieval at 128k Context

Attention Entropy: Why Softmax Blurs Retrieval at 128k Context

Comments
7 min read
Digit Tokenization: Why a Comma Changes Your LLM's Math Answer

Digit Tokenization: Why a Comma Changes Your LLM's Math Answer

Comments
6 min read
Token Healing: Why a Trailing Space Breaks Your LLM Output

Token Healing: Why a Trailing Space Breaks Your LLM Output

Comments
7 min read
Grammar-Constrained Decoding: Why Valid JSON Gets Wrong Answers

Grammar-Constrained Decoding: Why Valid JSON Gets Wrong Answers

Comments
7 min read
KV Cache Preemption: Why vLLM Throughput Collapses Under Load

KV Cache Preemption: Why vLLM Throughput Collapses Under Load

Comments
7 min read
Filtered Vector Search: Why Metadata Filters Break HNSW Recall

Filtered Vector Search: Why Metadata Filters Break HNSW Recall

Comments
7 min read
Why Repetition Penalty Breaks JSON and Code Generation

Why Repetition Penalty Breaks JSON and Code Generation

Comments
7 min read
RoPE Scaling: Why Raising rope_theta Breaks Short Context

RoPE Scaling: Why Raising rope_theta Breaks Short Context

Comments
7 min read
Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Comments
7 min read
Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Comments
7 min read
Why temperature=0 Still Gives Different Answers: Batch Invariance

Why temperature=0 Still Gives Different Answers: Batch Invariance

Comments
7 min read
Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Comments
8 min read
Why JSON Schema Field Order Breaks Structured Output Accuracy

Why JSON Schema Field Order Breaks Structured Output Accuracy

Comments
7 min read
Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Comments 1
7 min read
DPO Likelihood Displacement: Why Chosen Responses Get Rarer

DPO Likelihood Displacement: Why Chosen Responses Get Rarer

Comments
7 min read
ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

Comments
7 min read
Surface Form Competition: Why Log-Prob Answer Scoring Fails

Surface Form Competition: Why Log-Prob Answer Scoring Fails

Comments
7 min read
Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Comments
8 min read
Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Comments
8 min read
MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

Comments
7 min read
Why I measure interview silence with WebAudio, not the Speech API

Why I measure interview silence with WebAudio, not the Speech API

1
Comments
5 min read
Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

1
Comments
7 min read
Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Comments
5 min read
Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Comments
7 min read
Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Comments
7 min read
LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

Comments
7 min read
Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Comments
7 min read
Binary Quantized Embeddings: 32x Smaller, If You Center First

Binary Quantized Embeddings: 32x Smaller, If You Center First

Comments
8 min read
MoE Capacity Factor: Why Mixture-of-Experts Drops Your Tokens

MoE Capacity Factor: Why Mixture-of-Experts Drops Your Tokens

Comments
8 min read
Hubness in Vector Search: Why One Chunk Tops Every RAG Query

Hubness in Vector Search: Why One Chunk Tops Every RAG Query

Comments
7 min read
Min-p Sampling: Why top_p Truncates the Wrong Tail

Min-p Sampling: Why top_p Truncates the Wrong Tail

Comments
6 min read
Chunked Prefill: Why One Long Prompt Stalls Every Decode

Chunked Prefill: Why One Long Prompt Stalls Every Decode

Comments
8 min read
Sequence Packing Leaks Across Documents Unless You Mask It

Sequence Packing Leaks Across Documents Unless You Mask It

Comments
6 min read
KV Cache Quantization: Why Keys and Values Need Different Axes

KV Cache Quantization: Why Keys and Values Need Different Axes

Comments
6 min read
Multi-LoRA Serving: Why 100 Adapters Fit on One GPU

Multi-LoRA Serving: Why 100 Adapters Fit on One GPU

Comments
7 min read
Token Healing: Why a Trailing Space Wrecks LLM Completions

Token Healing: Why a Trailing Space Wrecks LLM Completions

Comments
7 min read
Digit Tokenization: Why Commas Fix LLM Arithmetic

Digit Tokenization: Why Commas Fix LLM Arithmetic

Comments
8 min read
Matryoshka Embeddings: Truncate Vector Dimensions, Keep Recall

Matryoshka Embeddings: Truncate Vector Dimensions, Keep Recall

Comments
7 min read
Why Filtered Vector Search Quietly Destroys HNSW Recall

Why Filtered Vector Search Quietly Destroys HNSW Recall

Comments
7 min read
LLM-as-Judge Position Bias: Measure It Before You Ship

LLM-as-Judge Position Bias: Measure It Before You Ship

Comments
8 min read
YaRN vs NTK-Aware RoPE Scaling: Why Long Context Breaks

YaRN vs NTK-Aware RoPE Scaling: Why Long Context Breaks

Comments
7 min read
Prompt Caching: How One Dynamic Token Kills the 90% Discount

Prompt Caching: How One Dynamic Token Kills the 90% Discount

Comments
7 min read
Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Comments
6 min read
Why Temperature 0 Doesn't Make Your LLM Deterministic

Why Temperature 0 Doesn't Make Your LLM Deterministic

Comments
6 min read
Speculative Decoding: Why a Great Draft Model Still Caps Speedup

Speculative Decoding: Why a Great Draft Model Still Caps Speedup

Comments
6 min read
Constrained Decoding: Force Valid JSON Without Wrecking Accuracy

Constrained Decoding: Force Valid JSON Without Wrecking Accuracy

Comments
6 min read
GRPO Explained: Why DeepSeek Dropped the Critic in RLHF

GRPO Explained: Why DeepSeek Dropped the Critic in RLHF

Comments
6 min read
Online Softmax: How FlashAttention Skips the N N Matrix

Online Softmax: How FlashAttention Skips the N N Matrix

Comments
7 min read
Late Interaction Retrieval: Why ColBERT Beats Single-Vector RAG

Late Interaction Retrieval: Why ColBERT Beats Single-Vector RAG

Comments
7 min read
Why repetition_penalty Quietly Corrupts Your Code Generation

Why repetition_penalty Quietly Corrupts Your Code Generation

Comments
6 min read
DPO Likelihood Displacement: When Preferred Answers Get Rarer

DPO Likelihood Displacement: When Preferred Answers Get Rarer

Comments
6 min read
Why Token Logprobs Beat Asking Your LLM How Confident It Is

Why Token Logprobs Beat Asking Your LLM How Confident It Is

Comments
7 min read
loading...