DEV Community

#evaluation

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Green Scores Without a Pinned Harness Are Screenshots

Green Scores Without a Pinned Harness Are Screenshots

1
Comments
4 min read
AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly

AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly

1
Comments 1
4 min read
Your Agent Has Observability. It Doesn't Have Evals.

Your Agent Has Observability. It Doesn't Have Evals.

Comments
10 min read
LLM Evaluation: How a Benchmark Produces Comparable Numbers

LLM Evaluation: How a Benchmark Produces Comparable Numbers

Comments
4 min read
How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

1
Comments
4 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Comments 2
5 min read
An LLM reviewer's "block" is a feature, not a verdict

An LLM reviewer's "block" is a feature, not a verdict

Comments 1
4 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

Comments
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

Comments
4 min read
Using Execution Traces to Evaluate AI Agent Behavior

Using Execution Traces to Evaluate AI Agent Behavior

2
Comments
5 min read
Cheap LLM code review is fine until it hits an authorization bug

Cheap LLM code review is fine until it hits an authorization bug

Comments
2 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

24
Comments 5
5 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
The model did the reverse-engineering. The validator was the hard part.

The model did the reverse-engineering. The validator was the hard part.

Comments 1
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.