DEV Community

#evals

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I Published a Perfect Recall. Then I Measured Dust.

I Published a Perfect Recall. Then I Measured Dust.

Comments
7 min read
OpenAI's dots need acceptance criteria, not just Custom Rules

OpenAI's dots need acceptance criteria, not just Custom Rules

1
Comments
6 min read
How to Evaluate AI Agents: Reliability, Cost, Latency, and Failure Modes

How to Evaluate AI Agents: Reliability, Cost, Latency, and Failure Modes

Comments
7 min read
Evals are the unit tests of prompts

Evals are the unit tests of prompts

Comments
2 min read
TypeSafe Jev Played Chess — And Landed Next to Reasoning Models

TypeSafe Jev Played Chess — And Landed Next to Reasoning Models

13
Comments 1
3 min read
One pass of my eval bills $9.14 on the API and $0 through the CLI

One pass of my eval bills $9.14 on the API and $0 through the CLI

Comments
2 min read
AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)

AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)

Comments
4 min read
I have been Vibecoding Evals (works better than I thought)

I have been Vibecoding Evals (works better than I thought)

Comments
3 min read
How to Design AI Evaluations You Can Actually Trust

How to Design AI Evaluations You Can Actually Trust

46
Picked as gem Comments 11
4 min read
Evals: How You Know the Agent Works in Production

Evals: How You Know the Agent Works in Production

Comments
3 min read
How to Write Reliable Rubrics for LLM-as-a-Judge Evaluations

How to Write Reliable Rubrics for LLM-as-a-Judge Evaluations

33
Comments 2
5 min read
Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once

Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once

Comments
6 min read
Rewriting research prose until the tests pass

Rewriting research prose until the tests pass

Comments
5 min read
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade

The 12-Prompt Eval I Run Before I Trust Any Model Upgrade

Comments
3 min read
How to Build AI Evals for Tool-Calling Agents

How to Build AI Evals for Tool-Calling Agents

1
Comments 2
17 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.