[go: up one dir, main page]

DEV Community

#evaluation

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 1
7 min read
Your New Eval Rule Is Untested Code Guarding Production

Your New Eval Rule Is Untested Code Guarding Production

Comments 1
5 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through

Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through

4
Comments 2
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
Your Agent's Confidence Score Is Not a Probability

Your Agent's Confidence Score Is Not a Probability

3
Comments 1
4 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

Comments
3 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
Your Agent's Deadline Is a Correctness Test, Not an SLO

Your Agent's Deadline Is a Correctness Test, Not an SLO

3
Comments 1
4 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

2
Comments 3
5 min read
Evaluating LLM Apps in Python

Evaluating LLM Apps in Python

Comments
9 min read
Evaluating LLM Apps in Java

Evaluating LLM Apps in Java

Comments
10 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

1
Comments
5 min read
Your AI judge might be reliable — and still be wrong

Your AI judge might be reliable — and still be wrong

Comments
3 min read
Reliable, and still wrong

Reliable, and still wrong

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.