Address
:
[go:
up one dir
,
main page
]
Include Form
Remove Scripts
Accept Cookies
Show Images
Show Referer
Rotate13
Base64
Strip Meta
Strip Title
Session Cookies
Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
#
rag
#
llm
#
embeddings
#
evaluation
Comments
1
 comment
7 min read
Your New Eval Rule Is Untested Code Guarding Production
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 2
Your New Eval Rule Is Untested Code Guarding Production
#
ai
#
agents
#
evaluation
#
testing
Comments
1
 comment
5 min read
OpenEval: Why LLM Evaluation Needs a Standard Format
Adha AK
Adha AK
Adha AK
Follow
Jul 30
OpenEval: Why LLM Evaluation Needs a Standard Format
#
llm
#
evaluation
#
ai
#
testing
Comments
Add Comment
1 min read
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 26
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through
#
ai
#
agents
#
evaluation
#
observability
4
 reactions
Comments
2
 comments
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
 reaction
Comments
Add Comment
6 min read
Your Agent's Confidence Score Is Not a Probability
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 29
Your Agent's Confidence Score Is Not a Probability
#
ai
#
agents
#
evaluation
#
observability
3
 reactions
Comments
1
 comment
4 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
#
llms
#
codegeneration
#
selfrepair
#
evaluation
Comments
Add Comment
3 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
Muhammed Rasin O M
Muhammed Rasin O M
Muhammed Rasin O M
Follow
Jul 10
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
#
evaluation
#
dataagents
#
benchmarks
#
syntheticdata
Comments
Add Comment
7 min read
Your Agent's Deadline Is a Correctness Test, Not an SLO
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 31
Your Agent's Deadline Is a Correctness Test, Not an SLO
#
ai
#
agents
#
evaluation
#
observability
3
 reactions
Comments
1
 comment
4 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 19
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
#
ai
#
agents
#
evaluation
#
observability
2
 reactions
Comments
3
 comments
5 min read
Evaluating LLM Apps in Python
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Python
#
python
#
ai
#
llm
#
evaluation
Comments
Add Comment
9 min read
Evaluating LLM Apps in Java
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Java
#
java
#
ai
#
llm
#
evaluation
Comments
Add Comment
10 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 2
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
#
ai
#
agents
#
evaluation
#
typescript
1
 reaction
Comments
Add Comment
5 min read
Your AI judge might be reliable — and still be wrong
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Your AI judge might be reliable — and still be wrong
#
evaluation
#
llmjudges
#
rlhf
#
methodology
Comments
Add Comment
3 min read
Reliable, and still wrong
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Reliable, and still wrong
#
evaluation
#
llmasjudge
#
benchmarks
Comments
Add Comment
3 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account