Jev Users
← All projects

decidebench

choyiny/decidebench MIT Python Eval & Align

rank#293
What it does

Benchmark of decision models (JEV and its open alternatives) and LLMs: accuracy, cost per task and latency on 400 contrastive decisions

Numbers
Stars
1
Forks
0
Open issues
0
Last push
today

Star history starts building from the first daily refresh.

Open on GitHub → Homepage

More in Eval & Align