← All projects
decidebench
choyiny/decidebench MIT Python Eval & Align
rank#293
Benchmark of decision models (JEV and its open alternatives) and LLMs: accuracy, cost per task and latency on 400 contrastive decisions
- Stars
- 1
- Forks
- 0
- Open issues
- 0
- Last push
- today
Star history starts building from the first daily refresh.