← All projects
Your-language-model-is-already-a-decision-model
ntlm1686/Your-language-model-is-already-a-decision-model Python Eval & Align
rank#233
We compared unfinetuned Qwen with Jev on decision accuracy, calibration, and latency, and found comparable accuracy and lower average calibration error without decision-specific fine-tuning.
- Stars
- 6
- Forks
- 0
- Open issues
- 0
- Last push
- today
Star history starts building from the first daily refresh.
agentaiai-agentsbenchmarkcalibrationjevjev-modelllmqwen