← All projects
jev-lab
llt22/jev-lab Python Eval & Align
rank#257
Hands-on research lab for TypeSafe's Jev (System One model): reproducible benchmarks of Noul/Choice/Score primitives, confidence gating, fan-out latency, agent control — plus a living audit of the Jev ecosystem.
- Stars
- 1
- Forks
- 0
- Open issues
- 0
- Last push
- today
Star history starts building from the first daily refresh.
ai-agentsbenchmarkconfidence-calibrationevalsjevllmllm-evaluationllm-observabilitymodel-evaluationpythonresearchstructured-outputs