Jev Users
← All projects

jev-vs-llm-tb-amr

Jiadalee/jev-vs-llm-tb-amr MIT Python Eval & Align

rank#310
What it does

Benchmark: a System One decision model (Jev) vs a generative LLM (Qwen3.8-27B) answering the same 27 TB drug-resistance questions per isolate — 98% agreement, 178–392× faster. Catalog-grounded AMR prediction with judgment = model, facts = code

Numbers
Stars
0
Forks
0
Open issues
0
Last push
today

Star history starts building from the first daily refresh.

Open on GitHub →

More in Eval & Align