A small, model-agnostic benchmark for small (<1B parameter) language models: tool-use (single-turn and an 8-turn natural-phrasing conversation) plus ARC-Easy and TruthfulQA MC1. No model — including the Q Project's own — gets forced-correct tool execution; every score reflects what the model actually generates on its own. Run it yourself and submit a result: see the GitHub repo.
Loading leaderboard…
Data: q-project/qbench-results on the Hub, updated whenever someone submits a run via the CLI.