Aceplore
Evalset — LLM Eval Sets, Scoring & Agreement
- Access details
- Secure checkout
- Product-specific terms
Evalset — LLM Eval Sets, Scoring & Agreement — a premium, 100% offline digital tool by Aceplore. Runs entirely in your browser: no accounts, no subscriptions, your data never leaves your device.
A private, offline scoring bench for model output — eval cases judged against a written rubric, a weighted score that prints its own arithmetic, a pass rate that excludes what nobody judged instead of calling it a failure, and raw agreement between two raters with the caveat attached — 31 tools in one, computed on your own machine.
31 built-in tools in one app. Suites that each own their cases, their criteria and their pass mark; a case written as one input plus what a passing answer must contain, with tags, a difficulty band, exact character, word and line counts and a length-derived token figure labelled an estimate wherever it appears; criteria with a weight, a scale top and a level descriptor for each number on that scale, and a weights view that gives each criterion its real share — weight over the actual total — and says so out loud when the weights do not sum to your target instead of rescaling them; runs with your own note on what was under test, a scoring screen taking one case, one rater and every criterion at a time, and a weighted-score view printing SUM(score × weight) ÷ SUM(weight) line by line; a pass rate with its denominator spelled out, averages by criterion, and pass rates per tag and per difficulty band; a run-to-run comparison matched by case id with separate regression and improvement lists; a failure log grouped into modes you name yourself; raters with their shared judgements and harshness, raw agreement over the judgements two people both made, and the disagreements themselves alongside the reasons each rater typed; plus a coverage matrix of every case × criterion × run cell, a CSV export for each table and a written report.
Weighted score. Each criterion becomes value ÷ scale × 100 before any weighting, which is the only way a 0–5 criterion and a pass/fail one can sit in the same total. On weights of 50, 30 and 20, scores of 4, 4 and 4 give a weighted 80%; 5, 3 and 4 give 84%. Every line of that working is printed under the case.
Pass rate. A case nobody judged is left out of the denominator and counted out loud — it is an unknown, not a failure, and a case with no scores at all gets no weighted score and no pass verdict. A case missing one criterion is divided by the weight actually scored, 80 rather than 100, giving 85% — and the unscored criterion is named rather than ignored.
Regressions. A real set comparison between two runs, matched by case id and counting only cases with a weighted score in both. A case missing a score in either run is excluded rather than being called a regression.
What's included:
- Evalset.html
- Quick-Start-Guide.html
- How-To-Use-Guide.html
- README.txt
- LICENSE.txt
Includes a personalized welcome wizard, an interactive guided tour, a ⌘K command palette, light/dark themes with 8 accents, and one-click JSON backup & restore.
Evalset is a planning, costing and record-keeping tool for people building with AI. Every figure it reports is computed by a stated formula from the rates, sizes, counts and thresholds YOU enter. It ships no model prices, no context windows and no tokeniser: those differ by provider and model and change without notice, so you enter them and the app shows which ones it used. Any token figure derived from the length of a text is an ESTIMATE, not a count — real tokenisers are model-specific and cannot run offline — so treat it as a planning figure and confirm actual usage against your provider's own reporting and billing before you commit to a budget. Nothing here is a guarantee of a cost, a saving, a quality level or a result, and none of it is business, financial, legal or engineering advice. You are solely responsible for your own decisions.
Each product ships with a clear license tier. Personal covers individual, non-commercial use; Commercial covers client and commercial work; Extended adds resale and SaaS rights. Pick your tier on this page.
Evalset — LLM Eval Sets, Scoring & Agreement
Usually bought with
Customers complete their toolkit with these — add the set in one click.
Customer reviews
Be the first to review
No reviews yet. Share your experience and help other shoppers choose with confidence.









