{"product_id":"evalset-llm-scoring","title":"Evalset — LLM Eval Sets, Scoring \u0026 Agreement","description":"\u003cp\u003e\u003cstrong\u003eEvalset — LLM Eval Sets, Scoring \u0026amp; Agreement\u003c\/strong\u003e — a premium, 100% offline digital tool by Aceplore. Runs entirely in your browser: no accounts, no subscriptions, your data never leaves your device.\u003c\/p\u003e\n\u003cp\u003eA private, offline scoring bench for model output — eval cases judged against a written rubric, a weighted score that prints its own arithmetic, a pass rate that excludes what nobody judged instead of calling it a failure, and raw agreement between two raters with the caveat attached — 31 tools in one, computed on your own machine.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003e31 built-in tools in one app.\u003c\/strong\u003e Suites that each own their cases, their criteria and their pass mark; a case written as one input plus what a passing answer must contain, with tags, a difficulty band, exact character, word and line counts and a length-derived token figure labelled an estimate wherever it appears; criteria with a weight, a scale top and a level descriptor for each number on that scale, and a weights view that gives each criterion its real share — weight over the actual total — and says so out loud when the weights do not sum to your target instead of rescaling them; runs with your own note on what was under test, a scoring screen taking one case, one rater and every criterion at a time, and a weighted-score view printing SUM(score × weight) ÷ SUM(weight) line by line; a pass rate with its denominator spelled out, averages by criterion, and pass rates per tag and per difficulty band; a run-to-run comparison matched by case id with separate regression and improvement lists; a failure log grouped into modes you name yourself; raters with their shared judgements and harshness, raw agreement over the judgements two people both made, and the disagreements themselves alongside the reasons each rater typed; plus a coverage matrix of every case × criterion × run cell, a CSV export for each table and a written report.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eWeighted score.\u003c\/strong\u003e Each criterion becomes value ÷ scale × 100 before any weighting, which is the only way a 0–5 criterion and a pass\/fail one can sit in the same total. On weights of 50, 30 and 20, scores of 4, 4 and 4 give a weighted 80%; 5, 3 and 4 give 84%. Every line of that working is printed under the case.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003ePass rate.\u003c\/strong\u003e A case nobody judged is left out of the denominator and counted out loud — it is an unknown, not a failure, and a case with no scores at all gets no weighted score and no pass verdict. A case missing one criterion is divided by the weight actually scored, 80 rather than 100, giving 85% — and the unscored criterion is named rather than ignored.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eRegressions.\u003c\/strong\u003e A real set comparison between two runs, matched by case id and counting only cases with a weighted score in both. A case missing a score in either run is excluded rather than being called a regression.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eWhat's included:\u003c\/strong\u003e\u003c\/p\u003e\n\u003cul\u003e\n\u003cli\u003eEvalset.html\u003c\/li\u003e\n\u003cli\u003eQuick-Start-Guide.html\u003c\/li\u003e\n\u003cli\u003eHow-To-Use-Guide.html\u003c\/li\u003e\n\u003cli\u003eREADME.txt\u003c\/li\u003e\n\u003cli\u003eLICENSE.txt\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eIncludes a personalized welcome wizard, an interactive guided tour, a ⌘K command palette, light\/dark themes with 8 accents, and one-click JSON backup \u0026amp; restore.\u003c\/p\u003e\n\u003cp\u003e\u003cem\u003eEvalset is a planning, costing and record-keeping tool for people building with AI. Every figure it reports is computed by a stated formula from the rates, sizes, counts and thresholds YOU enter. It ships no model prices, no context windows and no tokeniser: those differ by provider and model and change without notice, so you enter them and the app shows which ones it used. Any token figure derived from the length of a text is an ESTIMATE, not a count — real tokenisers are model-specific and cannot run offline — so treat it as a planning figure and confirm actual usage against your provider's own reporting and billing before you commit to a budget. Nothing here is a guarantee of a cost, a saving, a quality level or a result, and none of it is business, financial, legal or engineering advice. You are solely responsible for your own decisions.\u003c\/em\u003e\u003c\/p\u003e","brand":"Aceplore","offers":[{"title":"Personal","offer_id":51816848195724,"sku":"EVALSETPE","price":39.0,"currency_code":"CAD","in_stock":true},{"title":"Commercial","offer_id":51816848228492,"sku":"EVALSETCO","price":79.0,"currency_code":"CAD","in_stock":true},{"title":"Extended","offer_id":51816848261260,"sku":"EVALSETEX","price":149.0,"currency_code":"CAD","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0779\/3763\/9564\/files\/ev-00.png?v=1785800902","url":"https:\/\/aceplore.com\/products\/evalset-llm-scoring","provider":"Aceplore","version":"1.0","type":"link"}