Docsloth
Self-hostStart a preview

docsloth.dev

Benchmarks

This page publishes the evaluation method and the exact task definitions. Results appear here only when the benchmark has actually run with the stated models, budgets and seeds. No superiority claim is made before that, and overlapping intervals never establish one afterwards.

Actions

Actions

  • Run harness

    Jumps to the exact setup below: the no-credential harness verification (free, local, mechanics only) and the live evaluation (owner model key, bounded provider spend).

  • Download task definitions

    Downloads the real JSON task sets served from this site: the exact prompts, checker expressions, seeds and source script paths for both arms.

Method

  1. Same base model for both arms, plus the same source revision, tool rights and comparable budgets.
  2. Independent scoring: task success is decided by a harness that runs the produced artifact, never by the model grading itself.
  3. Multiple seeds per arm; means with 95% intervals; latency, spend and model versions reported.
  4. Frozen fixtures so runs are reproducible, and the full evaluation code is part of the public engine.

Run it

Run the harness verification — no credentials, no cost

This runs both arms against a local deterministic model stub. It verifies the statistics and fairness rules (unequal budgets or models are refused; overlapping intervals stay inconclusive) and writes nothing outside the process. It says nothing about model quality, and it never contacts a provider. The published evidence/benchmark-harness.json is exactly this output.

node docsloth-core/scripts/benchmark-harness.mjs

Cost: none. Permissions: a checkout of the public engine and Node 22+. Cases: fix-import, add-guide; seeds 1, 2, 3, 4, 5.

Run the live evaluation — owner model credentials required

The live benchmark sends the same prompts to your model endpoint for both arms and scores them with the independent checker. Without credentials it prints not_configured and exits without spending anything.

DOCSLOTH_MODEL_BASE_URL=<endpoint> \
DOCSLOTH_MODEL_API_KEY=<key> \
DOCSLOTH_MODEL_IDS=<approved-model-id> \
node scripts/benchmark-live.mjs
  • Cost: real provider spend through your key, bounded by BENCHMARK_BUDGET_MICRO_USD (default $5) with BENCHMARK_ARM_BUDGET_MICRO_USD (default $0.50) per call; Docsloth charges nothing for the run.
  • Permission: the operator supplies the credentials and approves publication; the script writes to BENCHMARK_OUT (default build-artifacts/benchmark-live.json).
  • Cases: runnable-example, failure-mode, breaking-change; seeds from BENCHMARK_SEEDS (default 1, 2, 3).
  • Results are published on this page only after that run; nothing is pre-filled or estimated.

Current results

Not yet run with live provider credentials. The harness is implemented and smoke-verified against a local deterministic model stub, which checks the statistics and fairness rules (unequal budgets or models are refused; overlapping intervals stay inconclusive) but says nothing about model quality. When the owner supplies credentials, results for the stated model versions will be published here with spend and latency.

What we will never publish

  • A benchmark where the open-source arm uses a weaker prompt than the managed arm.
  • Best-of-N runs presented as a single-seed result.
  • Benchmarks against models we cannot name, or with budgets we do not state.