← All projects

18trees-benchmark-pages-skill

An AI skill that renders a benchmark's evaluation data, paper, and analysis report into a GitHub Pages homepage and leaderboard, with every number on the page recomputed from the CSV by script.

On this page[4]

What it is

18trees-benchmark-pages-skill renders a benchmark’s evaluation data, a finished paper, and an analysis report into a homepage and leaderboard that can be hosted on GitHub Pages: the homepage explains what the benchmark is, and the leaderboard lets people compare models side by side.

It addresses a problem that gets rebuilt over and over. Everyone building a benchmark forks the same academic project page template and then hand-rolls the benchmark-specific parts — how tables are generated from experiment results, how the leaderboard changes when a model is updated, how someone else submits their results. The template has none of it, so every team solves the same problems again. This skill does those parts.

Its boundaries are stated plainly: it does not run evaluations, it does not do data analysis, and it does not do visual design — it only consumes finished analysis artifacts, and when data is missing it fails and asks for it rather than interpolating or estimating. It therefore needs an environment that can read and write files and run commands — Claude Code, Codex, or any agent that can execute commands. A web-only chatbot cannot run it: there the model would have to invent an HTML page from memory, which defeats the point, because every number on the page is recomputed from CSV.

How it is structured

18trees-benchmark-pages-skill Architecture A architecture diagram generated by Archify. Benchmark inputs Build & publish Skill package Raw Results · experiments/ + CSV · Benchmark inputs Raw Results experiments/ + CSV Paper & Report · abstract · figures · Benchmark inputs Paper & Report abstract · figures Site Source · site.yaml + data/ · Build & publish Site Source site.yaml + data/ Skill Rules · SKILL.md + references · Skill package Skill Rules SKILL.md + references Generator · build_site.py · Build & publish Generator build_site.py Page Template · paper page layout · Skill package Page Template paper page layout Static Site · index + leaderboard · Build & publish Static Site index + leaderboard Table Engine · Tabulator, vendored · Skill package Table Engine Tabulator, vendored GitHub Pages · published site · Build & publish GitHub Pages published site CI Workflow · pages.yml · Architecture component CI Workflow pages.yml CSV paper guides site.yaml layout table renders deploy rebuilds Legend Frontend Backend Database Cloud

The project has four layers; the inputs come from outside, and the other three live in the repository.

Input layer. Three things: raw evaluation data (experiments/<run>/*.jsonl plus the canonical per-item wide table), the paper (title, authors, abstract, and BibTeX), and the analysis report (REPORT.md and the figures in figures/). The per-item table matters most — one row per question, one column per model — because it is what lets every number on the leaderboard be recomputed. The diagram draws the paper and the analysis report in a single box, since both are narrative sources.

Rules layer. skills/benchmark-pages/SKILL.md is the single source of truth for the rules, and four documents under references/ cover the input contract, the leaderboard column spec, the site fields, and deployment with result submission. The rules fix the work into six phases: inventory the inputs, lock the evaluation terms, build the three tables, write site.yaml, render and self-check, then verify locally and deploy.

Source directory and generator. site.yaml is the only file a person writes by hand; data/leaderboard.csv holds one row per entrant, data/breakdown_*.csv splits the results by dimension, data/items.csv is the optional per-item detail, and data/metrics.md records the evaluation terms. scripts/build_site.py reads that source directory and, with the layout in assets/template/ and the vendored Tabulator table engine, renders _site/ (index.html, leaderboard.html, and static/). The whole chain has one Python dependency.

Delivery layer. The generated static site goes straight to GitHub Pages. The site/ directory in the repository is a real example — six models over 2,492 four-choice questions — rebuilt automatically by .github/workflows/pages.yml whenever the data changes; the leaderboard on that page is not a screenshot.

Design decisions worth noting

Numbers are never hand-written into HTML. Every number on the page comes from data/*.csv and is rendered by build_site.py; to change a number, change the CSV and re-run. This governs the mechanism rather than the content — it does not audit whether your data is right, only that the page is produced from the data and can be rebuilt at any time.

Structural problems block, data problems only warn. Missing columns or a referenced file that does not exist — anything that would keep the page from rendering — fails the build with no half-finished output. Plausibility questions such as correct / n disagreeing with the primary metric, or uneven denominators across entrants, only print a warning and the page is still produced; opt into strictness with --strict.

The upstream backlinks are a hard check. The page layout derives from the Academic Project Page Template and Nerfies (both CC BY-SA 4.0), so the generator uses REQUIRED_BACKLINKS to check that every generated site’s footer keeps those links, and there is no off switch — the licence obligation lives in the code rather than in good intentions.

One switch per section. Sorting and filtering, per-dimension heatmaps, confidence intervals, the random baseline anchor, the evaluation-protocol footnote, and the submission guide are each enabled in site.yaml; whatever is not configured simply does not render, so no empty boxes are left behind. The heatmap colour scales are normalised per column, because difficulty differs from column to column.

A passing generator does not mean a working page. The repository makes “open the generated site in a browser and click through every page” a hard requirement: the table renders, clicking a header sorts, narrow viewports do not overflow horizontally, and the model column stays leftmost and visible. This step is mandatory after any change to the template or the generator.

What problems it solves

  • Numbers are no longer copied into HTML by hand: change the CSV, re-run once, and the page follows while still matching the CSV.
  • Adding a model means adding a CSV row rather than editing HTML, and the tables and charts cannot drift apart.
  • The leaderboard already carries sorting and filtering, per-dimension heatmaps, confidence intervals, a random baseline anchor, the evaluation-protocol footnote, and a submission guide, so none of it has to be written from scratch.
  • “Is 75.6% good or bad?” has an answer — the random baseline anchor puts 25.0% and the best score side by side. So does “how were invalid answers counted?” — the terms come straight from data/metrics.md into the page footnote.
  • Generated sites are self-contained: one Python dependency, the table engine shipped with the site, opening offline with no CDN involved.
  • Structural problems are caught during the build, so no half-finished page is produced — which is how the repository’s own workflow uses it.
0