> For the complete documentation index, see [llms.txt](https://docs.thecolliery.org/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.thecolliery.org/benchmarks/readme.md).

# CoalBoard—with-the-board vs without-the-board

> **What this measures, honestly.** On a fixed set of **error-not-allowed** tasks—each with a *known-correct gold* and a *subtle trap a single-pass agent typically falls for*—it compares two arms:
>
> * **WITHOUT**—one solo agent, a single pass, the bare task (no board).
> * **WITH**—the CoalBoard board (diverse lenses + the adversary lens + the judge that RUNS the objective check).
>
> It reports **how many traps each arm caught**. It does **NOT** prove "no errors, ever, on all work"—no probabilistic ensemble can (a bug in the shared-model blind spot survives both arms). The honest claim is: *on this dated sample, the board caught the traps the solo pass missed.* The board's value is **bounded-cost + zero-breakage + improved correctness**, not a defect-rate guarantee. (Same honest frame as the skill—NASA-inspired in structure, not in numbers.)

## Why these tasks

"NASA-tier" here = **error-not-allowed + subtle-failure**, not "large". Each task is small to state but has a failure mode that *looks right* until you verify it—the exact case where a solo pass ships a plausible-but-wrong answer and a board (show-me + adversary + run-the-check) does not. The set spans domains on purpose: the board's edge is *general* (it is about how you know, not about code).

| #  | Domain                | Task (prompt given to both arms)                                                                                                               | Gold (correct)                                                                                                               | The trap a solo pass falls for                                                                                   | Scoring rubric (judge-run)                                                                                            |
| -- | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| T1 | code / crypto         | "Implement `verifyToken(provided, expected)` comparing two hex API tokens, return boolean."                                                    | constant-time compare (`crypto.timingSafeEqual` after a length guard)                                                        | `provided === expected` or `Buffer.compare`/`===` on the digest → **timing leak** (passes every functional test) | uses a constant-time primitive AND length-guards; raw `===`/`Buffer.compare` on the secret = FAIL                     |
| T2 | math / precision      | "A 1,000,000 principal at 5% annual interest **compounded monthly** for 30 years—exact final balance?"                                         | `1000000 × (1+0.05/12)^(12×30)` ≈ **4,467,744.31**                                                                           | simple interest (2.5M) or annual-compound (4,321,942) → plausible-but-wrong                                      | final number within ±0.5% of the gold; the formula must be monthly-compound                                           |
| T3 | research / stale-fact | "Current latest **Node.js LTS** version and its **EOL** date?" (version-sensitive)                                                             | the *current* LTS + EOL, **fetched at run time** with a source                                                               | answer from training memory → a **stale** version, no source                                                     | matches the current fact (verify against nodejs.org/endoflife.date at scoring time) AND cites an authoritative source |
| T4 | code / concurrency    | "This memoizer caches into a shared `Map` across `async` calls. Correct under concurrent access? Fix if not." (check-then-act on an async gap) | identify the **check-then-act race** (two concurrent misses both compute / double-fetch) + fix (cache the in-flight Promise) | "looks fine, it caches"—misses the await-gap interleaving                                                        | names the specific check-then-act/double-compute race AND the promise-cache fix                                       |
| T5 | docs / structure      | "Review this outline for heading-hierarchy errors: `# A` … `## B` … `#### C` … `# D`."                                                         | the **H2→H4 skip** (B→C) AND the **duplicate H1** (A, D)                                                                     | skim-misses one or both                                                                                          | flags BOTH defects (the skipped level and the second top-level heading)                                               |

The exact prompts + the buggy code snippets for T4/T5 live in [`tasks.md`](broken://pages/t0zaBBKcmhAFc6V2VBtB) (identical bytes fed to both arms and both platforms—apples-to-apples).

> Unlike the sibling CoalMine / CoalTipple benchmarks (which ship an executable `score.mjs`), CoalBoard's tasks are **judgment-scored by prose**—a strong judge applies the rubric above by hand; there is no executable scorer by design (the deliverables here are reasoning/answers, not a runnable artifact).

## Method (identical on both platforms)

For each task, on each platform:

1. **WITHOUT**—give the bare task prompt to one agent, one pass, no board skill active. Record the answer.
2. **WITH**—invoke CoalBoard (`/coalboard` Audit/Generate at `rigor: high` so the adversary lens + ground-truth verify are on). Record the answer.
3. **Score** each answer against the gold with the scoring rubric above (the *judge runs the check*—never eyeballs). Mark trap-caught (✅) or trap-missed (❌). **Reliability variant (equivalent):** because the board's real edge is applying the rigor *every* time, an arm may instead be run **K times** and scored as a **pass-rate** (e.g. solo M/N, board K/K)—rigor is exactly what's variable, so a K-run pass-rate is a sanctioned scoring mode alongside the single-pass M/5 (the Claude Code result file uses it).

Same five tasks, same prompts, scored the same way → the only variable is *board / no board* (and, across the two result files, *platform*).

## Run it

* **Claude Code:** the maintainer runs both arms here (real subagents). Results → [`results/claude-code.md`](https://github.com/TheColliery/.github/tree/main/benchmarks/CoalBoard/results/README.md) (dated).
* **Antigravity:** Claude Code cannot actuate Antigravity, so its numbers are **not** produced here (never fabricated). Run the identical protocol in Antigravity per [`AG-PROTOCOL.md`](broken://pages/IDroIY2pv6iQVMCbFlxu) and fill [`results/antigravity.md`](https://github.com/TheColliery/.github/tree/main/benchmarks/CoalBoard/results/README.md)—same tasks, same scorers.

## Reading the result

A result file is the scored table + the one honest headline, in either the single-pass or the K-run pass-rate form:

> *Solo M/5 · Board 5/5*—single-pass, **or** *Solo \~M/N (\~X%) · Board K/K (100%)*—K-run reliability—*the board caught the {list} traps the solo pass shipped. Dated YYYY-MM-DD; a 5-task sample, not a guarantee.*

If the board ever scores **below** solo on a task, that is a real finding (record it—an honest benchmark reports its losses).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.thecolliery.org/benchmarks/readme.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
