Evidence register / updated 28 September 2026

The evidence, including what is missing.

A measured local execution smoke test, published assumptions, and reproducible arithmetic are available now. Comparative execution benchmarks and production-fleet results are not yet available.

No measured speedup, customer savings, uptime SLA, certification, or production runner qualification is claimed on this page.
Evidence classification — 28 September 2026
ItemClassificationWhat it establishes
Break-even inequalityAnalyticalThe cost condition under explicitly stated assumptions; not measured savings.
Interactive calculatorIllustrative modelScenario arithmetic that changes when inputs change.
Provider comparisonDocumentary researchPublished capabilities and rates, with primary-source links.
Runtime benchmarksNot measuredNo comparative speed or tail-latency conclusion.
Customer production outcomesNot availableNo customer ROI or reliability claim.
Security qualificationNot completedDesign requirements are not a penetration test or an attestation.

Measured / local execution smoke test

A real task.
Three observed outcomes.

On 28 September 2026, we ran infinity/dedupe-orders through Harbor and Docker on an operator-owned machine. This is an internal TB3-profile execution fixture, not an official Terminal-Bench benchmark score.

The run was invoked locally through Harbor. It was not dispatched by this website, and it does not establish an operational Actions worker service, customer savings, or production security.
One trial per control · Harbor 0.18.0 · separate verifier configured without network access
ControlRewardVerifier testsFresh-input checksElapsed time
Reference solution1.09/9 passed3/3 passed179.26 s
No-op control0.01/9 passed0/3 passed55.67 s
Wrong ceiling (5,001)0.06/9 passed0/3 passed51.82 s

Zero trial errors. The negative controls are expected to fail the task: no work earned no reward, and changing the ceiling from 5,000 to 5,001 was detected. Each task environment was configured for one CPU and 1 GiB RAM. Timings include preparation and cache effects; they are not a provider comparison. No paid model calls were made, but local hardware, electricity, and operator time still have costs.

What passed

47/47 static checks, the reference solution’s nine verifier tests, and three freshly generated witness inputs.

What remains open

Traceability linked 16 of 17 requirements, with two review findings for the helper-directory requirement. The full assurance pipeline, model-backed stages, worker enrollment, and website-to-worker dispatch were not tested.

Self-reported development evidence, not independent attestation or a certificate. The source task was unchanged; the mutant ran from a separate staged copy. The downloadable record includes the source commit, task checksums, run IDs, and hashes of the locally retained raw results.

Scenario 001 / not a benchmark

A useful example.
Not a performance claim.

20,000 jobs. Assumed runtimes: six minutes on GitHub, three hosted, four local. 80% completed locally, $100 all-in local cost, and a $0 placeholder platform fee.

ScenarioMonthly modeled costBasis
GitHub-hosted$72020,000 × 6 × $0.006
Hosted alternative at Blacksmith’s listed rate$24020,000 × 3 × $0.004
Hybrid at $100 local cost$1484,000 × 3 × $0.004 + $100
Hybrid at $200 local cost$248More expensive than the hosted baseline

Provider rates are published inputs; runtimes and placement share are hypothetical. Free allowances, hosted storage, taxes, and unmodeled transfer/retry costs are excluded. A real platform quote must replace the $0 placeholder.

Primary sources

Read the original claims.

Reviewed 28 September 2026. Vendor statements are attributed, not independently verified benchmarks. Rates and capabilities can change.

GitHub Actions runner pricing(opens in a new tab)

Published Linux x64 2-core rate: $0.006/minute. GitHub rounds each job up to a whole minute.

GitHub self-hosted runners(opens in a new tab)

Self-hosted runner usage is currently free; infrastructure and associated services remain the customer's responsibility.

Blacksmith pricing(opens in a new tab)

Listed Linux x64 2-vCPU rate: $0.004/minute. Its headline savings assume faster execution; we have not independently measured that speedup.

Blacksmith Sticky Disks(opens in a new tab)

Persistent snapshot-backed build state, with branch-protection controls for cache writes.

Actuated(opens in a new tab)

Managed CI on customer infrastructure using disposable microVMs. A direct alternative to evaluate.

RunsOn pricing(opens in a new tab)

Customer-owned AWS infrastructure with an annual software licence; cloud infrastructure is billed separately.

Harbor framework(opens in a new tab)

Agent evaluation tasks, trials, verifiers and sandbox adapters. Different from the CNCF Harbor registry.

GitHub runner security(opens in a new tab)

Runner code can access referenced credentials. Log redaction is not a security boundary.

Firecracker requirements(opens in a new tab)

Linux/KVM requirements constrain where this microVM backend can run.

Docker resource constraints(opens in a new tab)

Containers need explicitly configured resource limits; defaults do not reserve the host's resources.

Start with your actual workload.

Explore the model, then apply for a scoped pilot. No card or compute commitment.

Request client access