Evidence register / updated 28 September 2026
The evidence, including what is missing.
A measured local execution smoke test, published assumptions, and reproducible arithmetic are available now. Comparative execution benchmarks and production-fleet results are not yet available.
| Item | Classification | What it establishes |
|---|---|---|
| Break-even inequality | Analytical | The cost condition under explicitly stated assumptions; not measured savings. |
| Interactive calculator | Illustrative model | Scenario arithmetic that changes when inputs change. |
| Provider comparison | Documentary research | Published capabilities and rates, with primary-source links. |
| Runtime benchmarks | Not measured | No comparative speed or tail-latency conclusion. |
| Customer production outcomes | Not available | No customer ROI or reliability claim. |
| Security qualification | Not completed | Design requirements are not a penetration test or an attestation. |
Measured / local execution smoke test
A real task.
Three observed outcomes.
On 28 September 2026, we ran infinity/dedupe-orders through Harbor and Docker on an operator-owned machine. This is an internal TB3-profile execution fixture, not an official Terminal-Bench benchmark score.
| Control | Reward | Verifier tests | Fresh-input checks | Elapsed time |
|---|---|---|---|---|
| Reference solution | 1.0 | 9/9 passed | 3/3 passed | 179.26 s |
| No-op control | 0.0 | 1/9 passed | 0/3 passed | 55.67 s |
| Wrong ceiling (5,001) | 0.0 | 6/9 passed | 0/3 passed | 51.82 s |
Zero trial errors. The negative controls are expected to fail the task: no work earned no reward, and changing the ceiling from 5,000 to 5,001 was detected. Each task environment was configured for one CPU and 1 GiB RAM. Timings include preparation and cache effects; they are not a provider comparison. No paid model calls were made, but local hardware, electricity, and operator time still have costs.
What passed
47/47 static checks, the reference solution’s nine verifier tests, and three freshly generated witness inputs.
What remains open
Traceability linked 16 of 17 requirements, with two review findings for the helper-directory requirement. The full assurance pipeline, model-backed stages, worker enrollment, and website-to-worker dispatch were not tested.
Self-reported development evidence, not independent attestation or a certificate. The source task was unchanged; the mutant ran from a separate staged copy. The downloadable record includes the source commit, task checksums, run IDs, and hashes of the locally retained raw results.
Scenario 001 / not a benchmark
A useful example.
Not a performance claim.
20,000 jobs. Assumed runtimes: six minutes on GitHub, three hosted, four local. 80% completed locally, $100 all-in local cost, and a $0 placeholder platform fee.
| Scenario | Monthly modeled cost | Basis |
|---|---|---|
| GitHub-hosted | $720 | 20,000 × 6 × $0.006 |
| Hosted alternative at Blacksmith’s listed rate | $240 | 20,000 × 3 × $0.004 |
| Hybrid at $100 local cost | $148 | 4,000 × 3 × $0.004 + $100 |
| Hybrid at $200 local cost | $248 | More expensive than the hosted baseline |
Provider rates are published inputs; runtimes and placement share are hypothetical. Free allowances, hosted storage, taxes, and unmodeled transfer/retry costs are excluded. A real platform quote must replace the $0 placeholder.
Primary sources
Read the original claims.
Reviewed 28 September 2026. Vendor statements are attributed, not independently verified benchmarks. Rates and capabilities can change.
Published Linux x64 2-core rate: $0.006/minute. GitHub rounds each job up to a whole minute.
Self-hosted runner usage is currently free; infrastructure and associated services remain the customer's responsibility.
Listed Linux x64 2-vCPU rate: $0.004/minute. Its headline savings assume faster execution; we have not independently measured that speedup.
Persistent snapshot-backed build state, with branch-protection controls for cache writes.
Managed CI on customer infrastructure using disposable microVMs. A direct alternative to evaluate.
Customer-owned AWS infrastructure with an annual software licence; cloud infrastructure is billed separately.
Agent evaluation tasks, trials, verifiers and sandbox adapters. Different from the CNCF Harbor registry.
Runner code can access referenced credentials. Log redaction is not a security boundary.
Linux/KVM requirements constrain where this microVM backend can run.
Containers need explicitly configured resource limits; defaults do not reserve the host's resources.
Start with your actual workload.
Explore the model, then apply for a scoped pilot. No card or compute commitment.