Prospective protocol / version 1.0

Measure the whole pipeline. Keep the receipts.

A fair comparison includes queue time, setup, execution, retries, storage, and human operations. The fastest successful sample is not the result.

Start with a comparable baseline.

Use customer-authorized repositories and immutable source commits. Record workflow revision, runner image digest, dependency lockfiles, CPU model, architecture, RAM, storage, region, network, and concurrency. Keep commands and correctness checks the same across providers.

Evaluate GitHub-hosted runners, Blacksmith, an unmanaged self-hosted baseline, and the candidate hybrid policy where eligible. Add Actuated or another managed customer-owned provider when relevant. Match resource classes, but also report hardware differences—two advertised vCPUs need not have equal performance.

Exercise both cache and capacity.

DimensionPlanned cases
Build familiesNode monorepo, Python tests, compiled-language build, Docker image build
Cache stateCold; unchanged warm run; source edit; dependency/toolchain change
Arrival patternSingle job; recorded business-hour arrivals; controlled peak burst
Host conditionDedicated idle worker; competing load; insufficient local capacity
Trust classApproved internal build; untrusted pull request; privileged release

Plan at least 30 repetitions per applicable build/cache/provider cell for an initial study. Randomize order, separate warm-up runs, report variation and confidence intervals, and collect substantially more samples for stable tail estimates. The initial sample size is not a guarantee of statistical power.

Make the bill and the wait visible.

  • Trigger-to-result time, with queue, provision, setup, and execution broken out.
  • Median and p95 completion time, with sample counts and uncertainty.
  • Total billed cost and cost per successful pipeline, including failed attempts.
  • Cache hit rate, transferred bytes, cache storage, and eviction behavior.
  • Local utilization, incremental power, and maintenance hours.
  • Developer interference and time lost to workstation contention.

Report credits and negotiated discounts separately. Show both marginal cost on existing hardware and fully allocated ownership cost. Do not add multiple savings percentages that describe the same avoided work.

A sleeping laptop is part of the experiment.

Deliberately test worker loss, network interruption, stale leases, low disk space, cache corruption, expired credentials, and exhausted cloud budgets. Confirm that jobs do not become successful merely because a required check never ran.

Keep publication and deployment tests in isolated test destinations. Verify duplicate-execution protection and explicit retry rules before allowing side-effecting workloads. Never run untrusted public code on personal machines merely to increase sample volume.

Publish an inspectable result package.

  1. Immutable workload manifest and exact commands.
  2. Hardware, runtime, provider rates, and billing rules.
  3. Raw timing records, failures, retries, and exclusions.
  4. Cache state, source hashes, and output correctness checks.
  5. Cost calculations and analysis code.
  6. Known limitations, negative results, and reproduction instructions.
Client source, credentials, personal information, and proprietary logs require permission and redaction before publication. Signing a record proves who signed it, not that an untrusted machine executed it correctly.

Set success criteria before seeing results.

A pilot should agree on a cost-reduction target, maximum completion-time regression, reliability floor, and allowed operator effort before measurement. The study should report whether all criteria were met, including cases where an existing hosted provider remains the better choice.

See what has actually been published

Start with your actual workload.

Explore the model, then apply for a scoped pilot. No card or compute commitment.

Request client access