Actions / Documentation

When a device or job stops

Tbench tends to fail loudly rather than guess. The worker prints a plain reason and refuses to claim work it cannot run, which is why most of this page is “read the message, then do the matching thing”.

← All documentation

Start with the diagnostics

Before changing anything, run the device’s own check. It verifies the four things a job depends on — device caps, the pinned Ubuntu 24.04 image, the Docker version, and broker enrollment — and names the eligible labels it found.

tbench-worker check

A pass prints the verified labels. A failure prints the specific refusal, which is usually one of the messages below.

The device won’t start

Preflight runs before anything is authorized, created or leased, so a refusal here means nothing was claimed. Match the message:

  • “requires an x86-64 Linux Docker engine” — Docker is not in Linux-container mode. On Windows, switch Docker Desktop to Linux containers (WSL2). Windows job images are not supported.
  • “Docker Engine 28 or newer is required” — upgrade the engine. The qualified isolated network mode needs it.
  • “Runner image OS is not the advertised Ubuntu 24.04” — the pinned image is missing or wrong. Pull it again and re-run the check.
  • “No runner class fits the opted-in device caps” — the configured caps are smaller than any label. Raise TBENCH_MAX_CPUS and TBENCH_MAX_MEMORY_MB, or give Docker more resources. See device capacity.
  • “Another process owns this worker” — a second worker is already serving on this device. Stop it, or let it finish.
  • “Device installation is incomplete” — an interrupted upgrade left a marker. Finish a reviewed reinstall; the worker will not claim jobs from a half-updated client tree.

A job isn’t running

If the device starts but work does not move, the cause is in this list:

  • Every device is offline. There is no paid cloud overflow, so a queued job waits until a device is awake with Docker running and a bounded session is serving.
  • The queue is at capacity. The worker prints “Queue is at capacity (N of M queued); no job claimed.” Admission is bounded per tenant; the count clears as jobs reach a terminal state.
  • The job’s label exceeds the device. A job asking for a larger class than the device opted into is refused, never silently reduced. Lower the job’s label or raise the device caps.
  • Policy does not admit the run. A client run whose workflow, event or branch is outside the tenant’s trigger policy is refused before it executes. Check the approved workflow file, event and branch with your operator.
  • The repository is outside the enrolled set. A device serves only the repositories its installation covers. Enroll the installation that includes this repository.
  • GitHub scheduled a sibling. GitHub’s own scheduler, not the queue order, picks which matching job runs first. A different job in the same repository may legitimately run first.

The worker polls for work and observes an assigned job on its own interval. A short delay between “queued” and “running” is normal and is not a failure.

Cleanup and attention

When a worker exits before it can remove what a job created, further jobs on that device are blocked on purpose: the platform would rather stop than run on a device with unclear local state.

  • “Owned-resource cleanup failed” — a container, network or runner could not be removed. The cleanup journal is retained; run recovery rather than deleting things by hand.
  • “Cleanup journal is not writable” — the device refused to create a job it could not record. Fix the permissions on the worker directory, then retry.
  • A job in attention — a sleeping or crashed worker’s job goes to attention, never to a fabricated success. Recovery decides what may return to the queue.

Recovery

Recovery checks the local resources the request owned and independently checks GitHub, then removes only the runner recorded for that request.

tbench-worker recover <request-id>

Two outcomes, and the difference matters:

  • A still-queued native job that GitHub never started and that has no assigned runner can return to the device queue under a fresh lease.
  • A started or ambiguous execution is non-retryable. It is reported honestly rather than re-run, because re-running work that may already have had side effects is worse than not running it.

The journal lists the exact names to inspect if recovery asks you to. It holds no credential, and the attestation recovery sends is identical whether or not a journal survives.

Enrollment problems

  • “Worker already enrolled” — the device already holds a bearer. Revoke it before enrolling again.
  • A single-use token that will not work — enrollment tokens are single-use. Ask your operator for a new one.
  • “Worker target is unset” — local state is missing or unreadable. Re-enroll, or repair the device state.

Never paste a credential into a page, a workflow, or a command argument. Enrollment uses a hidden prompt or a protected local file for exactly that reason. Installation alone grants no repository access.

Still stuck? Send the exact message and the output of tbench-worker check.