Docs · Performance

Measured, not guessed.

The early performance campaign focused on the canonical integration path because it is the most obviously serial part of the system. The numbers below are real measurements from the CP9 live gate, not synthetic targets.

Integration Queue baseline

ScenarioObserved wall timeApprox. per itemWhat it proved
3 clean PRs under churn18–22 s~6–7 sFIFO ordering, successful canonical integration and predictable serialization.
10 clean PRs~42 s queue→done~4.2 sThe queue remains correct under a larger burst; setup/clone work dominates.

Where the time goes

The current implementation intentionally chose the simple, isolation-friendly substrate first: preview and merge operations use scratch Git clones. That made the correctness gates easy to reason about, but it repeats object transfer and repository preparation.

queued PR
   │
   ├─ scratch clone → preview merge / contract validation
   │
   └─ scratch clone → canonical merge + non-force push
                         │
                         └─ Cloudflare Artifacts
CostWhy it existsLikely optimization
Repeated clone/fetch setupStrong isolation and a very simple correctness model.Per-repository bare mirror refreshed with git fetch.
Repeated object transferEach scratch clone starts independently.Shared object store/worktrees from the local mirror.
Strict FIFO integrationCanonical ordering must be deterministic.Keep serialization; reduce work performed while the queue item owns the integration slot.

Why strict FIFO is not the performance bug

Switchyard intentionally integrates one canonical change at a time. Parallelizing the final ref mutation would undermine the very freshness guarantee the queue provides. The optimization target is therefore the expensive preparation around each item, not making canonical last-writer-wins again.

What should be benchmarked next

ExperimentBaselineQuestion
Scratch clone vs bare mirror/worktree~4.2 s/item in the 10-PR gateHow much latency/object transfer disappears while preserving isolation?
Repository size sweepCurrent gates use small repositoriesHow do object count and history depth affect preview/integration?
1 / 3 / 10 / 25 queued PRs3 and 10 measuredQueue wait distribution and throughput as the burst grows.
Concurrent Agent executionsDeterministic runner onlyExecutor saturation and coordination overhead independently of LLM latency.
SSE fan-outSingle-node development useHow many simultaneous browser streams are practical before a hosted fan-out layer is warranted?
Performance principleMeasure the real bottleneck first. A cached mirror is justified by observed clone cost; redesigning the queue's correctness model is not.