Measured, not guessed.
The early performance campaign focused on the canonical integration path because it is the most obviously serial part of the system. The numbers below are real measurements from the CP9 live gate, not synthetic targets.
Integration Queue baseline
| Scenario | Observed wall time | Approx. per item | What it proved |
|---|---|---|---|
| 3 clean PRs under churn | 18–22 s | ~6–7 s | FIFO ordering, successful canonical integration and predictable serialization. |
| 10 clean PRs | ~42 s queue→done | ~4.2 s | The queue remains correct under a larger burst; setup/clone work dominates. |
Where the time goes
The current implementation intentionally chose the simple, isolation-friendly substrate first: preview and merge operations use scratch Git clones. That made the correctness gates easy to reason about, but it repeats object transfer and repository preparation.
queued PR
│
├─ scratch clone → preview merge / contract validation
│
└─ scratch clone → canonical merge + non-force push
│
└─ Cloudflare Artifacts
| Cost | Why it exists | Likely optimization |
|---|---|---|
| Repeated clone/fetch setup | Strong isolation and a very simple correctness model. | Per-repository bare mirror refreshed with git fetch. |
| Repeated object transfer | Each scratch clone starts independently. | Shared object store/worktrees from the local mirror. |
| Strict FIFO integration | Canonical ordering must be deterministic. | Keep serialization; reduce work performed while the queue item owns the integration slot. |
Why strict FIFO is not the performance bug
Switchyard intentionally integrates one canonical change at a time. Parallelizing the final ref mutation would undermine the very freshness guarantee the queue provides. The optimization target is therefore the expensive preparation around each item, not making canonical last-writer-wins again.
What should be benchmarked next
| Experiment | Baseline | Question |
|---|---|---|
| Scratch clone vs bare mirror/worktree | ~4.2 s/item in the 10-PR gate | How much latency/object transfer disappears while preserving isolation? |
| Repository size sweep | Current gates use small repositories | How do object count and history depth affect preview/integration? |
| 1 / 3 / 10 / 25 queued PRs | 3 and 10 measured | Queue wait distribution and throughput as the burst grows. |
| Concurrent Agent executions | Deterministic runner only | Executor saturation and coordination overhead independently of LLM latency. |
| SSE fan-out | Single-node development use | How many simultaneous browser streams are practical before a hosted fan-out layer is warranted? |