Docs · Operations

Running it, reliably.

Switchyard is a stateful control plane around two durable authorities: Git state in Cloudflare Artifacts and coordination state in Trestle. Operations are easier once those responsibilities are kept separate.

What starts with the process

ComponentResponsibilityFailure behaviour
HTTP/API serverAuthenticated product UI and JSON API.Clients reconnect/retry after restart; durable data is elsewhere.
ReconcilerCompare Artifacts refs with the durable view.Missed time only delays convergence; the next pass catches up.
Cloudflare Queue consumerFast-path pushed-event ingestion.At-least-once messages are retried; reconciliation remains the safety net.
Workflow runnerAdvance durable JavaScript workflow steps.Completed steps replay recorded results after restart; in-flight work retries.
Integration Queue workerSerialize validated changes into canonical Git.Queue records survive; stale canonical movement is reconciled rather than overwritten.

Service management

For a persistent host, run Switchyard under a service manager rather than an interactive shell. Keep the binary and static directory versioned together so an API/UI mismatch cannot be introduced during deployment.

sudo systemctl status switchyard
sudo systemctl restart switchyard
sudo journalctl -u switchyard -f

A reverse proxy such as Caddy should listen publicly while Switchyard itself remains on 127.0.0.1:8080. Do not expose the Trestle admin/control surface directly to the Internet.

Restart safety

Backup matrix

StateWhere it livesBackup consideration
Git objects, refs, commitsCloudflare ArtifactsAuthoritative repository state. Normal Git mirroring is also possible if desired.
Work, Attempts, PR coordination, workflow runs, findings, queue, auditTrestleBack up the configured Trestle database/storage.
Provider credential ciphertextTrestleMeaningful only with the matching encryption key.
Fallback encryption keySWITCHYARD_DATA_DIR/secret.keyBack up securely and separately from ordinary application files.
Scratch clones/worktreesLocal data/scratch directoryDisposable; never treat as authoritative state.

Recovery scenarios

FailureExpected recovery
Switchyard process diesRestart service; workflow/queue records resume and Git reconciliation catches up.
External human pushes while Switchyard is downArtifacts accepts the normal Git push; reconciliation observes the new ref after restart.
Queue event duplicatedNormalized transition identity deduplicates the domain effect.
Queue event never arrivesPeriodic reconciliation discovers the same Git movement.
Stale integrationNon-force Git/CAS semantics reject it; Switchyard re-previews/requeues or blocks explicitly.

Credentials and host compromise

Provider credentials are encrypted at rest with AES-256-GCM and never returned from metadata APIs. Plaintext is resolved only for an execution. The self-hosted fallback key may live on the same machine, so this protects against database-only or accidental disclosure — not an attacker who controls the whole host. A KMS/external secret backend is the appropriate stronger hosted/enterprise direction.

Observability

Today the control plane logs startup, reconciliation, queue, workflow and integration activity. The durable records themselves are also useful operational evidence: queue state explains why canonical has not moved, workflow steps explain replay, and audit records explain policy decisions. Structured metrics/alerting remain hardening work rather than being falsely claimed as complete.