Michael Panchenko 16dc83616a
Add the ability to launch parallel devenv instances (#9906)
* 🐳 Split devenv compose for parallel workspaces

Move shared services into an infra compose file and keep the main devenv container plus Valkey in a separate compose file driven by defaults.env. Parameterize host-side ports, container names, source path, and runtime env while keeping container-internal ports fixed for same-origin proxying.

Make tmux startup idempotent, add attach-devenv for the live instance, move shared MinIO user setup to infra startup, and let exporter scripts load backend _env.local overrides.

Co-authored-by: Codex <codex@openai.com>

* 🐳 Run parallel devenv instances against shared infra

Add support for running N parallel devenv instances under separate compose
projects sharing Postgres, MinIO, mailer, and LDAP. Each instance has its
own main container, Valkey, source checkout, tmux session, and host port
range offset by 10000 (3449 -> 13449 -> 23449, etc.).

./manage.sh run-devenv-agentic --n-instances N reconciles the running set
to exactly {ws0..ws(N-1)}: missing instances are created (workspace sync
from the live repo via git ls-files + per-instance env-file generation
under docker/devenv/instances/ + detached tmux startup), surplus instances
are stopped highest-first via compose down (never -v), already-running
instances are left untouched. ws0 binds the live repo at PWD; ws1+ are
scratch clones under ~/.penpot/penpot_workspaces/.

Backend workers (enable-backend-worker) are gated on PENPOT_BACKEND_WORKER
in backend/scripts/_env; ws1+ overlays disable them so async-task
notifications stay bound to a single Valkey Pub/Sub instance.

Compose helpers wrap docker compose with env -i so per-instance overlay
--env-file actually overrides defaults.env -- without the strip, the shell
env from sourcing defaults.env at startup would shadow the overlay (Compose
gives shell precedence over --env-file).

Other:
- Drop network aliases (- main, - redis); use container_name for
  cross-container DNS so multiple instances on the shared network don't
  fight over the same DNS name.
- Pin volume names via name: (PENPOT_*_VOLUME) so volumes survive project
  renames; ws0 keeps the pre-existing physical names (penpotdev_*).
- Remove cross-project depends_on from main.yml (postgres/minio-setup now
  live in penpotdev-infra); manage.sh ensure-infra-up docker-waits on the
  minio-setup one-shot.
- Strict arg parsing in run-devenv / run-devenv-agentic; --n-instances 0
  rejected.
- Remove unused Host-matched server block from the Caddyfile.

Memory mem:devenv/core and developer docs updated.

Co-authored-by: Codex <codex@openai.com>

*  Document and stabilise the parallel-workspace CLI; wire AI agents

Improve parallel-workspaces developer CLI,
and add an opt-in layer that lets four AI
coding agents (Claude Code, opencode, VS Code Copilot, OpenAI Codex CLI)
drive a specific workspace through a single launcher command.

Parallel-workspace semantics
----------------------------

each run-devenv-agentic call brings up one wsN;
--ws N (integer; default 0) targets a specific workspace and auto-starts
ws0 first when N>=1 so the worker invariant holds. --sync is forbidden on
ws0 and re-seeds the workspace from the live repo for ws1+. Stop semantics
mirror the start invariant -- ws0 is the last to stop, shared infra stops
with it, --all walks every instance highest-first. The worker policy
section explains why workers run only on ws0 (Postgres FOR UPDATE
SKIP LOCKED is safe across many workers but the cron dedup primitive is
best-effort, and :telemetry / :audit-log-archive are not idempotent).
Per-instance Valkey Pub/Sub isolation, msgbus topology, and the
"async task notifications miss ws1+ tabs" caveat are stated explicitly.

The mem:prod-infra/core memory captures the same external-services and
task-queue / Pub-Sub topology in agent-readable form, and
mem:backend/core and mem:critical-info now cross-link it so backend work
surfaces the horizontal-scaling constraints from the start.

AI coding agent integration
---------------------------

New top-level .devenv/ directory holds committed templates
(templates/{claude-code,opencode,vscode}.json and templates/codex.toml,
each with \${PENPOT_MCP_PORT} and \${SERENA_MCP_PORT} placeholders) plus
committed shared entries (matching shared/* files for Playwright, the
only workspace-independent server we ship today).

./manage.sh start-coding-agent <claude|opencode|vscode|codex> [--ws N]
launches the chosen client against one workspace. It cd's into the
target's directory (the live repo for ws0; workspace-path "wsN" for ws1+)
and refuses to launch unless (a) the binary is on PATH, (b) the
workspace directory exists for ws1+, and (c) the instance is up
(devenv-main-running) -- the MCP servers only exist while the devenv is
running. The agentic-devenv guide is restructured around this Quick
start path, with a per-client table and a Manual configuration fallback
for clients we don't cover.

Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ♻️ Scope the shadow devtools to the dev build

---------

Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-03 15:48:25 +02:00

4.0 KiB

Production infrastructure (services Penpot depends on)

Backend (app.config, PENPOT_* env vars) is parameterized; deployments choose providers.

Services

  • PostgreSQL: durable store. Profiles, teams, files, sessions, audit, storage_object metadata, the task queue, scheduled_task cron registry, migrations. File-data also lives here when the file-data backend is legacy-db/db. One shared DB across all backends.
  • Redis (Valkey-compatible): per-backend message bus and cache. Concrete uses: msgbus Pub/Sub for collaborative-editing broadcasts and team/profile-org notifications fired by RPC handlers (app.rpc.notifications, files_update, teams, websocket); file-summary cache gated by enable-redis-cache; rate-limit counters; and the dispatcher→runner work hand-off list penpot.worker.queue:<tenant>:<queue>. PENPOT_REDIS_URI.
  • Object storage: backends :s3 and :fs. S3 in prod; devenv uses MinIO. Holds uploaded media, file-data when the file-data backend is storage, exports. Backend-side details (resolve, dedup, bucket set, file-data backends): mem:backend/http-storage-filedata-subtleties.
  • SMTP mailer: invitations, password resets, email verification (sent via the :sendmail worker task).
  • LDAP (optional auth provider): helpers in app.auth.*, gated by enable-login-with-ldap.

Task queue and worker model

Async tasks are enqueued via wrk/submit! (app.worker), which inserts a row into the shared Postgres task table tagged with queue = "<tenant>:<queue-name>". Submission is fire-and-forget — RPC handlers never poll, never wait, and workers never publish to msgbus. The only completion signal is the task row's status / completed_at columns, which nothing in rpc/ reads. Soft-delete RPCs return immediately after marking the top-level row, leaving the cascade and reaping to workers.

Workers run on backends with enable-backend-worker in PENPOT_FLAGS. Each worker-enabled backend has a dispatcher (polls task with FOR UPDATE SKIP LOCKED, marks status='scheduled', RPUSHes claimed task IDs into its own Redis list) and one or more runners per queue (BLPOP from that same local list, execute, update the Postgres row). The Redis hand-off list is purely intra-backend — cross-backend coordination happens at the Postgres row level.

Cross-backend safety

Postgres row locking is the only correctness primitive: task claims via FOR UPDATE SKIP LOCKED, cron firing via FOR UPDATE SKIP LOCKED on the scheduled_task row, plus task-handler-internal locks (e.g. file_gc_scheduler locks candidate file rows). This makes the work-claim path safe across any number of worker-enabled backends.

Two known race patterns survive multi-backend operation:

  • Cron dedup is best-effort. The lock on scheduled_task is released when the task body finishes. If two backends' cron timers fire for the same scheduled instant with a gap larger than the task body's runtime, both execute it. Penpot's cron entries are idempotent (session-gc, objects-gc, storage-gc-*, tasks-gc, upload-session-gc, file-gc-scheduler); the exceptions are :telemetry (would double-report) and :audit-log-archive (depends on archive target idempotency).
  • wrk/submit! ::dedupe true does a non-atomic DELETE then INSERT. Concurrent cross-backend submits can both bypass the DELETE (each sees the other's uncommitted insert as absent) and end up with duplicate 'new' rows. Each row claims and runs once independently, so the underlying work is fine; the "at most one pending" guarantee weakens.

Penpot in production lives with both: horizontal-scale deployments accept "exactly-once" as "essentially-once for idempotent operations." Devenv parallel instances handle it by running workers only on ws0 (see mem:devenv/core).

See also

  • Devenv composition and the ws0-only worker placement: mem:devenv/core.
  • Storage backend resolution, dedup, file-data lifecycle: mem:backend/http-storage-filedata-subtleties.