125 Commits

Author SHA1 Message Date
Andrey Antukh
e6983a1e4e
✨ Enforce idle and absolute session expiration (#11654)
* ✨ Enforce idle and absolute session expiration

Sessions now expire on two server-side conditions: an idle window
(PENPOT_AUTH_TOKEN_COOKIE_MAX_AGE, default 7d) and an absolute cap
from creation (PENPOT_AUTH_TOKEN_COOKIE_MAX_AGE_ABSOLUTE, default
30d, enforced by the token :exp claim). The daily session-gc task
deletes rows that exceed either window, so idle sessions can no
longer be replayed and active sessions are not deleted at the idle
window.

Also remove the legacy v1 HTTP sessions: the http_session table and
the string-id / :ver 0 token code paths are gone. Any v1 cookie now
requires a fresh login.

Document the session expiration configuration in the technical guide
and add a backend memory describing the token, renewal and GC model.

Closes #11646

AI-assisted-by: deepseek-v4.1-flash

* 🐛 Address session-expiration review findings F1-F4

Fix the unreadable test (a stray paren broke whole-suite
discovery), enforce idle expiration on every request in
wrap-authz, fail boot fast when the absolute cap sits below
the idle window, and align config defaults with the memory
rule while fixing its migration number and stale reference.

Closes #11646

AI-assisted-by: muse-spark-1.3-contributor
2026-10-01 10:12:05 +02:00
Andrey Antukh
d67a00c1d5
✨ Normalize storage metadata with a closed schema (#11987)
* ✨ Add Malli schema for storage metadata with dual decode

Phase 1 of the storage_object.metadata migration: reads accept both
Transit and plain JSON (sniffed by the marker) and always return the
normalized shape; writes validate against a closed per-bucket Malli
schema and still serialize as Transit unless the new
:storage-metadata-as-json config flag is set.

The 0155 migration normalizes existing rows inside Transit (reference
to bucket, default bucket, drop of chunk leftovers) and is
idempotent; large instances should fake it and run the batched
script instead.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Address review findings on storage metadata Phase 1

Collapse the dead :reference leg of the gc-touched bucket fallback
(the decode always sets :bucket on non-nil metadata, so it is only
reachable with a NULL column) and fix its comment.

Pin the write flag off in the transit-assuming metadata tests so the
suite proves the same with the flag set, and add coverage for the
flag rollback contract, JSON hash survival, NULL metadata in
gc-touched, and the 0155 normalization statements.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Backfill NULLs, canonical buckets, comment fix

Backfill NULL metadata columns in 0155 via coalesce (the key-missing
rule already matches them), derive valid-buckets from the Malli schema
dispatch entries so the list lives in one place, and correct the
lookup-bucket fallback comment to NULL columns.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Defer corrupt metadata rows in storage gc-touched

Decode touched rows individually so one non-map metadata value no
longer aborts the whole chunk: corrupt rows are logged and deferred
exactly one day in the same transaction, keeping their metadata
intact for a later repair, while healthy rows process normally.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Address storage metadata phase 1 review findings

Address the review findings on the storage metadata phase 1 branch:

- Fix put-and-delete-object: it stored the object with
  ::sto/expired-at, so the row was already deleted and del-object!
  returned false. Add delete-expired-object-returns-false to keep
  the expired-delete case covered.
- Cache the Malli decoder and encoder per process. Building them
  compiles the closed multi-dispatch schema, and decode-metadata
  runs on every read path (get-object, dedup probes, GC batches).
- Catch Exception instead of Throwable in try-decode-row so JVM
  Errors are not deferred as corrupt metadata.
- Add penpot_storage_gc_poison_total, emitted from
  storage-gc-touched; wire ::mtx/metrics into its handler.
- Anchor the encoding sniff to the start of the document so a
  plain JSON value that begins with a Transit-looking prefix is
  not read as Transit.
- Cover every bucket on both encodings, a JSON roundtrip through
  the jsonb column, nil metadata, the canonical bucket set and a
  poison-only GC chunk.
- Rename private check-metadata! to check-metadata.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Simplify the storage metadata schema to a single map

Replace the per-bucket :multi dispatch with a single closed map: the
bucket is validated with ::sm/one-of over metadata-buckets (now a plain
set) and the remaining keys are typed optional fields. Per-bucket
enforcement shrinks to a one-line :fn guard requiring :file-id and :id
for file-data, whose ids the GC reads to resolve references.

- Drop the dead (sm/register! ::metadata ...): nothing references the
  schema by keyword.
- Define tempfile-bucket and upload-session-bucket in the schema and
  alias them from app.storage, removing duplicated literals.
- Keep content-type required and the map closed, so an unknown bucket
  or key still fails fast on write.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Drop input coercion from encode-metadata

encode-metadata no longer runs the json-transformer decoder before
validation. On the write path its only effect was coercing string
UUIDs to UUID, and every producer already passes native UUIDs (the
RPC profile-id, uuid/random, or binfile ids decoded as ::sm/uuid).
Reads keep decoding, so stored Transit or JSON values still come
back as native types.

- Replace encode-accepts-string-uuids with
  encode-rejects-string-uuids, pinning the stricter contract.
- Pass native UUIDs in encode-writes-plain-json-with-flag.

AI-assisted-by: deepseek-v4.1-flash

* 📚 Document each statement in the storage metadata migration

Move the per-statement rules out of the header and add a comment to each
UPDATE explaining what it does: drop chunk leftovers, promote the legacy
"~:reference" to "~:bucket", drop residual "~:reference", and backfill the
default bucket. The header keeps the scope, the encoding note, the `->`
vs `?` note and the large-instance warning.

AI-assisted-by: deepseek-v4.1-flash

* 📚 Unwrap wrapped lines in the backend storage memory

One line per bullet or paragraph, as mem:memory-maintenance requires.
Only formatting; no content change.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Defer storage GC poison rows in their own transaction

process-chunk! no longer takes poison-ids; it only processes the healthy
chunk. The deferral moves to defer-poison! and process-touched! runs it in
its own transaction, separate from the freeze/delete work. The loop still
drains while there is chunk or poison, so a batch made only of poison rows
does not leave healthy rows behind the LIMIT 10 waiting for the next run.

Add a regression test: ten poison rows plus one healthy row with a later
touched_at are all handled in the same run.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Declare per-bucket metadata requirements in one map

Replace the file-data-specific predicate with bucket-requirements, a map
from bucket to the extra keys it must carry. metadata-buckets is derived
from its keys and a single generic :fn enforces presence, so a new bucket
and its contract are one entry. organization now requires
:organization-id; file-data keeps requiring :file-id and :id.

Update the http-assets test helper to set organization-id for its
organization objects.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Drop the ! suffix from storage GC helpers

Rename the internal helpers in app.storage.gc-touched (process-chunk,
defer-poison, mark-freeze-in-bulk, ...) to drop the trailing !.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Drop the ! suffix from storage GC deleted helpers

Rename the internal helpers in app.storage.gc-deleted (clean-deleted,
delete-sobjects, delete-give-up, ...) to drop the trailing !.

AI-assisted-by: deepseek-v4.1-flash

* 🐛 Fix dedup lookup for JSON-encoded storage metadata

get-database-object-by-hash only matched the Transit keys, so once the
:storage-metadata-as-json flag wrote plain JSON rows the dedup stopped
finding them and duplicated blobs. Match both encodings with a UNION ALL
of two indexable branches.

- Add migration 0156 with the plain-key dedup index; the legacy 0068
  index stays until Transit support is removed.
- Cover it with a JSON dedup test and a Transit -> JSON cross test.

AI-assisted-by: deepseek-v4.1-flash
2026-10-01 07:19:21 +02:00
Andrey Antukh
94d6f5a25c Merge remote-tracking branch 'origin/main' into staging 2026-09-29 23:07:26 +02:00
Andrey Antukh
3b886995ba 🐛 Emit the accept-organization-invitation audit event once
The event was written twice per accepted organization invitation. The
backend submitted it, and the browser then re-submitted a copy of the
props that the backend had already put in the response under
`:organization-invitation-audit` (`handle-token :team-invitation` in
`verify-token.cljs`). Both rows carried the same name with different prop
vocabularies, and the browser copy only existed when the browser finished
the flow.

Emit the event from the backend only. It now also carries the three props
that lived in the browser copy: the organization member count before the
add, the add source, and whether the invitee also joined a team. The
origin moves to the event context as `:event-origin`. The response no
longer includes `:organization-invitation-audit`, so the browser stops
emitting the event and `verify-token.cljs` drops its
`app.main.data.event` require.

The `accept-*` events of this command now share one prop vocabulary:
`:profile-id` for the accepting profile, `:invited-by` for the inviter
and `:profile-email` for the email, replacing the mix of
`:user-id`/`:user-who-send-invitation` and `:email`.

Audit consumers of `accept-organization-invitation` now see one row per
acceptance instead of two, and must read the new prop names.

AI-assisted-by: space-bunny-free
2026-09-29 22:40:03 +02:00
Andrey Antukh
0de865748d
🐛 Keep camelCase svg attribute names when importing a penpot file (#11974)
The json reader of the binfile v3 import rewrites every key of every
entry to kebab-case, nested maps included. Shape `:svg-attrs` is the one
map whose keys are camelCase react prop names, because the svg import
path runs them through `attrs->props`, so an attribute exported as
`fillRule` came back as `:fill-rule` and was stored that way.

The renderer looks the attribute up by its camelCase name, does not find
it and falls back to the default fill rule, so a shape exported with
`fillRule: evenodd` was painted without its hole, and the attributes
panel showed `fill-rule`.

`clean-shape-post-decode` already runs on every shape right after the
schema decode, for page shapes and component objects alike, so the
repair goes there: run `:svg-attrs` back through `attrs->props`, the
same transform that built the keys. It is idempotent, so shapes that
arrive correct are left untouched.

The new tests import a real export that carries the attribute, for
page shapes and component shapes.

Closes #11954

AI-assisted-by: space-bunny-free
2026-09-29 14:15:21 +02:00
Andrey Antukh
dc0ea3a69c
🐛 Attribute audit events to the caller, not the response (#11952)
* 🐛 Attribute audit events to the caller, not the response

prepare-rpc-event took the event profile-id from the result map
whenever it carried one, before falling back to the caller. Any
command returning a response with a :profile-id key silently
credited the action to somebody else.

get-error-report returns the report with its decoded content
merged in, and that content holds the profile that owned the
report, so privileged reads were logged against the users whose
crashes were being inspected. verify-token on a team invitation
returns the inviter's profile-id, so accepting an invitation was
logged against the inviter.

Resolution is now ::audit/profile-id metadata, then
::rpc/profile-id, then the zero uuid; the response is never
consulted. The two verify-token branches that relied on it now
declare the profile in the result metadata. Every other command
either already declared it or returns no :profile-id; all 30
registered command namespaces were checked.

The tests that pinned the old behavior are replaced by ones
covering the new contract.

AI-assisted-by: space-bunny-free

* 🐛 Coerce the audit profile-id override to a uuid

The only sanctioned way for a command to override the profile of an
audit event is the ::audit/profile-id metadata, and the value is set
by hand in a dozen commands, some of them reading it from token
claims or other sources we do not type.

schema:event requires a uuid and submit* swallows the validation
error, so a string did not fail loudly: the row was dropped silently.
Values that cannot become a uuid are now discarded with a warning
and the event falls back to the caller, which is always a valid uuid.
A uuid, the common case, exits on the first check.

AI-assisted-by: space-bunny-free
2026-09-29 09:47:07 +02:00
Andrey Antukh
f38c7dd639
⬆️ Update deps (#11960)
* 🐛 Migrate openUIApi schema to Zod v4 function syntax

Zod 4 removed z.function().args(), which broke the
plugins-runtime build with implicit-any errors on every
openUIApi parameter and knock-on possibly-null errors on
the modal in plugin-manager.

Declare the inputs with z.function({ input: [...] }) so the
parameter and return types infer again; behavior is unchanged.

Add a regression spec covering delegation, optional args and
rejection of invalid theme and title values.

AI-assisted-by: muse-spark-1.3-contributor

* 📚 Fix deprecated markdown-it-anchor permalink option in docs

Migrate docs Eleventy config to the markdown-it-anchor v10 API.
Replace the deprecated boolean permalink option with
linkInsideHeader, keeping the same symbol and class.
Bump markdown-it-anchor to v10 and related docs deps.

AI-assisted-by: muse-spark-1.3-contributor

* ⬆️ Update deps

* ⬆️ Update base docker images

* 📎 Fix mcp tests
2026-09-29 09:07:34 +02:00
Andrey Antukh
b890b94d27
🐛 Add a migration dropping the obsolete stroke-per-side attr (#11946)
Individual stroke widths shipped with a `:stroke-per-side` boolean on every
stroke, telling the renderer and the CSS generator whether the four
per-side widths were meaningful. Comparing the sides answers that on its
own, so the attribute was later dropped from the closed stroke schema
and the toggle became ephemeral editor state.

Files written between those two changes still carry the attribute, and
the closed schema rejects the unknown key, so loading such a file fails
`check-file-data` with a `:malli.core/extra-key` error and surfaces as
an internal error in the editor.

Add migration `0030-remove-stroke-per-side-attr`, which drops the
attribute from every stroke of every page and component shape. The
per-side widths are the saved design data and are kept untouched, so a
file whose boolean was `true` renders exactly as before.

The migration is naturally idempotent and a no-op for files that never
carried the attribute, since `dissoc` on a map without the key returns
an equal map.

Cover it with tests for the schema rejection, the per-side and global
widths surviving, component shapes, idempotency, and the run through
`migrate-file`.

Closes #11943

AI-assisted-by: space-bunny-free
2026-09-28 14:57:17 +02:00
Alonso Torres
9d08e26cb3
✨ Add japanese translations (#11922) 2026-09-25 14:57:56 +02:00
Andrey Antukh
eec06a5987
🐛 Allow registration with disabled public registration (#11912)
* 🐛 Allow invitation-based registration when disable-registration is set (#5178)

Per documentation, disable-registration 'disables registration
(still enabled for invitations only)'. Two bugs prevented this:

1. verify_token.clj: when processing an invitation token for a
   non-logged-in user with no member-id, the redirect included
   registration-disabled? in its condition, sending invited users
   to the login page instead of the register page.

2. auth.clj validate-register-attempt!: the registration-disabled
   check fired unconditionally before the invitation-token check,
   rejecting the actual register RPC even with a valid invitation.

Fix: in verify_token.clj remove registration-disabled? from the
redirect condition for new-user invitations. In auth.clj restructure
the check as an if/else: with an invitation token, validate the token
and allow registration; without one, enforce the flag as before.

* 🐛 Allow registration with disabled public registration

Allow valid team invitations to create new profiles when public
registration is disabled, while keeping password login and invitation
validation required.

Add backend regression coverage for flag combinations and verify-token
redirects, frontend route coverage, and configuration documentation.
Closes #5178

AI-assisted-by: space-bunny-free

* 🐛 Revalidate active invitation during registration

Require a live, unexpired team invitation before using the
registration exception, and recheck it before creating a profile.
Reuse the same lookup in invitation token verification.

Add regression tests for canceled and expired invitations, the
registration race, and explicit redirect contracts. Update docs
and backend auth guidance.

AI-assisted-by: Space Bunny Free

* 🐛 Lock and normalize invitation registration checks

Lock active invitation rows during transactional registration and
acceptance so cancellations cannot race with profile or membership
creation.

Normalize invitation emails before comparisons and database lookups.
Add concurrency, email casing, and final flag regression tests.

AI-assisted-by: Space Bunny Free

---------

Co-authored-by: Sumit Ridhal <sridhal@redhat.com>
2026-09-25 14:13:12 +02:00
Andrey Antukh
4b978767ea
🌐 Clean up en translations (#11853)
Drop 277 keys nothing references from en.po (verified against
frontend/src and common/src) and let sync propagate the
deletions to every locale. Clear all 10 fuzzy entries: fill
the 5 empty translations, keep the 4 valid ones, and drop the
duplicated max-quote-reached in favor of max-quota-reached
(the backend code stays, the UI maps it to the quota text).

Recover 22 used-but-missing keys with translations: the 19
shortcuts section/subsection labels plus connected-to,
pixel-grid-color and tokens.add-set. Make the rest
statically visible to rehash instead: :label fns on shortcut
commands, sections and subsections (one debug-only and one
colorpicker-local id exempt); case branches in place of
dm/str-built keys (export modal, text decoration and
transform, undo history with raw-key fallback); hoist
conditionals out of tr calls; pre-translate modal props and
role labels; replace the lone (i18n/tr ...) site with tr.
Turn static :error/code data into eager :error/fn calls in
the common schemas and the auth/password forms. Rename the
two keys containing spaces and point team leave at
max-quota-reached. Backend-driven keys stay dynamic by
design, declared with (tr ...) comments: the five
weak-password details, team and organization notifications.

Tooling: rehash also scans common/src and no longer treats
a missing -l as no locale; new clj-kondo tr-dynamic warning
flags non-literal tr args (lint scripts use --fail-level
error so it never fails CI); tr docstring states the
literal-only rule. Tests cover the shortcut label wiring,
the undo-history fallback and the :error/fn schemas.
Translations memory rewritten to match; es check word list
gains three entries.

Rebased onto develop: adopt the register field-error UX
(the weak-password declarations move onto the :options
code), keep develop's newer keys (connection-error,
account-locked, save-retrying, tokens-source strings) with
fresh references, and reword the shortcuts.cljs prose
comment so rehash does not invent a "literal" key.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-25 11:11:44 +02:00
Andrey Antukh
5ef70c7284
✨ Serialize render-wasm builds per checkout (#11903)
Protect shared setup, compilation, artifact copy, and target cleanup with
one flock lock per checkout. Route watch builds and frontend cleanup
through the protected scripts.

Document the lock contract and normalize the frontend and exporter
build:wasm commands.

Closes #11901

AI-assisted-by: Space Bunny Free
2026-09-25 09:42:27 +02:00
Andrey Antukh
de14311ce7
⚡ Add xf:add-index and memoize interactions menu rendering (#11915)
Introduce a shared xf:add-index transducer in app.common.data that
attaches the position to each item, and cover it with unit tests.

Use it in the workspace interactions menu: the indexed interactions
list is now derived in a memoized step keyed on the interactions
prop, so it is not rebuilt when the section is collapsed or
expanded. The previous code called d/enumerate on every render.

Update the frontend UI conventions memory with the pattern and the
constraint that the transducer only works on associative items.

AI-assisted-by: deepseek-v4.1-flash
2026-09-24 18:09:49 +02:00
Andrey Antukh
dad6026132
🐛 Test and document MCP plugin terminal-close handling (#11906)
Extract the WebSocket close-code decision into a pure ReconnectPolicy
module and cover it with node:test. The plugin now names the 1008
policy-violation code instead of using a magic number, and the terminal
behavior of a 1008 close is pinned by a regression test so a future
refactor cannot reintroduce the reconnect storm.

Document the contract in mem:mcp/core: 1008 is terminal, it comes from a
duplicate user token or a missing token in multi-user mode, and recovery
is manual via "Connect here".

Closes #11510

AI-assisted-by: deepseek-v4.1-flash
2026-09-24 15:57:23 +02:00
Andrey Antukh
acd146f6f4
🐛 Restore rate-limit headers and add Retry-After on 429 (#11895)
* 🐛 Restore rate-limit headers and add Retry-After

The account-lockout change replaced the header-forwarding 429 handler
with a body-only one, so existing RPC rate-limit responses lost their
x-rate-limit-remaining and x-rate-limit-reset headers. Account lockout
never sent Retry-After either.

Make handle-error :rate-limit preserve ::http/headers and add a
retry-after header when the exception carries a non-nil :ttl in
seconds, keeping the current JSON body. Add focused tests for both the
lockout and the RPC limiter paths.

Document activation, defaults, password/LDAP scope, Redis fail-open
behavior, and the lockout risk, and record the final HTTP contract in
the backend auth memory.

Refs #11397

AI-assisted-by: deepseek-v4.1-flash

* 🐛 Add Retry-After to RPC 429 and expose headers in CORS

Address review follow-ups on the account-lockout 429 contract:

- The RPC limiter now sets retry-after in its 429 headers (seconds
  until the longest rejecting limit resets), so it matches the
  account-lockout response and the HTTP standard.
- CORS exposes retry-after, x-rate-limit-remaining, and
  x-rate-limit-reset so browser clients can read them.
- Use backticks for Retry-After and account-locked in the docs for
  consistency with nearby sections.

Refs #11397

AI-assisted-by: deepseek-v4.1-flash
2026-09-24 12:19:29 +02:00
Andrey Antukh
fe8a305811 🐛 Fix RPC test helper dropping injected request metadata
The push-audit-events initiator test injects a request with
:app.http/auth-key-id via :app.http/request metadata. prepare-rpc-params
always replaced that metadata with a fresh DummyRequest, so the key id
was lost and the initiator fell back to "app".

Honor a caller-supplied request map, merging body params into its
:params, and keep the dummy request only for non-map IRequest stubs
(reify requests used by the audit tests, which cannot be assoc'd).

Add focused regression tests for the three cases and document the
helper contract in the backend testing memory.

AI-assisted-by: deepseek-v4.1-flash
2026-09-24 11:42:14 +02:00
Andrey Antukh
25eff238ae 🔧 Remove the link-issue verification step from gh.py
GitHub does not report mutation-created issue-to-PR links through
closedByPullRequestsReferences(userLinkedOnly: true), so the
verification in `gh.py link-issue` failed even when
addCloseIssueReferences succeeded and the link existed.

Drop the re-query and trust the successful mutation: the command now
fails only when a link target is missing or the mutation does not
return the issue. Update the tests, the gh helper memory, the PR/issue
workflow memories, and the create-pr skill so they no longer promise
verification.

AI-assisted-by: deepseek-v4.1-flash
2026-09-24 09:14:08 +00:00
Andrey Antukh
b3c1aab720 Merge remote-tracking branch 'origin/staging' into develop 2026-09-23 19:19:18 +02:00
Andrey Antukh
d6abd2fecd 🔧 Add explicit issue-to-PR linking command
Add a link-issue command that creates GitHub's Development reference
and verifies both sides. Keep Closes in descriptions for context, but
make the API link the source of truth, including for merged PRs.

Add tests for successful links, missing verification, output, and
failures. Update the PR workflow memories and create-pr skill to use
the command.

AI-assisted-by: space-bunny-free
2026-09-23 18:53:25 +02:00
Andrey Antukh
992170a9a4 🐛 Remove OpenCode V1 plugin support
Remove the legacy @opencode-ai/plugin import and V1 server export so
the local plugin loads under OpenCode V2 without project dependencies.

Add a focused smoke test for the V2 export and tool registration, and
update the related script memories.

AI-assisted-by: space-bunny-free
2026-09-23 18:53:25 +02:00
Andrey Antukh
5b3844c37a ✨ Add observability improvements (#11854)
* 🐳 Add upstream diagnostics to nginx access log

Enrich every access-log line with the internal journey of the request:
the status the backend answered (us), the time spent connecting to it
(uct), the time spent waiting for its answer (urt) and the internal
address that served the request (ua).

A plain 502 line used to say nothing about where the request died. With
this format, the tail of the line classifies the failure: connection
rejected, backend accepted and hung (uct + urt under 1s), or backend
stuck until read timeout. This was the missing witness in the Sep 20
incident, where nginx received connection resets with zero timeouts and
zero rejections.

Applied both to the production image template and the devenv config.
With proxy_pass on variables there is no upstream keepalive, so uct
measures one real TCP connection per request.

Parsing the new fields (us, uct, urt, ua) on the log shipper is left to
ops, so they can be filtered in Loki.

AI-assisted-by: glm-5.3-flash

* 🐳 Add stub_status endpoint for nginx metrics

Add a dedicated localhost-only server (listen 127.0.0.1:8082) exposing
/stub_status next to every other location of the public server. Ops can
run the official nginx-prometheus-exporter as a sidecar against
http://127.0.0.1:8082/stub_status and get nginx_connections_active,
accepted vs handled, reading/writing/waiting and request rates in
Prometheus.

Binding it to localhost and its own server keeps it unreachable from
outside the host and out of the public surface, and access_log off
avoids polluting Loki with one line per Prometheus scrape. The base
image already ships stub_status compiled in, so no image rebuild is
needed.

Applied both to the production image template and the devenv config.

AI-assisted-by: glm-5.3-flash

* ✨ Expose http server gate metrics (worker and connector)

The backend already measured dispatch latency but nothing reported the
state of the "house door": the xnio worker queue and threads, and the
monitor-level listener counters. This was the exact blind spot of the
Sep 20 incident, where the server kept answering health checks while it
accepted connections and dropped them without response.

Add a periodic metrics sampler that lives and dies with the http
server (single daemon thread, 15s interval, each sample guarded so an
unexpected error does not cancel subsequent runs) and publishes:

- worker (xnio MXBean gauges): penpot_http_worker_queue_size,
  busy_threads, pool_size and max_pool_size. Negative samples are
  discarded: the MXBean transiently reports -1 on the busy thread
  count (verified live), and a stale negative would read as zero.
- listener (Undertow connector statistics, enabled via the new
  :server/statistics yetti option): penpot_http_connector_active*
  _connections gauge and requests_total / errors_total counters.
  Undertow exposes absolute totals, so the sampler keeps a watermark
  atom and publishes deltas, skipping (and moving forward past) a
  counter reset.

The connector-level part depends on yetti v11.11, which now accepts
a :server/statistics server option (patch authored and released
upstream; before it, ListenerInfo#getConnectorStatistics always
returned nil).

New tests cover the samplers with fake MXBean/collector statistics
against real prometheus collectors, including the negative-sample
filter, the delta/watermark logic and the sampler lifecycle.

AI-assisted-by: glm-5.3-flash

* 🐛 Include jdk.management in the backend runtime JRE

The production image builds a trimmed JRE with jlink and omitted
jdk.management. Without that module the OS MXBean is
sun.management.BaseOperatingSystemImpl, which has no
getProcessCpuTime, getOpenFileDescriptorCount nor
getMaxFileDescriptorCount. The prometheus client StandardExports
reads those getters reflectively and collect() swallows the
NoSuchMethodException, so process_open_fds, process_max_fds and
process_cpu_seconds_total silently disappeared from /metrics while
the other process_* families kept flowing.

Verified against Prometheus: the app job only ever exposed
process_start_time_seconds, process_virtual_memory_bytes and
process_resident_memory_bytes; the fd and cpu families were absent.
Reproduced locally by running the backend metrics registry on a JRE
built with the same jlink module list (false/false/false) and on one
with jdk.management added (true/true/true).

Add the module to --add-modules and pin the metric contract with
backend-tests.metrics-test.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Build the http metrics sampler on promesa.exec

Replace the hand-rolled ScheduledThreadPoolExecutor and ThreadFactory
with promesa.exec primitives: px/scheduled-executor with a daemon
thread factory, and a px/schedule chain that reschedules the next
sample when the current one finishes.

Beyond fitting the existing periodic-task pattern (worker/cron,
rpc/rlimit), the chained schedule makes the docstring promise real:
with scheduleAtFixedRate an exception escaping the runnable cancelled
the following executions, while the reschedule now happens in a
finally block.

The sampler shutdown uses px/shutdown-now (shutdown! is deprecated in
promesa 12.0.0) to cancel the pending sample, keeping the previous
halt semantics.

The lifecycle test moves to the promesa predicates and a new test
covers the error-resilience promise: the first sample runs, throws,
and the next one is still scheduled.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Tighten the http metrics samplers

The samplers are leaf functions: they receive what they need and
publish it. Drop the internal nil guards (if there is no metrics
instance or no mxbean there is nothing to call them for) and move the
checks to the boundary, where the optional data is resolved:
sample-http-metrics now short-circuits with some-> and when-let.

Write the four worker gauges as four static operations instead of a
vector of pairs walked by doseq: the set is fixed, so the collection
only adds an allocation and hides each operation.

Drop the ! suffix from the sample-*-metrics family: ! marks a function
whose contract is to mutate state, while these report, and the mutation
happens in the mtx/run! they call. The constant true return, which only
existed so the removed guard tests could assert it, goes away too.

Tests follow the move: the internal-guard tests are replaced by one
boundary test (a nil server publishes nothing).

AI-assisted-by: deepseek-v4.1-flash

* 📚 Add the function design rules memory

Document the rules that came out of the http metrics sampler review:
preconditions are checked at the boundary instead of re-checked in the
core, optional-by-design data is guarded where the optionality is born,
a fixed set of operations is written statically, ! marks mutation and
not reporting, and production code is not shaped for tests.

Also state in the memory maintenance guide that memories must not use
manual line wrapping.

Linked from critical-info so it is read when designing a solution or
an API, not only when touching the samplers.

AI-assisted-by: deepseek-v4.1-flash

* 📚 Unwrap the critical-info memory lines

The memory maintenance guide forbids manual line wrapping, so rewrite
critical-info with one line per bullet and paragraph. A stray `*` at
the start of one continuation line is dropped.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Drop the redundant guard in the http server halt

create-metrics-sampler always returns the scheduler, so the sampler is
always present when integrant calls halt-key!; the nil check was dead
code, same as the yt/stop! call next to it.

AI-assisted-by: deepseek-v4.1-flash

* ✨ Add srepl helper to delete profiles by email

Add `delete-profiles-by-email!` to app.srepl.main. It accepts a
single email, a comma separated list of emails or a coll of emails,
resolves each profile, logs it to audit and enqueues the
delete-object task. The deleted-at is backdated with the configured
deletion-delay so profiles and their owned teams are purged on the
next gc pass.

Extract the per-email deletion logic into a private fn and reuse it
from `delete-profiles-in-bulk!`. Add tests for the new
`parse-emails` helper.

AI-assisted-by: glm-5.3-flash
2026-09-23 18:03:03 +02:00
Andrey Antukh
452f38cf5d
✨ Add Prometheus metrics for storage operations (#11700)
* ✨ Add storage operation metrics for S3 and buckets

Expose Prometheus metrics for the object storage subsystem.

The S3 backend now attaches an AWS SDK MetricPublisher that counts
API calls, retries and latency per operation and target. The storage
layer counts logical operations and deduplication outcomes per Penpot
bucket, and the assets handlers count served requests per route.

Closes #11676

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Fix storage metrics labels, errors and test gaps

Address the review findings on the storage metrics commit.

Label reads with the object's own backend, count failed asset
serving as errors without swallowing them, and cover the failed
S3 call, S3 asset path and permission-denied branches with tests.
Also share the label helper and reuse the metrics test helper.

Closes #11676

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Harden storage metrics and fill test gaps

Address the second-round review findings on storage metrics.

Unknown backends now fail explicitly and count as errors, exists
stays paired with its dedup outcome, and the thumbnail, missing
storage, expired reads, unknown buckets and write failure paths
are covered by tests. Label coercion goes through the shared
metrics helper.

Closes #11676

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Harden storage metrics accuracy and coverage

Address the full-branch review findings on storage metrics.

Touch and delete emit only on changed rows, reads emit after the
backend fetch, unknown backends fail explicitly, and tempfile
mismatches count as unauthorized. Publisher nil policy, pairing
rules and attempt semantics are documented and covered by tests.

Closes #11676

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Address full-branch review findings on storage metrics

Touch and delete resolve labels from the row, reads stay paired,
failures are covered by tests, and logging, ranges and docs are
tightened. Includes the label helper unit tests and the retries
wording clarification.

Closes #11676

AI-assisted-by: muse-spark-1.3-contributor

* ⚡ Label touch and del metrics from UPDATE RETURNING

The storage metrics change resolved metric labels for touch-object!
and del-object! with an extra SELECT per id-based call. Since app.main
instruments storage unconditionally, every GC collector and binfile
import paid that extra round trip: deleting a team with 10k media
objects doubled the storage_object statements exactly on the paths
that already process the most rows.

touch-object! and del-object! now take only the object id (UUID) and
read the labels from the updated row itself via RETURNING id, backend,
metadata: one statement, no pre-read, and labels that always match the
row actually mutated. del-object! additionally guards on deleted_at
IS NULL, so a repeated delete returns false and emits no metric.

Also from the review of the full branch: extract the duplicated
serve/emit/rethrow block in app.http.assets into one helper; give
penpot_storage_s3_timing explicit histogram buckets up to 60s (the
default cap at 7.5s hid the slow S3 calls the metric exists for);
drop the unused ::target-id config key from the S3 backend and
hardcode the :default target label until per-bucket routing lands.

AI-assisted-by: glm-5.3-flash

* ✨ Harden storage metric recording and definitions

The metric definition schema is now closed and declares every key
the collectors read: buckets, quantiles, max-age and reg. A typo
such as a misspelled ::mdef/buckets used to compile and silently
fall back to the default histogram buckets; it now fails at
startup.

The asset result-label fallback coerced an absent status to 500,
so a future serve path without a status would have counted
successes as errors. The mapping is now explicit and documented:
served below 400, unauthorized for 401/403, not-found for 404,
and error for everything else, including an absent status.

The never-fail try/catch around metric recording existed four
times with drift. One app.metrics/run-safe! helper replaces them:
it no-ops on a nil metrics instance and logs the first failure
per hint at warn level, then at debug, so a broken setup surfaces
once without flooding the log. The S3 publisher keeps its outer
try/catch: it is the SDK MetricPublisher contract boundary.

AI-assisted-by: glm-5.3-flash

* ✨ Make metrics mandatory and run! safe by default

Recording a metric must never change the behavior of the operation
being measured, so `run!` now catches recording failures itself: the
first failure per metric id logs at warn, later ones at debug. This
replaces the `run-safe!` helper, whose four copies had drifted, and
applies the guarantee to every emit site instead of only storage.

The metrics instance precondition is a plain assert, and the collector
lookup stays outside the recording guard, so a missing instance fails
hard even when asserts are disabled. Metrics is therefore no longer
optional: the storage, s3-backend and db-pool schemas require
`::mtx/metrics`, and the assets handler cfg always carries it.

`wrap-publisher` no longer returns nil for a nil instance, and the db
pool wires the prometheus tracker unconditionally.

AI-assisted-by: deepseek-v4.1-flash
2026-09-23 15:25:21 +02:00
Andrey Antukh
89e91ba372
✨ Add observability improvements (#11854)
* 🐳 Add upstream diagnostics to nginx access log

Enrich every access-log line with the internal journey of the request:
the status the backend answered (us), the time spent connecting to it
(uct), the time spent waiting for its answer (urt) and the internal
address that served the request (ua).

A plain 502 line used to say nothing about where the request died. With
this format, the tail of the line classifies the failure: connection
rejected, backend accepted and hung (uct + urt under 1s), or backend
stuck until read timeout. This was the missing witness in the Sep 20
incident, where nginx received connection resets with zero timeouts and
zero rejections.

Applied both to the production image template and the devenv config.
With proxy_pass on variables there is no upstream keepalive, so uct
measures one real TCP connection per request.

Parsing the new fields (us, uct, urt, ua) on the log shipper is left to
ops, so they can be filtered in Loki.

AI-assisted-by: glm-5.3-flash

* 🐳 Add stub_status endpoint for nginx metrics

Add a dedicated localhost-only server (listen 127.0.0.1:8082) exposing
/stub_status next to every other location of the public server. Ops can
run the official nginx-prometheus-exporter as a sidecar against
http://127.0.0.1:8082/stub_status and get nginx_connections_active,
accepted vs handled, reading/writing/waiting and request rates in
Prometheus.

Binding it to localhost and its own server keeps it unreachable from
outside the host and out of the public surface, and access_log off
avoids polluting Loki with one line per Prometheus scrape. The base
image already ships stub_status compiled in, so no image rebuild is
needed.

Applied both to the production image template and the devenv config.

AI-assisted-by: glm-5.3-flash

* ✨ Expose http server gate metrics (worker and connector)

The backend already measured dispatch latency but nothing reported the
state of the "house door": the xnio worker queue and threads, and the
monitor-level listener counters. This was the exact blind spot of the
Sep 20 incident, where the server kept answering health checks while it
accepted connections and dropped them without response.

Add a periodic metrics sampler that lives and dies with the http
server (single daemon thread, 15s interval, each sample guarded so an
unexpected error does not cancel subsequent runs) and publishes:

- worker (xnio MXBean gauges): penpot_http_worker_queue_size,
  busy_threads, pool_size and max_pool_size. Negative samples are
  discarded: the MXBean transiently reports -1 on the busy thread
  count (verified live), and a stale negative would read as zero.
- listener (Undertow connector statistics, enabled via the new
  :server/statistics yetti option): penpot_http_connector_active*
  _connections gauge and requests_total / errors_total counters.
  Undertow exposes absolute totals, so the sampler keeps a watermark
  atom and publishes deltas, skipping (and moving forward past) a
  counter reset.

The connector-level part depends on yetti v11.11, which now accepts
a :server/statistics server option (patch authored and released
upstream; before it, ListenerInfo#getConnectorStatistics always
returned nil).

New tests cover the samplers with fake MXBean/collector statistics
against real prometheus collectors, including the negative-sample
filter, the delta/watermark logic and the sampler lifecycle.

AI-assisted-by: glm-5.3-flash

* 🐛 Include jdk.management in the backend runtime JRE

The production image builds a trimmed JRE with jlink and omitted
jdk.management. Without that module the OS MXBean is
sun.management.BaseOperatingSystemImpl, which has no
getProcessCpuTime, getOpenFileDescriptorCount nor
getMaxFileDescriptorCount. The prometheus client StandardExports
reads those getters reflectively and collect() swallows the
NoSuchMethodException, so process_open_fds, process_max_fds and
process_cpu_seconds_total silently disappeared from /metrics while
the other process_* families kept flowing.

Verified against Prometheus: the app job only ever exposed
process_start_time_seconds, process_virtual_memory_bytes and
process_resident_memory_bytes; the fd and cpu families were absent.
Reproduced locally by running the backend metrics registry on a JRE
built with the same jlink module list (false/false/false) and on one
with jdk.management added (true/true/true).

Add the module to --add-modules and pin the metric contract with
backend-tests.metrics-test.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Build the http metrics sampler on promesa.exec

Replace the hand-rolled ScheduledThreadPoolExecutor and ThreadFactory
with promesa.exec primitives: px/scheduled-executor with a daemon
thread factory, and a px/schedule chain that reschedules the next
sample when the current one finishes.

Beyond fitting the existing periodic-task pattern (worker/cron,
rpc/rlimit), the chained schedule makes the docstring promise real:
with scheduleAtFixedRate an exception escaping the runnable cancelled
the following executions, while the reschedule now happens in a
finally block.

The sampler shutdown uses px/shutdown-now (shutdown! is deprecated in
promesa 12.0.0) to cancel the pending sample, keeping the previous
halt semantics.

The lifecycle test moves to the promesa predicates and a new test
covers the error-resilience promise: the first sample runs, throws,
and the next one is still scheduled.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Tighten the http metrics samplers

The samplers are leaf functions: they receive what they need and
publish it. Drop the internal nil guards (if there is no metrics
instance or no mxbean there is nothing to call them for) and move the
checks to the boundary, where the optional data is resolved:
sample-http-metrics now short-circuits with some-> and when-let.

Write the four worker gauges as four static operations instead of a
vector of pairs walked by doseq: the set is fixed, so the collection
only adds an allocation and hides each operation.

Drop the ! suffix from the sample-*-metrics family: ! marks a function
whose contract is to mutate state, while these report, and the mutation
happens in the mtx/run! they call. The constant true return, which only
existed so the removed guard tests could assert it, goes away too.

Tests follow the move: the internal-guard tests are replaced by one
boundary test (a nil server publishes nothing).

AI-assisted-by: deepseek-v4.1-flash

* 📚 Add the function design rules memory

Document the rules that came out of the http metrics sampler review:
preconditions are checked at the boundary instead of re-checked in the
core, optional-by-design data is guarded where the optionality is born,
a fixed set of operations is written statically, ! marks mutation and
not reporting, and production code is not shaped for tests.

Also state in the memory maintenance guide that memories must not use
manual line wrapping.

Linked from critical-info so it is read when designing a solution or
an API, not only when touching the samplers.

AI-assisted-by: deepseek-v4.1-flash

* 📚 Unwrap the critical-info memory lines

The memory maintenance guide forbids manual line wrapping, so rewrite
critical-info with one line per bullet and paragraph. A stray `*` at
the start of one continuation line is dropped.

AI-assisted-by: deepseek-v4.1-flash

* ♻️ Drop the redundant guard in the http server halt

create-metrics-sampler always returns the scheduler, so the sampler is
always present when integrant calls halt-key!; the nil check was dead
code, same as the yt/stop! call next to it.

AI-assisted-by: deepseek-v4.1-flash

* ✨ Add srepl helper to delete profiles by email

Add `delete-profiles-by-email!` to app.srepl.main. It accepts a
single email, a comma separated list of emails or a coll of emails,
resolves each profile, logs it to audit and enqueues the
delete-object task. The deleted-at is backdated with the configured
deletion-delay so profiles and their owned teams are purged on the
next gc pass.

Extract the per-email deletion logic into a private fn and reuse it
from `delete-profiles-in-bulk!`. Add tests for the new
`parse-emails` helper.

AI-assisted-by: glm-5.3-flash
2026-09-23 12:50:18 +02:00
Andrey Antukh
1c7a73ec16
✨ Add account lockout after failed login attempts (#11402)
* ✨ Add account lockout after failed login attempts

Implement per-account brute-force protection using a Redis-backed
failed-login counter. After 5 failed attempts within 15 minutes, the
account is temporarily locked out and all login attempts (including
with the correct password) are rejected with a 429 response.

Closes #11397

AI-assisted-by: longcat-2.0

* 🐛 Bind LDAP session to directory-verified profile

The account-lockout change added a shortcut that preferred the
profile matching the typed email over the one returned by the LDAP
directory. These can differ with aliases, UPNs, or multi-valued mail
attributes, letting a user with valid LDAP credentials bind a session
to another Penpot account.

Keep the typed-email profile only for lockout checks. After LDAP
succeeds, resolve the session profile from the directory identity as
before and clear failed attempts on that profile.

AI-assisted-by: deepseek-v4.1-flash
2026-09-23 12:31:50 +02:00
alonso.torres
e62546a03a ✨ Wait previous fail state before retry 2026-09-23 12:05:24 +02:00
Andrey Antukh
34b24a9d9d ✨ Retry transient saves with backoff and reconnect notice
Classify save failures as transient or terminal (`transient-error?`
over the repo retryable types plus `:invalid-save-response`).
Transient failures keep the head commit queued under a new `:retrying`
status and resend it with backoff (2s/8s/20s, then terminal):
stamp rotation reuses the same `:commit-id`, the in-flight guard
prevents double-sends, and episode tokens silence stale timers.
One tagged reconnect notice per episode (hidden on save and on
terminal failure, silent recovery) plus a `:retrying` save-indicator
state; the browser `online` event and new edits resume the episode.
Terminal failures keep the exact `:error` path. Covers tasks 4, 6
and 7 with 31 persistence tests; updates the persistence memory.

Relates to #11724

AI-assisted-by: muse-spark-1.3-contributor
2026-09-23 12:05:24 +02:00
Andrey Antukh
95e551697f 🐛 Report environment failures as compact audit events
Connectivity and gateway failures (network, offline, 502/503 and
nitrate configuration) are not application defects, but offline fell
through to :default and 502/503 rendered exception-page, so they
reached the internal error reports and alerts with the full payload
(stack plus the last events). They are now classified as environment
failures and reported as audit-only handled-exception events.

generate-report accepts an explicit :format, as keyword arguments or as
a trailing map. :compact keeps the context header plus type, code and
uri, and skips the stack, the ex-data dump (which may contain request
headers) and the last-events list. flash derives the payload format from
the cause, so environment failures get a compact report; the audit event
name stays the canonical one requested by the caller
(handled-exception/unhandled-exception) because external tooling filters
on those names. Environment fingerprints drop the stack frame, so
grouping does not depend on the internal call site.

submit-report now requires an exception cause: a report without one is
ignored instead of using a separate fallback fingerprint, so a single
fingerprint format governs every report.

:offline gets its own handler and both connectivity handlers show the
new errors.connection-error message instead of the generic toast.

Closes #11743

AI-assisted-by: deepseek-v4.1-flash
2026-09-23 12:05:24 +02:00
Andrey Antukh
ee651b86d8 🐛 Bound error report amplification with a dedup governor
Add a report governor in app.main.errors: each report carries
a fingerprint, the first occurrence is always emitted, and
repeats within 2 minutes are counted and included in the next
emitted report as :occurrences. The fingerprint cache is
bounded by evicting the oldest entry.

flash reserves the report before generating it, so suppressed
occurrences do not build a report. static.cljs now passes the
cause so the exception page gets a full fingerprint.

Closes #11726

AI-assisted-by: deepseek-v4.1-flash
2026-09-23 12:05:24 +02:00
Andrey Antukh
cc0941343d 📚 Add backend audit-log memory documentation
Add a new memory file documenting the backend audit log system:
purpose, storage schema, RPC producers, frontend ingestion,
webhooks/error-reporter/telemetry consumers, and Nexus archival.

Also wire a reference to it from the backend core memory so it
is discoverable through the memory graph.

AI-assisted-by: longcat-2.0
2026-09-22 15:49:46 +00:00
Andrey Antukh
2f679eaa0e
💥 Remove client-provided id from creation RPC commands (#11784)
The seven creation commands no longer accept an optional client
id: create-file, create-project, create-team,
create-team-with-invitations, upload-file-media-object,
create-file-media-object-from-url and assemble-file-media-object.
The server always generates the identifier; a sent id is ignored.

Malli maps are open and the RPC layer never strips unknown params,
so the handlers that would still honor an id (create-file,
create-project) now drop it explicitly. Internal callers that pass
remapped ids (project duplicate, binfile import) keep working.

Closes #11783

AI-assisted-by: muse-spark-1.3-contributor
2026-09-22 15:51:50 +02:00
Andrey Antukh
2041473cc4
🐛 Add regression test for viewer zoom url loop (#11821)
Lock in the fix from 31b73460c3 (#11803) with a regression
test for the exact reported scenario: loading the viewer
with a URL that already contains `zoom=fill`.

At 2.18.0-RC5 `update-zoom-querystring` navigated without
any comparison, so the load sequence bundle-fetched →
zoom-to-fill → update-zoom-querystring → nav → navigated
re-ran forever and crashed the page with React error #185
("maximum update depth exceeded"). The guard added in
31b73460c3 breaks the cycle; the new test asserts that a
bundle fetch against a `zoom=fill` route emits no
navigation events.

Also updates the dashboard/viewer frontend memory to
document the guard and the loop it prevents.

AI-assisted-by: glm-5.3-flash
2026-09-22 14:58:32 +02:00
Andrey Antukh
bc3cb4bddf
🌐 Complete Catalan translations in frontend (#11741)
* 🌐 Complete Catalan translations in frontend

Complete the Catalan (ca.po) locale to 100% coverage against en.po,
using es.po as support reference. Adds the 1439 missing entries
across workspace, dashboard, labels, shortcuts, subscription,
errors, modals and onboarding, keeping vosaltres treatment and
IEC/Termcat terminology consistent with the existing strings.
Normalizes placeholders and plural forms, drops the 14 stale
obsolete entries and canonicalizes the file with the repo
translations script.

Closes #11739

AI-assisted-by: muse-spark-1.3-contributor

* 📚 Add frontend translations memory with Catalan criteria

Record the PO workflow, the sync fuzzy-flag gotcha and the
Catalan glossary and tone agreed upon while completing ca.po,
and link the new memory from the frontend core routing.

AI-assisted-by: muse-spark-1.3-contributor

* 🔧 Add gettext to devenv image

Provide msgfmt and msgattrib in the dev environment for
checking PO translation files.

AI-assisted-by: muse-spark-1.3-contributor

* 🌐 Fix Catalan translations and add PO checker

Review of the missing-whitespace pattern found ~90 glued words
across 75 entries, plus 4 lost plural forms and 2 placeholder
mismatches verified against tr call sites. All fixed in ca.po.

Adds frontend/scripts/check-translations.js (vocabulary-free PO
QA: glued words, punctuation, placeholders, plurals) with
--self-test, wired as pnpm run check-translations and
documented in mem:frontend/translations.

AI-assisted-by: muse-spark-1.3-contributor

* 🌐 Multi-locale PO checker with word catalogs

Split the checker engine from its word lists: ca/es catalogs now
live in scripts/check-translations/words.<locale>.txt and all
messages are in English. Adds an es seed (calibrated to zero
errors) and fixes 7 typos it found in es.po. Universal checks
(placeholders, plurals, punctuation) run without a catalog.

AI-assisted-by: muse-spark-1.3-contributor

* 🌐 Merge PO checker into translations.js

Fold check-translations.js into translations.js as a check
subcommand reusing its locale helpers; word lists stay in
scripts/check-translations/words.<locale>.txt. Also fixes the
getopts stopEarly bug that made -l useless after the command
(sync -l ca synced every locale), drops dead lodash import
and code, unifies help and exit codes. Removes the
check-translations package alias; use translations.js
check -l <locale> with explicit -l.

AI-assisted-by: muse-spark-1.3-contributor

* 🌐 Keep unused placeholders out of the gate

Reverts the %s-stripping on unused auth.terms-privacy-agreement:
the links mirror its markdown sibling and a reactivation may
need them. Placeholder mismatches on #, unused keys now warn
instead of failing, and the rule is recorded in
mem:frontend/translations.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-22 14:10:01 +02:00
Andrey Antukh
2255266d45
✨ Enable closed schemas for RPC methods (#11136)
* ✨ Enable closed schemas for RPC methods

* 🐛 Fix duplicate make-dummy-request test helper definition

The branch added a variadic DummyRequest/make-dummy-request pair but
left the pre-existing single-arg definition in place. Because it was
loaded last, zero-arg (make-dummy-request) calls added by
prepare-rpc-params and rpc-nitrate-test threw ArityException, which
broke 384 tests and caused 14 downstream assertion failures.

Remove the stale duplicate so the variadic definition is the only
one, and drop the now-unused yrq alias and duplicate yres alias.

AI-assisted-by: deepseek-v4.1-flash

* ✨ Add focused tests for make-dummy-request helper

Pin the call contract of make-dummy-request, which the suite uses
in three styles: no arguments, a single options map, and keyword
arguments. The helper's redefinition shadowing in 8ca95adb98 was
only caught by a full-suite run with hundreds of unrelated errors;
these tests fail locally in a focused --focus run.

Cover the zero-arg defaults, map and keyword overrides, the
:body-bytes -> ByteArrayInputStream wrapping, :body-stream
precedence, and cookie readback. Also clarify the docstring to
list all supported call styles.

AI-assisted-by: deepseek-v4.1-flash

* 🚑 Prevent RPC client params from overriding auth context

Strip qualified keys from decoded request params before merging
them with the server-built auth context, so transit bodies can
no longer override ::profile-id, ::auth-type or ::token-perms.
Adds a regression test proving the override and the fix.

AI-assisted-by: muse-spark-1.3-contributor

* 📚 Merge backend subtleties memories under generic name

Rename rpc-db-worker-subtleties to subtleties and fold in
http-storage-filedata-subtleties, so the name no longer
enumerates topics. Update all mem: references accordingly.

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Add realistic tests for RPC auth override

Cover the transit wire vector and the real wrapped :get-profile
method with two database profiles, proving a session cannot read
another profile by smuggling :app.rpc/profile-id in the body.

AI-assisted-by: muse-spark-1.3-contributor

* ✨ Add e2e test for RPC auth context override

Parametrize rpcPost with contentType, accept and query so e2e
can send hand-written transit bodies without new dependencies.
The new test proves a transit-smuggled :app.rpc/profile-id no
longer overrides the session in get-profile. Also fix the demo
email assertion in auth-flow to the current uuid format.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-22 14:03:02 +02:00
Andrey Antukh
117c8db0bb Merge remote-tracking branch 'origin/staging' into develop 2026-09-22 10:24:30 +02:00
Andrey Antukh
5c22f5bfb7
⚡ Build the frontend bundle once for all E2E suites (#11792)
* ⚡ Build the frontend bundle once for all E2E suites

Merge tests-integration, tests-composable-suite and tests-plugin-api-suite
into one "CI: E2E" workflow. Each of the three ran its own full
frontend/scripts/build on every PR, so one PR paid the build three times.

The new build-bundle job restores actions/cache key frontend-bundle-<sha>,
runs frontend/scripts/build only on a miss and saves the key before the
job ends. The integration shards, the composable suite and the mocked
Plugin API suite now all need build-bundle and restore the same key with
fail-on-cache-miss, so none of them builds. A workflow re-run of the same
SHA reuses the cached bundle instead of rebuilding it.

Triggers become the union of the previous paths (frontend, common,
render-wasm, plugins): the bundle embeds the built plugins, so a plugins
change runs the whole set. workflow_dispatch keeps running the
integration job only, as before.

Job names are kept identical on purpose: they are the GitHub check
contexts and branch protection may match them by name.

Docs: new mem:frontend/e2e-ci-workflow records the build-once contract,
referenced from mem:frontend/core and mem:frontend/testing; the composable
memory and both suite READMEs are updated.

AI-assisted-by: deepseek-v4.1-flash

* 🐛 Fix mocked plugin suites crashing without frontend deps

The mocked CI drivers shelled out to frontend/scripts/e2e-server.js,
which imports express from frontend/node_modules. CI jobs install
only plugins/ deps, so the import failed with ERR_MODULE_NOT_FOUND
and the run timed out waiting for localhost:3000.

Serve the prebuilt bundle with a zero-dependency static server
built into each driver (ci/static-server.ts, kept in sync in both
suites) plus node:test coverage for it.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-22 10:23:15 +02:00
Andrey Antukh
d68531b783
⬆️ Update devenv dependencies (#11790)
* ⬆️ Update devenv dependencies

Update Node.js, OpenCode, clj-kondo, Babashka, Pixi, GitHub CLI, uv,
and Serena to their current stable releases.

AI-assisted-by: gpt-5.6-sol

* ⬆️ Update devenv to Java 27

Use Zulu JDK 27 in the development image for compatibility testing.
Update the official checksums for both supported architectures.

AI-assisted-by: gpt-5.6-sol

* 🐳 Replace MinIO with RustFS in devenv

Run RustFS as the development S3 service and wait for its health check.
Install a pinned AWS CLI with checksums and use it to create the bucket
idempotently from each backend entry point.

Keep the old MinIO volume untouched and use a new RustFS volume.

AI-assisted-by: gpt-5.6-sol

* 🐳 Replace MailCatcher with persistent Mailpit

Run Mailpit as the devenv SMTP sink while preserving mailer:1025 and the
localhost:1080 UI.

Store its SQLite inbox in a named volume and wait for the readiness
endpoint before starting runtime containers. Bind the web UI to loopback so
development emails stay local.

AI-assisted-by: gpt-5.6-sol

* ⬆️ Update Node.js to 24.21.0

Align the host NVM version with the Node.js version used by devenv.

AI-assisted-by: gpt-5.6-sol

* ⬆️ Update devenv to PostgreSQL 18.6

Run PostgreSQL 18 with its versioned volume layout and a TCP readiness
check that ignores the temporary initialization server.

Install the matching client, create penpot_nexus, and preserve the old
PostgreSQL 16 volume for rollback or logical migration.

AI-assisted-by: gpt-5.6-sol

* 🐳 Expose RustFS ports in devenv

Publish the RustFS S3 API and management console on localhost port 9000
and 9001.

Keep both bindings on loopback so object storage is not exposed to the local
network.

AI-assisted-by: gpt-5.6-sol

* 🐳 Install standalone pnpm in devenv

Install pnpm 12.5.0 from architecture-specific release archives and
verify their published checksums.

Remove the Corepack setup while allowing pnpm to honor the project
packageManager pins.

AI-assisted-by: gpt-5.6-sol

* 🔥 Remove corepack, use system pnpm everywhere

Corepack is gone from Node 25+, so every `corepack enable` call
fails. pnpm now ships as a system binary (devenv, CI runners and
Docker images install it directly) and auto-downloads the version
pinned in `packageManager` on mismatch.

Scripts, workflows and Dockerfiles call `pnpm` straight away; the
three deploy workflows use a single `pnpm/setup@v2` step; and the
new `scripts/sync-pnpm-version` stamps all 35 `packageManager`
fields from the system pnpm, replacing the `corepack use` sweep.

AI-assisted-by: muse-spark-1.3-contributor

* 🐛 Fix exporter watch missing render-wasm build step

The exporter watch compiled CLJS requiring the generated
src/app/wasm/shared.js, which only render-wasm/build export
produces. Without it shadow-cljs failed with a cryptic missing
./shared.js dependency. Run build:wasm before watching, as
the frontend watch:app and exporter scripts/build already do.

AI-assisted-by: muse-spark-1.3-contributor

* 🔧 Add opencode V2 support and adapt plugins

Register the penpot tools for both opencode V1 (server())
and V2 (setup() with JSON Schema inputs) from a single
dependency-free plugin file, sharing the psql and
paren-repair runners between both paths.

Install the opencode2 binary side-by-side with V1 in the
devenv image and document the dual registration in the
paren-repair and psql memories.

AI-assisted-by: muse-spark-1.3-contributor

* ⬆️ Update pnpm and opencode
2026-09-22 10:22:31 +02:00
Andrey Antukh
fe89e9e52c Merge remote-tracking branch 'origin/staging' into develop 2026-09-17 20:23:41 +02:00
Andrey Antukh
5b3e36489c 📚 Document int?/integer? predicate coverage in Clojure memory
Review assumed int? was 32-bit; it covers Long/Integer/Short/Byte.
Note it in mem:clojure/idioms so the mistake is not repeated.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-17 20:23:22 +02:00
Andrey Antukh
30e52af22e 📎 Backport creating-issue serena memories from develop 2026-09-16 19:48:50 +02:00
Andrey Antukh
df383be6b2 Merge remote-tracking branch 'origin/staging' into develop 2026-09-15 19:59:54 +02:00
Andrey Antukh
8ff99f0766 📚 Document how to add issues as sub-issues
Add the verified REST procedure for linking an issue as a
sub-issue of an umbrella/EPIC: get the REST id, POST to the
parent's sub_issues endpoint with a typed -F field, and verify
both directions. Route it from the create-issue skill.

AI-assisted-by: deepseek-v4.1-flash
2026-09-15 19:56:48 +02:00
Andrey Antukh
91ecfefa31 📚 Add EPIC issue type to creating-issues memory
The repository defines an EPIC issue type, but the memory only
listed Bug, Enhancement, Feature, Task, Question and Docs. Add
the EPIC row with its type id and its mapping entry so umbrella
tracking issues are typed consistently.

AI-assisted-by: deepseek-v4.1-flash
2026-09-15 16:12:02 +00:00
Andrey Antukh
ad7e035b63
✨ Add indirection for upload-chunks storage via new table (#11651)
* ✨ Add upload_session_chunk indirection for chunked uploads

Chunks now live in the upload_session_chunk table with non-deleting
foreign keys to storage_object and upload_session, instead of tempfile
objects with session metadata. Reads go through a JOIN, so chunk state
never scans storage_object.

Uploads validate the live session, reject duplicate indexes, and store
objects in the new upload-session bucket without extra metadata.
Assemble removes mappings and marks the session consumed; objects-gc
procedurally purges consumed and stalled sessions, touching referenced
objects first. Touched-gc and deleted-gc handle the new bucket, and
upload-session-gc is removed.

Closes #11644

AI-assisted-by: muse-spark-1.3-contributor

* 🐛 Fix quota, give-up and coverage for session chunks

Exclude consumed sessions from the sessions-per-profile quota so
finished uploads free their slot at once. Remove chunk mappings before
the gc-deleted give-up delete to respect the NO ACTION keys. Catch
java.sql.SQLException for duplicate chunks. Cover the profile-owned
session purge and the UNIQUE race backstop with tests.

AI-assisted-by: muse-spark-1.3-contributor

* 🐛 Align chunked upload tests with upload_session_chunk

Drop the duplicate-index tests written against metadata-backed
chunks; the UNIQUE mapping makes those cases unrepresentable and
the new tests cover them. Rewrite the rejected-duplicate tests to
expect :validation/:chunk-already-exists and assert against the
upload_session_chunk table, and scope the chunk-too-large "nothing
stored" check to the mapping table.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Use NO ACTION DEFERRABLE session FKs in single migration

Fold the profile FK change into 0154 so the feature ships one
migration. All three upload session FKs use ON DELETE NO ACTION
DEFERRABLE: identical to RESTRICT in normal operation, but deferrable
for tooling that relies on SET CONSTRAINTS ALL DEFERRED. Extend the
RESTRICT test to the direct profile delete.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Reserve chunk slot before writing blob in upload-chunk

Make object_id nullable and insert the mapping with NULL inside the
session-locking transaction, then write the blob outside it and link
it with a conditional update. A failed write removes the mapping and
reraises; a mid-flight death leaves a NULL row and the client starts
a new session.

AI-assisted-by: muse-spark-1.3-contributor

* 🔥 Remove redundant session_id index on upload_session_chunk

The UNIQUE(session_id, chunk_index) btree already serves
session_id-only lookups and the session FK check through its
leftmost column, so the standalone index only taxed the
per-chunk INSERT path. Verified with EXPLAIN on an equivalent
table shape.

AI-assisted-by: muse-spark-1.3-contributor

* ⚡ Merge chunk touch and delete into single RETURNING query

Replace the SELECT-then-DELETE round-trip in
delete-upload-sessions! with DELETE ... RETURNING object_id,
touching each returned object. Same semantics, one less query
per purged session. Follows the RETURNING pattern already used
in file-gc.

AI-assisted-by: muse-spark-1.3-contributor

* ♻️ Let objects-gc own chunk mapping deletion

Assemble-chunks now only marks the session as consumed; the
chunk mappings stay until objects-gc purges them (touching the
chunk objects first), leaving a single procedural deletion
path for consumed, stalled and profile-purge sessions.

AI-assisted-by: muse-spark-1.3-contributor

* 🐛 Fix font-deletion GC expectations for chunk objects

Update final storage-gc-touched counts to include the two
chunk objects touched by objects-gc when purging consumed
upload sessions (8/5/5 instead of 6/3/3).

AI-assisted-by: muse-spark-1.3-contributor

* 🐛 Release chunk reservation when the link UPDATE fails

Review feedback on #11651: the link UPDATE in upload-chunk could
leave a NULL reservation behind, blocking retries of the same
index with :chunk-already-exists. Remove the reservation when
the link fails so the client can retry in the same session;
the orphaned blob stays touched for touched-gc. Also realign
the process-bucket! cond branches in gc-touched.

Tests: chunked-upload-link-failure-releases-slot and
chunked-upload-null-reservation-blocks-retry.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-15 17:26:43 +02:00
Andrey Antukh
b660ea9d53 ✨ Route the app by query string with screen key
Move SPA routing out of the URL fragment into the normal query
string. The screen travels in a reserved `screen` key holding
the route name (`?screen=workspace&team-id=…`); every other
param keeps its name. `rt/nav` and `rt/resolve` keep their
signatures.

This deletes the fragment-mirroring URL surgery, simplifies
link-preview (the server sees everything) and nginx (single
path, no SPA fallback rules needed), and migrates OIDC
redirects, email links, e2e helpers and plugin test utils to
the new format.

Legacy `#/…` URLs translate client-side for one Penpot version
(`legacy-routes`, marked TODO(next-version)); non-SPA paths
are untouched.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-15 17:10:54 +02:00
Andrey Antukh
39ca4c264c
🔥 Remove onboarding A/B test and welcome file creation (#11707)
Drop the onboarding-03 experiment consulted through
external-feature-flag, keeping the false-branch behavior: registration
never requests a welcome file and the workspace never shows the
onboarding modals. Remove the now-unused welcome-file machinery on the
backend (RPC wiring, welcome_file namespace, welcome-file-id prop and
the post-login redirect). Keep the external-feature-flag helper as the
seam for future experiments and note it in mem:frontend/core.

Closes #11705

AI-assisted-by: muse-spark-1.3-contributor
2026-09-15 14:14:32 +02:00
Andrey Antukh
bda8459d89 Merge remote-tracking branch 'origin/staging' into develop 2026-09-11 12:57:51 +02:00
Andrey Antukh
09736aa4c9 ✨ Enforce commit body line wrapping
Add a body line-length validator to scripts/check-commit. It
fails when a body line exceeds 76 characters, exempting
trailers, URLs, and unbreakable tokens. The 76 limit leaves
room for git log's four-space indent in an 80-column
terminal.

Align the subject limit with the documented 70 characters;
the checker allowed 90 before.

Document the rule as a hard, verifiable requirement in
AGENTS.md, CONTRIBUTING.md, the create-commit skill, and
the workflow memory, and point at scripts/check-commit.

Add tests for the validator and the subject length rule.

AI-assisted-by: deepseek-flash
2026-09-11 08:10:49 +00:00
Andrey Antukh
37dab75e1a Merge remote-tracking branch 'origin/staging' into develop 2026-09-10 20:29:44 +02:00
Andrey Antukh
eca1d81692 🔧 Remove legacy pnpm build key and clarify updating doc
Drop the ignored-since-pnpm-11 onlyBuiltDependencies entry from
render-wasm/pnpm-workspace.yaml, keeping allowBuilds as the single
source of build approvals. Clarify the updating-pnpm gotcha so it no
longer claims pnpm writes ignoredBuiltDependencies.

AI-assisted-by: muse-spark-1.3-contributor
2026-09-09 11:45:00 +02:00
Andrey Antukh
831953c41e Merge remote-tracking branch 'origin/staging' into develop 2026-09-09 11:02:06 +02:00