mirror of
https://github.com/penpot/penpot.git
synced 2026-10-02 16:56:16 +00:00
Add a new memory file documenting the backend audit log system: purpose, storage schema, RPC producers, frontend ingestion, webhooks/error-reporter/telemetry consumers, and Nexus archival. Also wire a reference to it from the backend core memory so it is discoverable through the memory graph. AI-assisted-by: longcat-2.0
7.7 KiB
7.7 KiB
Backend Audit Log
Penpot records what users do as events in the Postgres audit_log table. There are two producers (the backend RPC layer and the frontend app) and four consumers (webhooks, error reporters, telemetry shipping, and the Nexus archive). Everything below follows that flow: purpose, storage, producers, consumers, archival.
Purpose
- The audit log answers "who did what, when, from where": every RPC mutation and selected frontend actions become a row with
name,type,profile-id,ip-addr,propsandcontext. Product analytics, abuse investigation and compliance exports all read from here, so keep events truthful and never put secrets inprops. - It is also the trigger bus for side effects: the same event object fans out to webhooks, error reporting and telemetry without the RPC handler knowing. New features should reuse this bus instead of building parallel notification paths.
Storage
- Live
audit_logcolumns:iduuid PK defaultgen_random_uuid();name/typetext NOT NULL;created_attimestamptz NOT NULL defaultnow()(server time, the source of truth);tracked_attimestamptz defaultnow()(client-claimed time, corrected on ingest);profile_iduuid NOT NULL;sourcetext telling full rows (backend/frontend) apart from anonymized copies (telemetry:backend/telemetry:frontend);ip_addrinet;props/contextjsonb holding transit-encoded maps;archived_attimestamptz set once Nexus acknowledges the row. - Indexes: PK on
(id); partialcreated_at WHERE archived_at IS NULLserving the archive scan; partialarchived_at WHERE archived_at IS NOT NULLserving the GC;(source, created_at)serving the telemetry scan. Each consumer has its own index, so a slow consumer never blocks the others.
Backend producers (app.loggers.audit)
- Most backend events need no manual code:
wrap-auditinapp.rpcruns after every RPC handler when:webhooks,:audit-logor:telemetryis on (unless the command sets::audit/skip) and builds the event viaprepare-rpc-event. The event name defaults to the command name (prefixed with<module>-outsidemain), props default to the request params, and timestamps come from the server request time. - Commands customize through result metadata (
rph/with-meta):::audit/replace-propsswaps the props wholesale (auth commands useprofile->propsso a register event carries the profile, not the password),::audit/propsmerges extras,::audit/context/profile-id/name/typeoverride the defaults.clean-propsalways strips nils, qualified keys and:session-id/:password/:old-password/:token/:client-secretas a last line of defense. submitis the normal entry point (fills defaults, validatesschema:event, runs insidetx-run!, logs failures without failing the RPC).insertis the low-level one for CLI/helpers and the webhook subsystem: direct write, no webhook/telemetry fan-out, silent unless:audit-logis on. Boot emitstrigger/instance-startfromsetup/propsso every restart is visible in the log.
Consumers I: webhooks (app.loggers.webhooks)
- Webhooks are the first dependent: when an event carries
::webhooks/event?,process-event(worker task:process-webhook-event) finds the team's active webhooks from the event props (team-id, elseproject-id, elsefile-id), records atrigger webhookrow, and enqueues one:run-webhookdelivery per match. Batching and dedupe come from the audit event itself (batch-key+batch-timeout), not from webhook config. :run-webhookPOSTs the event in the webhook'smtype(JSON camelCase, transit, or form-encoded), logs each attempt inwebhook-delivery, and disables the webhook after 3 consecutive errors. Delivery problems never touch the audit row itself.
Consumers II: error reporters (app.loggers.database, app.loggers.mattermost)
- Both reporters listen for backend
:errorlog records and for frontend crash events (recognized by the::audit/eventmarker), through a sliding-buffer channel so a flood of errors cannot stall the app. The database reporter persists them intoserver-error-report(source 4 = audit-log origin); the Mattermost reporter forwards a short notification to:error-report-webhookwhen configured. - Consequence for producers: crash reports only exist if the frontend collector is running and
push-audit-eventsacceptsunhandled-exception/exception-pageevents. Disabling the whole pipeline also blinds error reporting from the frontend.
Consumers III: telemetry (app.loggers.audit + app.tasks.telemetry)
- Telemetry reuses the same table with anonymized shadow rows (
source LIKE 'telemetry:%'): day-truncated timestamps,0.0.0.0IPs, props reduced to uuid/boolean/number values plus a few allowlisted fields (lang,auth-backend, derivedemail-domain, never raw emails), and a minimal context allowlist. Both full and shadow rows can coexist per event; that duplication is intentional. - The telemetry cron ships shadow rows to
:telemetry-urias JSON in 10k batches, deletes them on success, and purges leftovers older than 7d. Nothing is collected or sent on official hosts (telemetry-excluded?coverspenpot.app/penpot.dev).
Frontend ingestion (app.rpc.commands.audit, app.main.data.event)
- The browser cannot write to the table directly; it POSTs transit batches to
push-audit-events, which stamps serverid, sessionprofile-id, request ip and servercreated-at, and distrusts the client clock (future or >1h-laggingtracked-atis reset, original preserved in context). The endpoint is a no-op without:audit-log/:telemetryor on a read-only pool. - The in-browser collector (
app.main.data.event) only starts afterget-enabled-flagsconfirms the backend wants events. It turns Potok events and explicitev/eventcalls (nitrate membership changes, workspace file stats, crash reports) into a capped buffer (1024, chunks of 100, 2s debounce, current profile only) and sends fire-and-forget.skip-audit?exists for resumed dashboard actions so one user gesture is not counted twice. - Because collection is best-effort and includes
PerformanceObservernoise (performance-*triggers), backend tests must never assert exact frontend event counts.
Archival to Nexus and retention
- Long-term storage lives outside Penpot in Nexus. Every 5m the
:audit-log-archivecron takes chunks of 128 unarchived rows (FOR UPDATE SKIP LOCKED), POSTs them as transit{:events [...]}authenticated withx-shared-key: "nexus <key>"(:nexus-shared-key, else derived from the instance secret), and marksarchived_at=now()only on HTTP 204, in the same transaction. Anything else is retried on the next run; a missing URI with the flag on raises:task-not-configured. - Every 5m the
:audit-log-gccron deletes all archived rows (no age filter), so archive must run before GC or data ships never. Cron dedup is best-effort (mem:prod-infra/core): two backends can fire the archiver twice, which is why the Nexus endpoint must be idempotent and the DB only marks acknowledged rows. - Flags live in
common/flags.cljcvaria and are enabled asPENPOT_FLAGS=enable-<name>::audit-log,:audit-log-archive,:audit-log-gc,:audit-log-logger(structuredapp.auditlog).:telemetry-enabledconfig auto-adds:enable-telemetry.
Tests
backend_tests/rpc-audit-test.cljexercises the whole backend path (full-row insert, telemetry-only and dual-row modes,submit*, no-op without flags,insertgating,prepare-rpc-eventresolution) withwith-redefs [cf/flags #{...}]against realaudit_logrows.- Other RPC suites mock
app.loggers.audit/submit(nil return;helpers.cljstubs it globally) and assert on:call-args-list; any new command that must (or must not) emit an event needs the same treatment.