ui-ux-pro-max-skill/scripts/generate-catalog-summary.py
speedy75015-crypto e2effd5775
fix(data): make catalog snapshot hashes line-ending independent (#478)
* fix(data): regenerate stale catalog-summary snapshot hashes

catalog-summary.json was not regenerated after google-fonts.csv,
google-font-licenses.json, icons.csv and phosphor-icons-upstream.json
changed, so `npm --prefix cli run verify:data` fails on a clean
checkout of main:

  validate:semantic          4 stale snapshot errors
  validate:catalog-summary   "catalog-summary.json is stale"
  test:python                1 failure / 153
  check:assets               2 files out of sync

Regenerated with the existing --verified-at 2026-08-26: only the four
sha256 fields change. The date is a human attestation that the font
catalog was checked against the upstream google/fonts repository, so it
is deliberately left untouched -- no such verification was performed
here.

verify:data now exits 0.

Note: prepublishOnly runs sync:assets before verify:data, which
regenerates the snapshot at publish time. That is why released packages
are unaffected and the drift stayed invisible on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015UidECV1wVBD8SW6Abuj71

* fix(data): make catalog snapshot hashes line-ending independent

Root cause of the stale snapshot restored in the previous commit.

bd19ab9 (#462) regenerated catalog-summary.json from a CRLF checkout.
Every recorded sha256 was the CRLF hash of its source file, so the
check failed on every LF platform. The four committed values are
exactly sha256(crlf_bytes):

  google-fonts.csv              committed d03194d2…  = CRLF hash
  google-font-licenses.json     committed 7c35e410…  = CRLF hash
  icons.csv                     committed 272ccf0e…  = CRLF hash
  phosphor-icons-upstream.json  committed 81c37fb3…  = CRLF hash

Two conditions had to combine: the digest hashed raw bytes, and no
.gitattributes pinned these files to LF, so Windows checkouts get CRLF
by default. Restoring the hashes alone would let the next contributor
on Windows reproduce the same commit.

Three changes:

- normalize line endings in generate-catalog-summary.py's digest(), so
  the snapshot no longer depends on the checkout
- apply the same normalization in validate_data.py, which independently
  recomputes the hashes and has to agree with the generator
- add .gitattributes pinning src/ui-ux-pro-max/data/*.{csv,json} to LF,
  so a Windows checkout matches the committed bytes in the first place

sync-assets.mjs already normalizes to LF, so this only extends an
existing project convention to the two places that were missing it.

Adds test_catalog_summary_line_endings.py: LF and CRLF inputs must
digest identically, the committed snapshot must match the normalized
sources, and a simulated CRLF checkout must still produce the recorded
hashes. The third case fails against the pre-fix digest.

verify:data exits 0; the Python suite goes from 153 to 156 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015UidECV1wVBD8SW6Abuj71

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 14:14:59 +07:00

146 lines
5.4 KiB
Python

#!/usr/bin/env python3
"""Generate or verify deterministic catalog counts and snapshot hashes."""
import argparse
import csv
import hashlib
import json
import sys
from datetime import date
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
DATA = ROOT / "src/ui-ux-pro-max/data"
OUTPUT = DATA / "catalog-summary.json"
def rows(name):
with (DATA / name).open(encoding="utf-8", newline="") as handle:
return list(csv.DictReader(handle))
def digest(path):
# Normalize line endings so the snapshot does not depend on whether the
# working tree was checked out with LF or CRLF.
return hashlib.sha256(path.read_bytes().replace(b"\r\n", b"\n")).hexdigest()
def load_json(name):
return json.loads((DATA / name).read_text(encoding="utf-8"))
def checked_date(value):
try:
parsed = date.fromisoformat(value)
except (TypeError, ValueError) as exc:
raise ValueError("verified-at must use YYYY-MM-DD") from exc
if parsed.year <= 1970 or parsed > date.today():
raise ValueError(f"verified-at has suspicious date {value!r}")
return value
def build(verified_at):
styles = rows("styles.csv")
stack_paths = sorted((DATA / "stacks").glob("*.csv"))
licenses = load_json("google-font-licenses.json")
icons = load_json("phosphor-icons-upstream.json")
excluded = licenses.get("excludedFamilies", [])
counts = {
"styles": {
"total": len(styles),
"searchable": sum(row.get("Status") != "deprecated" for row in styles),
"active": sum(row.get("Status") == "active" for row in styles),
"supplemental": sum(row.get("Status") == "supplemental" for row in styles),
"deprecated": sum(row.get("Status") == "deprecated" for row in styles),
},
"products": len(rows("products.csv")),
"palettes": len(rows("colors.csv")),
"reasoningProfiles": len(rows("ui-reasoning.csv")),
"fontPairings": len(rows("typography.csv")),
"googleFonts": len(rows("google-fonts.csv")),
"curatedIcons": len(rows("icons.csv")),
"upstreamPhosphorIcons": icons.get("iconCount"),
"uxGuidelines": len(rows("ux-guidelines.csv")),
"motionPresets": len(rows("motion.csv")),
"chartTypes": len(rows("charts.csv")),
"stacks": len(stack_paths),
"stackGuidelines": sum(len(rows(f"stacks/{path.name}")) for path in stack_paths),
}
return {
"schemaVersion": 1,
"verifiedAt": checked_date(verified_at),
"counts": counts,
"snapshots": {
name: {"sha256": digest(DATA / name)}
for name in (
"google-fonts.csv", "google-font-licenses.json",
"icons.csv", "phosphor-icons-upstream.json",
)
},
"promotionPolicy": {
"changedFamilySetRequiresExplicitApproval": True,
"relevanceGateRequired": True,
"unlicensedFamiliesExcluded": True,
},
"pendingCandidates": sorted(
({"family": item["name"], "reason": item["reason"]} for item in excluded),
key=lambda item: item["family"].casefold(),
),
}
def check_readme_counts(summary):
counts = summary["counts"]
exclusions = len(summary["pendingCandidates"])
expected = {
"README.md": (
f"**{counts['googleFonts']:,} approved Google Fonts**",
f"**{exclusions} review exclusions**",
f"**{counts['curatedIcons']} curated rows**",
f"**{counts['upstreamPhosphorIcons']:,}-icon upstream Phosphor manifest**",
),
"README.zh.md": (
f"**{counts['googleFonts']:,} 个已批准的 Google Fonts**",
f"**{exclusions} 个待审核的排除项**",
f"**{counts['curatedIcons']} 条精选记录**",
f"**{counts['upstreamPhosphorIcons']:,}-icon Phosphor upstream manifest**",
),
}
for name, tokens in expected.items():
text = (ROOT / name).read_text(encoding="utf-8")
missing = [token for token in tokens if token not in text]
if missing:
raise ValueError(f"{name} catalog counts are stale: {', '.join(missing)}")
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--verified-at")
parser.add_argument("--check", action="store_true")
args = parser.parse_args()
try:
verified_at = args.verified_at
if args.check and not verified_at:
if not OUTPUT.exists():
raise ValueError("catalog-summary.json is missing")
verified_at = json.loads(OUTPUT.read_text(encoding="utf-8")).get("verifiedAt")
if not verified_at:
raise ValueError("--verified-at is required when generating the summary")
summary = build(verified_at)
content = json.dumps(summary, ensure_ascii=False, indent=2) + "\n"
if args.check:
if not OUTPUT.exists() or OUTPUT.read_text(encoding="utf-8") != content:
raise ValueError("catalog-summary.json is stale; regenerate it")
check_readme_counts(summary)
else:
OUTPUT.write_text(content, encoding="utf-8")
except (KeyError, OSError, ValueError, json.JSONDecodeError) as exc:
print(f"generate-catalog-summary: {exc}", file=sys.stderr)
return 2
print("Catalog summary is current." if args.check else "Generated catalog summary.")
return 0
if __name__ == "__main__":
raise SystemExit(main())