Files
Epicure/SECURITY_AUDIT.md
T
Arnaud 7626f1b496 fix: close TOCTOU race + missing checks in tier limits, add AI rate limits, patch CVEs (v0.60.0)
Security audit fixes (see SECURITY_AUDIT.md):
- lib/tiers.ts: checkAndIncrementTierLimit's recipe/storage branches did a
  live count/sum check with no lock tying it to the later write, so two
  concurrent requests near the cap could both pass and both write past the
  limit. Added checkTierLimitInTransaction — holds a per-user Postgres
  advisory lock for the duration of the caller's transaction, so the check
  and the write are atomic together. Applied to every recipe/storage write
  path (recipes create/update, fork, rate, avatar, and 7 AI recipe-creation
  routes).
- Adversarial re-verification of that fix caught a bigger gap in the same
  area: four AI routes (photo import, idea generation, batch-cook,
  translate-to-new-draft) had no recipe-limit check at all. Fixed.
- Added missing per-user rate limits to 5 AI endpoints (adapt, drinks,
  pairings, translate, variations) that had none, unlike their siblings.
- search page now checks for a session server-side, matching every other
  page under (app)/ — defense in depth; the underlying API was already
  public by design and proxy.ts already blocked unauthenticated requests.
- admin/settings route now uses the shared requireAdmin instead of a local
  duplicate.
- Documented two admin support-ticket endpoints missing from the OpenAPI
  spec; verified full route/OpenAPI parity otherwise.
- Bumped drizzle-orm to fix a SQL-identifier-escaping CVE, and overrode two
  transitive deps (esbuild, postcss) with known CVEs. pnpm audit is clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 22:21:51 +02:00

9.1 KiB

Epicure Security Audit

Two-round audit: round 1 = 5 parallel domain sweeps (auth, API routes, DB/tier-limits, storage, AI/chatbot); round 2 = independent adversarial verification of every finding that survived round 1. Date: 2026-07-20.

Status: all findings below fixed and re-verified by a fresh adversarial pass (2026-07-20, later same day) — see "Fix verification" under each.

Confirmed findings

1. TOCTOU race in recipe/storage tier-limit enforcement — Medium — FIXED

File: apps/web/lib/tiers.ts (recipe/storage branches of checkAndIncrementTierLimit) Call sites: apps/web/app/api/v1/recipes/route.ts, apps/web/app/api/v1/upload/presign/route.ts, and (found during fix-verification, see below) several AI recipe-creation routes that had no check at all.

Check did a live SELECT COUNT/SUM, compared to limit, returned — no transaction or lock tied it to the later insert. Two concurrent requests near the cap could both read the pre-insert count, both pass, both write — exceeding lifetime maxRecipes/storageMb via a trivial Promise.all script.

Fix: added checkTierLimitInTransaction(tx, userId, fallbackTier, key, amount) to lib/tiers.ts — runs pg_advisory_xact_lock(hashtext(userId)) as the first statement inside the caller's db.transaction, then re-checks the live count/sum against that same transaction, before the insert/update runs. A concurrent request for the same user blocks on the lock until the first commits/rolls back, then sees its committed row. Applied as the first statement inside the write transaction in:

  • recipes/route.ts (POST — recipe count + photo storage)
  • recipes/[id]/route.ts (PUT — storage delta on photo add/remove)
  • recipes/[id]/fork/route.ts, recipes/[id]/rate/route.ts, users/me/route.ts (avatar)
  • ai/adapt/[id]/route.ts, ai/generate-meal/route.ts, ai/meal-plan/generate/route.ts, ai/translate/[id]/route.ts, ai/import-photo/route.ts, ai/generate-from-idea/route.ts, ai/batch-cook/generate/route.ts

Fix verification: a fresh adversarial pass confirmed the lock+check+write are genuinely atomic per-user in every listed call site, and also caught that four AI recipe-creation routes (import-photo, generate-from-idea, batch-cook/generate, translate/[id]) had no tier check at all (not even the old unlocked pre-flight) — an unconditional bypass of maxRecipes, broader than the original race. All four now call checkTierLimitInTransaction inside their insert transaction. Presign routes (upload/presign, upload/avatar-presign) still only do the old unlocked pre-flight check — this is fine, since storage is derived live from real rows (not a counter), so presigning alone never moves the accounted total; the authoritative, locked check now happens at write time (recipe save, rating save, avatar save) instead.

2. No rate limit on 5 AI endpoints, bypassable further by BYOK — Medium — FIXED

Files: apps/web/app/api/v1/ai/adapt/[id]/route.ts, ai/drinks/[id]/route.ts, ai/pairings/[id]/route.ts, ai/translate/[id]/route.ts, ai/variations/[id]/route.ts

None of these called applyRateLimit, unlike siblings ai/generate, ai/scale, ai/substitute, ai/recipe-chat, ai/import-url, which all throttle per-user via Redis.

Fix: added applyRateLimit(\rl:ai:${userId}`, 10, 60)` to all 5, placed immediately after the session check, matching the sibling routes' convention and rate class (one-shot single-recipe AI transforms).

Fix verification: confirmed correct placement (before any DB/AI work), correct per-user keying, and consistent limit choice against siblings. Re-swept the entire ai/ route tree — no other endpoint is missing a rate limit. BYOK still bypasses the usage quota (skipQuota: isByok) but not the rate limit itself — same behavior as every other AI route in the codebase, not a gap introduced or left open by this fix.

3. search page has no server-side session check — Low — FIXED

File: apps/web/app/(app)/search/page.tsx

Rendered SearchPageContent with zero session check, unlike sibling pages. Low-impact even before the fix: the underlying API route is intentionally public (query hard-filtered to public recipes regardless of session), and proxy.ts (the actual edge guard — middleware.ts was removed in an earlier commit, contrary to what CLAUDE.md claimed at audit time) already redirects unauthenticated requests to /login for non-public paths. The missing page-level check was a defense-in-depth/consistency gap, not a live data-leak.

Fix: added the same auth.api.getSession / if (!session) return null; check used by every sibling page. Also corrected CLAUDE.md's stale auth section to describe proxy.ts (not the removed middleware.ts) plus the per-route checks as defense-in-depth.

Fix verification: confirmed the added check matches the established convention exactly (checked against 9 sibling pages). No regression.

Minor / non-blocking — FIXED

  • apps/web/app/api/v1/admin/settings/route.ts reimplemented requireAdmin() locally instead of importing the shared lib/api-auth.ts version. Fixed: now imports and uses the shared requireAdmin. Response code for the no-session case changed from 403 to 401 (shared helper's behavior) — verified this is a correctness improvement, not a breaking change: no frontend caller branches on that specific status code, and /admin/* is already gated upstream by proxy.ts. Updated lib/openapi.ts's documented responses to include 401. Swept all other admin routes — none hand-roll their own admin check; this was the only duplicate.
  • apps/web/lib/ai/user-bio.ts's buildUserBioContext interpolates the user's own privateBio directly into the system prompt — self-injection only, not a cross-user vector. Flagged for completeness only, no fix needed.

Clean (checked, no issues found)

  • Auth/authz: role/tier always re-read from DB (requireAdmin, checkAndIncrementTierLimit, moderator-delete branch), never trusted from session cookie cache. No IDOR, no open redirect, no hardcoded secrets, no timing-attack-relevant comparisons.
  • API routes: ownership scoping consistent across recipes, meal-plans, shopping-lists, collections, comments, webhooks. No mass assignment (explicit field allowlists). SSRF closed via IP-pinning + redirect re-validation in safeFetch/validate-webhook-url.ts (covers RFC1918/loopback/link-local/multicast, IPv4 + IPv6). File uploads validate content-type/size server-side, server-generated keys (no path traversal). Auth rate limiter backed by Redis (not per-replica in-memory). Error responses sanitized, no stack-trace leakage.
  • DB/SQL: all raw sql\`usage is parameter-bound, not string-concatenated.searchroute explicitly escapes%/_/` before ilike. ~30 spot-checked update/delete sites all scope by owner or a pre-verified findFirst. Stripe webhook updates users.tier keyed by webhook-verified identifiers, not client input.
  • Storage/S3: keys always server-built (UUID-based, never from client filenames). Presigned POST (not bare PUT) with signed content-length-range/Content-Type conditions, 300s expiry, single-key scope. Ownership re-checked via isOwned*Key helpers before any writeback of a client-submitted key. Deletes only ever use DB-read-back keys after owner-scoped fetch. Secrets never reach the client bundle.
  • AI/chatbot: tool-calls (create-recipe-tool, add-to-shopping-list-tool) don't write to the DB directly — persistence goes through normal authz-checked API routes after user confirmation. No "fetch/modify by ID" tool exists to exploit. Admin AI settings allowlisted + audit-logged. Secrets masked in site-settings.ts. Non-BYOK AI calls are quota-checked and rate-limited (all AI routes, after finding #2's fix). Model output rendered as plain React text (no dangerouslySetInnerHTML in chat paths) — no XSS via model output.

Round-2 verification notes

All three confirmed findings were independently re-verified by a second agent reading the code fresh, with one correction: the original search-page write-up attributed the low severity to the API route being "auth-protected" — round 2 found that's wrong, the API is intentionally public by design, not gated. Net severity assessment unchanged.

Fix-verification round (2026-07-20, later same day)

After implementing fixes for all three confirmed findings plus the minor requireAdmin duplication, ran a second independent adversarial pass specifically against the fix diff (two parallel reviewers, one per fix area). Outcome: findings #2, #3, and the minor requireAdmin item confirmed cleanly fixed with no regressions. Finding #1's fix was confirmed correct for every call site it touched, but the same pass caught a broader pre-existing gap in the same area — four AI recipe-creation routes with no tier-limit check whatsoever — which has since also been fixed and included in the "Fix" description above. No further issues found; typecheck, lint, full test suite (191/193 passing — the 2 failures are pre-existing and unrelated to this session, see git history), and production build all pass clean.