# CURIO — PRODUCT & BUILD SPECIFICATION *An AI mastery tutor. Follow your curiosity.* --- ## 1. OVERVIEW Curio is a curiosity-led learning platform where a learner states what they want to learn and receives a personalized **Journey** of short, bounded **Lessons**. Lessons are AI-generated, AI-verified, and — crucially — **AI-tutored**: the learner does not passively consume content, they *produce* (predict, explain, solve), and the AI diagnoses what they misunderstand and adapts. The product is built on one idea: **a generative model can do what a course or a feed cannot — respond to the individual learner.** Curio is organized entirely around that capability. --- ## 2. CORE CONCEPT & DIFFERENTIATOR Curio is not a course catalog and not a content feed. Learning is **generated, not selected**, and **diagnosed, not just delivered**. The defensible capability — the moat — is **misconception diagnosis**: inferring *what* a learner misunderstands and *why* from what they actually say, and targeting exactly that. Generating explanations is now commodity; diagnosing a specific learner's wrong mental model and correcting it is not. **The product is built around the misconception model, not the content stream.** --- ## 3. THE LEARNING UNIT: THE LESSON The atomic unit is a **Lesson**: a self-contained microlearning unit, completable in one sitting (~10 minutes), built from 2–4 **segments**. ``` Lesson (~10 min, bounded, one sitting) ├─ Segment 1: short readable explanation (~300–500 words, reads like an article) │ └─ Checkpoint: learner produces (predict / explain-back / solve) │ └─ AI diagnosis + targeted feedback (live) ├─ Segment 2: explanation → Checkpoint → diagnosis └─ Segment 3: explanation → Checkpoint → diagnosis → Mastery gate ``` Design rules: - **Reads like an article, not a quiz.** Explanation prose is continuous and calm; checkpoints *punctuate* (2–4 per Lesson) rather than interrupting every sentence. - **Bounded.** Clear start and end, finishable in ~10 minutes. The session cap bounds learner attention *and* live token cost. - **Active within the bound.** A passive and an active 10-minute session cost the same minutes; the active one retains far more. Microlearning brevity is preserved; passivity is removed. - **Open-ended where possible.** Checkpoints favor free-text production over multiple choice — free text is what makes diagnosis possible. A **Journey** is an ordered sequence of Lessons for a topic, with adaptive branching driven by the memory engine. --- ## 4. THE TUTORING LOOP Per checkpoint: 1. **Read** — a short AI-generated explanation segment. 2. **Produce** — the learner predicts, explains in their own words, or solves a small problem. 3. **Diagnose** — the AI evaluates the *actual response* against a pre-verified rubric, reference answer, and per-concept misconception library; it names the specific misunderstanding. 4. **Adapt** — targeted feedback; re-explain differently if it didn't land; escalate difficulty if the learner is solid. 5. **Gate** — advancement is earned by demonstrated production, not by clicks. The outcome feeds the memory engine (§11). --- ## 5. CURIOSITY RAILS The entry point is open: *"What do you want to learn today?"* Learners may pull tangents mid-Lesson ("wait, why does that happen?"). A tangent spawns a short sub-explanation and then **returns to the checkpoint**, so exploration never decays into aimless dabbling. Curiosity is encouraged; it is always brought back onto the attempt→diagnose loop. --- ## 6. LEARNING MODES - **Journey mode** — structured progression through a topic's Lessons with adaptive branching and mastery tracking. - **Daily Review** — surfaces concepts due for reinforcement from the spaced-repetition schedule. This is the substantive retention loop and the intended source of habit, in place of dopamine-driven gamification. --- ## 7. DESIGN LANGUAGE & EXPERIENCE Curio should feel **modern, human, inspired, and beautiful** — not as decoration, but because the product is a reading-and-thinking environment, and the craft of that environment is part of the teaching. The interface should feel like a brilliant, warm tutor who has set your work down in front of them and is reading it carefully. ### 7.1 Thesis The experience is **a calm reading surface that quietly comes alive when you respond.** Reading is unhurried and editorial; the moment of diagnosis is where the product breathes. Nothing competes with those two states. This is explicitly *not* a gamified, dashboard-dense, or chatbot-bubble app. ### 7.2 Signature: marginalia The one memorable, bold element is the **diagnosis-as-marginalia** moment: when a learner submits a checkpoint response, the tutor's feedback arrives threaded against *their own words* — the way a great teacher annotates your answer in the margin, pointing at the exact phrase where the misconception lives (the `evidence_span` from §10.3). The feedback references what the learner actually wrote. Spend the design's boldness here; keep everything around it quiet. ### 7.3 Direction (a starting point, not a mandate) Build by working the two-pass design process (plan a token system → critique it against this brief for genericness → then build). The following is the intended direction; refine it, don't default to it. Deliberately avoid the current AI-design clichés (cream + high-contrast serif + terracotta; near-black + acid accent; broadsheet hairline columns). - **Palette — "ink & light."** A soft, slightly cool paper (not cream) as the reading ground; an inky near-black for text; a considered deep indigo as the focus/brand accent (depth, study); a warm ochre reserved *only* for the human tutor moments (marginalia, encouragement). Diagnosis states are calm and legible, never an alarm: mastered = a quiet sage/teal, partial = ochre, misconception = a soft clay/rose that says *let's look again*, never buzzer-red. A genuinely comfortable dark mode (warm charcoal, not pure black). - **Typography — reading is the product.** A humanist, low-contrast **reading serif** for body (e.g. Newsreader / Source Serif / Literata as candidates), set at a generous measure (~62–68ch) and line-height (~1.6); a characterful **display face used with restraint** for the prompt and headings; a humanist **mono** for mastery-map data and labels. Make the type a memorable part of the page, not a neutral delivery vehicle. - **Motion — meaning, not decoration.** No animation library. Use CSS transitions and the **View Transitions API**. Reserve the one orchestrated motion moment for diagnosis: a brief "considering" beat, then feedback easing into the margin — it should feel like a person responding, not a form validating. Everything else is still. `prefers-reduced-motion` is honored everywhere. - **Layout — one thing at a time.** Single reading column, generous margins, minimal chrome. The checkpoint is inset in the reading flow; feedback appears as marginalia, not a modal. The mastery map is a separate, richer view — it does not crowd the learn surface. ### 7.4 Voice (copy is design material) Feedback copy *is* the tutor's character. Warm, specific, plain, and on the learner's side — names the exact thing that tripped them, never condescending, never gamified ("Not quite — the bit that slipped is *…*", never "❌ Wrong! −10 XP"). Active voice; actions say what they do. Empty and error states give direction in the interface's voice, not apologies. ### 7.5 Quality floor (non-negotiable) Responsive to mobile; visible keyboard focus; reduced motion respected; WCAG-AA contrast on text and on every diagnosis-state color. Loading is part of the experience — the cold-start outline (§8.3) is a designed state, never a spinner. ### 7.6 Anti-patterns No XP counters, badges, streak confetti, or leaderboards; no dense dashboard on the learn surface; no generic component-library default look; no chat-bubble UI; no autoplay or attention-grabbing motion. --- ## 8. CONTENT ARCHITECTURE Curio is a **growing library of validated, generated content that gets personalized** — not a pure live-generation engine. The first learner to request a topic pays the generation cost; everyone after reuses validated content. This makes real-time UX, accuracy, and cost simultaneously feasible. ### 8.1 Three tiers 1. **Blueprint** — the canonical, fully verified Journey for a normalized intent: an ordered list of concepts with Lesson templates. Generated and verified **once**, reused by all learners. 2. **Variants** — pre-generated explanation/checkpoint bodies per concept at 2–3 difficulty levels, generated and verified at blueprint time. 3. **Live** — for long-tail intents with no blueprint hit: generated fresh through a lighter synchronous gate, then queued for full verification and **promoted into the blueprint cache** so the next learner gets a hit. Two cost classes follow: | Cached + verified once, reused by everyone | Live, per learner (pay tokens exactly here) | |---|---| | explanation segments | evaluation of the learner's free-text response | | checkpoints + reference answers + rubrics | misconception diagnosis + targeted feedback | | worked examples | adaptive re-explanation | | per-concept misconception libraries | tangent handling | ### 8.2 Intent normalization (drives cache hit rate) "Learn Python", "teach me python basics", and "i wanna learn python" must collapse to one key: - **Structured key:** one cheap classification call → `{domain, subtopic, goal_type: skill|exam|concept, level_hint}` → canonical string. - **Embedding fallback:** embed the raw intent; pgvector nearest-neighbor search against existing blueprint embeddings; reuse if cosine > threshold. - Miss only when both fail. `goal_type` and `level_hint` are **parameters into** a blueprint, not separate blueprints. ### 8.3 Serve flow (skeleton-first + buffer) Perceived latency is won by streaming *structure* before *content*: the Journey outline appears in under a second; the learner can't move faster than the first segment arrives. ``` on submit(intent): key, ctx = normalize(intent) # ~100ms bp = blueprint_cache.lookup(key, ctx) # exact key, then embedding NN if bp: # WARM lesson = hydrate(bp, next_concept, user) # variant SELECTION, ~150ms buffer.push(session, lesson); stream(lesson) else: # COLD bp = gen_blueprint(intent, ctx) # structured call → concept list stream(bp.outline) # outline visible now (designed state) lesson = gen_lesson(bp, concept[0], rag) # light sync gate before serve buffer.push(session, lesson); stream(lesson) enqueue(promote_to_blueprint, bp) # full verification + cache, async on advance / buffer_low(session): enqueue(generate_next, session, next_concept) # background job ``` The serve path reads from a Redis buffer fronting Postgres; a background job refills it. Concurrent generations per session are capped for backpressure. Postgres remains the source of truth — a Redis outage degrades latency, not availability. --- ## 9. AI GENERATION PIPELINE Generation is **RAG-grounded** to constrain hallucination at the source. Per concept the generator produces, all grounded in retrieved trusted sources and stored with provenance (`source_chunk_ids`): - explanation segments - checkpoints (prompts) + **reference answers** + **grading rubrics** - worked examples - a **misconception library**: likely wrong mental models for the concept, each with a tag and a diagnostic signature Rubrics and misconception libraries are first-class generated artifacts — they are what make live grading constrained and verifiable. All prompts live in a versioned registry; generation output is **structured (schema-constrained) JSON** via the model's native structured-output mode. --- ## 10. AI VERIFICATION SYSTEM Curio verifies with AI on two surfaces. Both are made reliable by the same discipline: **ground the check against real artifacts, decorrelate from the generator, structure the output, and allow abstention.** ### 10.1 Content verification — the cascade A naive "second LLM checks the card" shares the generator's blind spots and rubber-stamps confident-wrong content. Instead: - **Grounded, not free-judged.** The verifier checks each atomic claim against the retrieved source chunks ("is this supported by these sources?"), not against its own priors. - **Independent solving.** The verifier never sees the marked answer; it **solves checkpoints/quizzes blind** (n-sample, majority vote) and compares to the reference answer. Mismatch → flag; sample disagreement → uncertainty. - **Decorrelation.** A different model/provider and an adversarial prompt, so failure modes don't overlap the generator's. Run as a **cascade**, so expensive scrutiny is amortized: | Tier | Runs | Cost | Checks | |---|---|---|---| | **T0 Structural** | every item, always | ~free | schema valid; exactly one correct option; reference answer references a real option; refs resolve; non-empty; length bounds | | **T1 Cheap-model** | every item at generation | low | claim grounding; blind solve; obvious safety. Clears the easy majority, else escalate | | **T2 Adversarial** | blueprint promotion + escalations + low confidence | medium, **amortized once per blueprint** | full rubric, claim-by-claim grounding, n-sample blind solve, pedagogical soundness, cross-item consistency | | **T3 Human** | top-traffic templates + persistent uncertainty | high, rare | spot review, golden-set labeling | Because content is cached at the template level, T2 runs **once per blueprint** and is reused by thousands of learners — near-zero per-serve cost. ### 10.2 Response-grading verification The AI also grades free-text learner responses. Making this trustworthy: - **Grade against pre-verified artifacts.** At runtime the grader checks the response against the already-verified rubric + reference answer + misconception library — constrained, not freeform. - **Misconception matching.** The grader maps the response to a known misconception tag or flags `novel` (→ escalate). Closed-set classification beats open-ended critique. - **Two-sided error awareness.** A false-pass wastes learning; a **false-fail (telling a correct learner they're wrong) corrodes trust and is worse.** The grader is biased toward generous grading and always reveals its reasoning so the learner can contest. - **Abstention + fallback.** The grader may return `uncertain` → reveal the model answer for self-assessment rather than confidently mis-grading. - **Cascade.** A cheap model grades the rubric-matched majority; a strong model handles `novel`/`uncertain` only. ### 10.3 Output contracts Content verifier: ```json { "verdict": "pass | fail | uncertain", "claims": [{"text": "", "status": "supported|unsupported|contradicted", "evidence_chunk_ids": ["..."]}], "solve_check": {"independent_answers": ["B","B","C"], "majority": "B", "reference": "B", "agreement": 0.67, "status": "match|mismatch|ambiguous"}, "pedagogy": {"teaches_concept": true, "prereqs_respected": true}, "safety": {"in_scope": true, "mode_compliant": true}, "confidence": 0.0 } ``` Response grader: ```json { "verdict": "mastered | partial | misconception | uncertain", "rubric_hits": ["criterion_id"], "misconception_tag": "tag | novel | none", "evidence_span": "", "feedback_directive": "advance | re-explain: | probe:", "confidence": 0.0 } ``` Contracts are defined as **Zod schemas** shared between the LLM layer, the API, and the frontend. ### 10.4 Keeping verifiers honest - **Golden sets** of human-labeled cases for both surfaces; measure precision/recall continuously, tracking **false-fail rate** separately for grading. A model upgrade triggers an offline re-sweep of the cached library. - **User-signal outer loop (free):** collective checkpoint failure (many learners miss the same item → suspect the reference answer), a report button, and dwell-time anomalies auto-trigger re-verification. - **Versioning:** content keyed by `(intent_key, content_version, model_version)`; invalidation is lazy; a flagged item regenerates just itself. --- ## 11. MEMORY ENGINE A **mastery-tracking engine**, not a knowledge graph: - **Per-concept mastery score** updated from demonstrated production at checkpoints and from misconception tags, via Bayesian Knowledge Tracing or an Elo-style update. - **Spaced-repetition schedule** per concept (SM-2-like) driving Daily Review. This drives difficulty scaling, review timing, reinforcement, and the mastery gate. A relational knowledge graph is deferred until there is data to justify it. --- ## 12. PEDAGOGY The design implements the mechanics behind the **2-sigma** finding (1:1 tutoring + mastery learning outperforms classroom instruction by ~two standard deviations) — exactly what generative AI can deliver and a feed structurally cannot: - **Active production / retrieval practice** at every checkpoint. - **Mastery gating** before advancement. - **Spaced repetition** via Daily Review. - **Interleaving** of related concepts rather than blocking. --- ## 13. PERSONALIZATION - **Variant selection** by mastery: pick the pre-generated difficulty variant matching the learner's score. Deterministic, sub-millisecond, zero live LLM cost. - **Live personalization** is the feedback and re-explanation themselves — which is the core value, not overhead, and is where live tokens are deliberately spent. --- ## 14. SAFETY & QUALITY CONTROL Three strictness modes — **open**, **controlled** (default), **enterprise-safe** — applied as parameters to generation, content verification, and response grading (including how off-topic or unsafe free-text input is handled). Filtered domains are enforced at intent normalization and at T0/T1. --- ## 15. TELEMETRY & EVALUATION Because the product *is* diagnosis, instrumentation is a first-class requirement, not an afterthought: - **Diagnosis telemetry:** every `response → grade` is logged; track false-fail rate, misconception-hit rate, and mastery gain per concept. You cannot improve a diagnosis engine you don't measure. - **LLM observability:** per-call tracing of latency, tokens, and cost (Langfuse or OpenTelemetry). Without it the false-fail and cost metrics are aspirational. - **Eval harness in CI:** golden sets for content verification and response grading run on every prompt change. Prompts are code — versioned and tested. The harness doubles as the build's definition of "good enough." - **Cost/latency guards:** per-session token cap (the Lesson bound enforced in code), a provider circuit breaker, an **edge rate limiter on the cold-start generation endpoint** (it spends money on demand), and graceful degradation to abstain→reveal. --- ## 16. FEATURE SET **Core (in the product):** intent entry; Journey mode; bounded Lessons with checkpoints; live diagnosis with abstention; mastery + spaced repetition; Daily Review; mastery map (learner-facing view of mastered / shaky / due); source citations (surfacing the provenance already stored). **Loop reinforcers (near-term):** confidence rating before reveal (catches *confidently wrong*); Socratic follow-up probe on a `partial` verdict. **Deferred:** learn from your own material (PDF/notes → personalized Journey — strong differentiator, leverages existing RAG infra, but a large surface); voice. **Explicitly out of scope:** XP / badges / leaderboards, infinite feed, social features, multi-provider fallback orchestration, knowledge graph. Microlearning + Daily Review + mastery map supply motivation intrinsically; bolt-on gamification risks reintroducing passive, low-retention "junk-food" learning. --- ## 17. DATA MODEL ``` source_chunk(id, doc_ref, embedding, text) blueprint(id, intent_key, embedding, title, model_version, content_version, status[draft|verifying|published|flagged]) concept(id, blueprint_id, ord, name, retrieval_ctx_ref) lesson(id, blueprint_id, ord, est_minutes) segment(id, lesson_id, ord, body_json, source_chunk_ids[], verify_status, content_version) checkpoint(id, segment_id, prompt, kind[predict|explain|solve], reference_answer, rubric_json, verify_status) misconception(id, concept_id, tag, signature, verify_status) verify_report(id, target_id, target_type, verdict, confidence, payload_json, model_version) user(id, ...) journey_instance(id, user_id, blueprint_id, position) mastery(user_id, concept_id, score, last_seen, next_review) response(id, user_id, checkpoint_id, text, ts) grade(id, response_id, verdict, misconception_tag, confidence, payload_json) job(id, type, idempotency_key, status, payload_json, attempts, created_at) -- durable ledger for paid background work -- Redis: buffer:{session} -> [hydrated lesson/segment ids]; generation-concurrency locks ``` --- ## 18. TECH STACK **Single language, end to end (TypeScript).** One type system, one toolchain, and request/response/LLM contracts shared across the boundary. Nothing here requires a Python ML runtime — the work is hosted-LLM calls, embeddings, and pgvector queries, all first-class in TypeScript. - **App framework:** Next.js (App Router, TypeScript) full-stack — React Server Components for the reading surface, Route Handlers for the API, native streaming/SSE for skeleton-first delivery. Split out a thin service (e.g. Hono) only if the API outgrows Next. - **LLM orchestration:** Vercel AI SDK — provider-agnostic, native streaming and **schema-constrained structured output** (Zod). Generator and grader configured to **different providers/models** (decorrelation). Exploit two cost levers: native structured output for verifier/grader contracts, and **prompt caching** for the large, session-reused grounding context (sources + rubric + misconception library). - **Contracts/validation:** Zod, shared across API I/O, LLM outputs, and DB-derived types. - **Database:** PostgreSQL + pgvector (single store, source of truth). - **ORM & migrations:** Drizzle ORM + drizzle-kit (type-safe, SQL-first, clean pgvector support). - **Cache / buffer:** Redis — a cache and serve-buffer in *front* of Postgres, never load-bearing for availability. - **Background work:** for the vertical slice, Route Handlers + the `job` ledger table suffice. When scale work arrives: **BullMQ** (Redis-backed, TS), with **idempotency keys and durable Postgres job state** — these jobs spend money, so a retry must never regenerate-and-rebill. - **Styling & motion:** Tailwind CSS + CSS custom properties for design tokens; motion via CSS transitions and the **View Transitions API** (no animation library). `prefers-reduced-motion` respected. - **Observability:** Langfuse (or OpenTelemetry) for LLM call tracing; standard app metrics/logs. - **Rate limiting:** edge limiter (e.g. Upstash Ratelimit on Redis) on the cold-start generation endpoint. - **Testing:** Vitest (unit; LLM mocked) + a golden-eval runner (real-model, CI-gated) + Playwright for the core learn-flow e2e. - **Hosting:** Next on Vercel (or Node host); managed Postgres+pgvector (Neon / Supabase / RDS) and managed Redis (Upstash). Keep the cold-start generation path on a runtime with adequate timeout (Node serverless/long-running, not edge). --- ## 19. COST MODEL - **Scaffolding** (explanations, checkpoints, rubrics, misconception libraries) is generated and T2-verified **once per blueprint**, amortized to near-zero per serve. - **Live cost** is grading + feedback turns: per Lesson ≈ 2–4 checkpoints × (grade + feedback) + occasional re-explanation/tangent. The ~10-minute Lesson bound **directly caps live turns per session** — microlearning bounds cost as well as attention. - **Levers:** a grading **cascade** (cheap model for the rubric-matched majority, strong model only on `novel`/`uncertain`) and **prompt caching** of the reused grounding context keep per-turn cost low. - Net: with a high cache-hit rate, live LLM spend concentrates on cold long-tail topics (throttleable) and on the irreplaceable act of responding to the learner. Model this against the chosen provider's pricing at assumed hit rates and checkpoint counts before building. --- ## 20. MVP SCOPE & BUILD ROADMAP **The MVP answers one question:** *can an AI tutor diagnose what a learner misunderstands and fix it, well enough that they learn and return?* **MVP loop:** prompt → one Lesson (2–3 segments) → read + open-text checkpoints → live grading against verified rubric + misconception library → mastery gate. **Phases:** - **P0 — Scaffold & infra.** Next.js + TypeScript + Tailwind app; docker-compose (Postgres+pgvector, Redis); Drizzle schema + migrations; Vitest; provider-agnostic LLM layer (Vercel AI SDK) with a mockable client; design tokens stubbed. - **P1 — Vertical slice (single hardcoded topic).** Content schema + seeded corpus + pgvector retrieval → RAG-grounded Lesson generation (segments, checkpoints, reference answers, rubrics) → serve + render the reading surface and checkpoint → **response grading with abstention** → **golden eval harness gating false-fail rate in CI.** Build the marginalia feedback moment here — it is the core experience. This milestone validates diagnosis quality. - **P2 — Diagnosis depth.** Misconception-library quality + the full content-verification cascade. - **P3 — Adaptivity & scale.** Mastery + spaced repetition; blueprint cache, intent normalization, BullMQ workers with idempotent paid jobs. - **P4 — Roadmap features**, one at a time: confidence rating, Socratic probe, Daily Review, mastery map, source citations, variant personalization, accounts. Anonymous-session-first; accounts arrive in P4. --- ## 21. OPEN DECISIONS 1. **Checkpoint count per Lesson** — fixed (e.g., 3) or adaptive to mastery? Sets the cost ceiling and the est-minutes contract. 2. **False-fail tolerance** — how generous is the grader, and when does it abstain vs. reveal? Sets the trust/rigor balance; prototype with real responses. *Settle before writing the grader.* 3. **Cold-start consistency** — may a cold long-tail Lesson be graded against a T1-only rubric (full T2 async), or does first serve block on T2? 4. **Concept granularity** — one concept per Lesson, or per Lesson-cluster? 5. **Source corpus policy** — define allowed, non-infringing sources before generation begins. --- ## 22. NORTH STAR A learner types *"Teach me astrophysics."* A short, readable Lesson opens — finishable in ten minutes, set in calm, beautiful type. They read a few hundred words, then are asked to predict what happens to a collapsing star. They type a guess. The tutor reads it, pauses for a beat, and writes back in the margin — pointing at the exact phrase where their thinking went wrong, explaining the missing piece a different way, and asking again. They get it. Mastery updates; the next review is scheduled. Every explanation, problem, and rubric was generated and verified once and reused — only the part that read their actual answer ran live, which is the whole point.