19 KiB
Zotero KB Design
Overview
This project builds a local Zotero-to-card pipeline for agent-friendly research workflows. It reads a local Zotero data directory, converts selected literature into lightweight knowledge cards, supports project-scoped writing assistance, and prepares a file-based integration path for Claude Code / Codex via SKILL.
The first implementation pass prioritizes a local API + Web service, not a Zotero-native product. Zotero remains the source of truth for bibliographic data and attachments. The generated card library becomes the reusable asset layer for humans and agents.
Goals
- Read Zotero data directly from a local data directory such as
/mnt/c/Users/WSX/Zotero - Support importing the literature currently selected in Zotero into the local system
- Preserve Zotero collection hierarchy in the generated library structure
- Convert literature into lightweight Markdown knowledge cards with a JSON index
- Support project creation and project-scoped literature management
- Support adding and removing literature from a project at any time
- Support two independent writing-mode actions:
- citation candidate recommendation
- first-pass citation plan generation
- Prepare a file-based SKILL integration so Claude Code / Codex can read only the content of a specified project
Non-Goals For V1
- Full knowledge-base mode UI or workflows across the entire library
- General-purpose chat over all literature
- Final paper text generation
- Editing Zotero data in place
- Replacing Zotero as the bibliographic source of truth
Knowledge-base mode is intentionally deferred. V1 only reserves the underlying global library structure and API surface needed for future implementation.
Context And Constraints
Local Zotero Access
The user confirmed that the Zotero data directory is available locally and contains both zotero.sqlite and storage/.
V1 uses a mixed ingestion model:
- Read bibliographic data, notes, tags, collections, and attachments from the local Zotero data directory
- Use a minimal Zotero bridge only to retrieve the keys of the items currently selected in the Zotero UI
This split is required because direct filesystem access can read the library contents but cannot reliably observe Zotero's current UI selection state.
Reference Projects
zotero-ragis the main implementation reference for local web service structure, attachment extraction, and search-oriented pipeline compositionzotcardis the main reference for the Zotero-side bridge pattern and selected-item access inside the Zotero desktop runtime
The new system is intentionally not identical to either:
- unlike
zotero-rag, the durable asset is a card library rather than only a retrieval index - unlike
zotcard, the primary product surface is a local service and file-based workspace rather than a Zotero-first note plugin
Product Model
The system has two storage scopes with different purposes.
1. Global Library Scope
This stores the full set of imported literature and generated cards across all imports. It is the foundation for the future knowledge-base mode.
Responsibilities:
- maintain the canonical local card library
- preserve collection hierarchy across all imported items
- store normalized item metadata and card indexes
- avoid regenerating duplicate cards for the same Zotero item unless the source content changes
2. Project Scope
Projects are scoped working sets built on top of the global library. A project contains references to selected literature rather than owning the canonical copy of every card.
Responsibilities:
- define which literature belongs to the project
- support adding and removing selected items at any time
- provide a project-scoped view for writing assistance
- constrain Claude Code / Codex to read only the project's content
This split avoids redundant card copies while preserving future flexibility for project-specific views and exports.
Architecture
The service is divided into five units.
1. Zotero Bridge
Purpose:
- ask Zotero for the currently selected item keys
Characteristics:
- minimal surface area
- no card logic
- no library parsing
- no knowledge of projects
The bridge can be a small Zotero plugin or local bridge script embedded into a plugin package. Its only required output is a list of selected item_key values.
2. Zotero Reader
Purpose:
- read local Zotero data from
zotero.sqliteandstorage/ - load item metadata, creators, abstract, tags, notes, collections, and attachments
- extract attachment text for supported files
Characteristics:
- local-data-first
- deterministic
- independent from UI
3. Card Builder
Purpose:
- normalize raw Zotero material into a
source bundle - call a configurable LLM to transform source material into an agent-readable card
- write Markdown cards and update JSON indexes
Characteristics:
- one card per Zotero regular item
- source-aware hash for rebuild detection
- card schema stable enough for direct SKILL consumption
4. Project Manager
Purpose:
- create projects
- track which Zotero items belong to which project
- expose project-scoped listings and stats
- allow project item removal without affecting the global library
5. Writing Service
Purpose:
- run project-scoped writing assistance over project cards
V1 includes two independent actions:
- recommend citation candidates
- generate a first-pass citation plan
These actions consume project-scoped card views and never read outside the target project.
Directory Layout
The workspace uses a shared library plus project views.
workspace/
library/
collections/
Theory/
Subtopic/
Paper A [ABCD1234].md
index/
items.json
cards.json
collections.json
cache/
source-bundles/
ABCD1234.json
projects/
thesis-ch2/
project.json
selected-items.json
project-index.json
Library Scope Files
library/collections/: canonical Markdown cards arranged to mirror Zotero collection hierarchylibrary/index/items.json: normalized item metadata keyed by Zotero item keylibrary/index/cards.json: card metadata used for filtering, retrieval, and writing-mode preparationlibrary/index/collections.json: normalized collection tree and item membershiplibrary/cache/source-bundles/: raw normalized source material captured before LLM card generation
Project Scope Files
projects/<project-id>/project.json: project metadata and model configurationprojects/<project-id>/selected-items.json: ordered set of Zotero item keys assigned to the projectprojects/<project-id>/project-index.json: precomputed project view for fast loading and SKILL consumption
Projects do not need to duplicate every canonical card. The project index maps selected items to canonical card paths in the global library.
Data Model
Project Metadata
Example:
{
"id": "thesis-ch2",
"name": "Thesis Chapter 2",
"zotero_data_dir": "/mnt/c/Users/WSX/Zotero",
"selection_mode": "zotero-bridge",
"llm": {
"provider": "openai",
"model": "gpt-5-mini",
"base_url": null
},
"created_at": "2026-04-15T00:00:00Z"
}
Item Index Entry
library/index/items.json stores one normalized entry per Zotero item.
Required fields:
item_keytitlecreatorsyearitem_typeabstracttagscollection_pathsnote_idsattachment_keysattachment_statuscard_pathsource_hashupdated_at
Card Index Entry
library/index/cards.json stores fields optimized for reading and writing assistance.
Required fields:
item_keycard_pathtitlesummarykeywordsclaimsquotable_spanswriting_hintsupdated_at
Collection Index Entry
library/index/collections.json stores normalized collection information.
Required fields:
collection_keynameparent_keypathitem_keys
Card Format
Each regular Zotero item generates one Markdown card in the canonical library.
Example:
---
item_key: ABCD1234
title: Retrieval-Augmented Writing in Humanities Research
year: 2024
collections:
- Theory/Subtopic
authors:
- Alice Smith
- Bob Li
tags:
- llm
- retrieval
attachment_status: ok
source_hash: "sha256:6f6f5d9de1b5a9b98d7ce1c76d7a8d5c441d72e3"
---
# Summary
This paper argues that card-oriented retrieval pipelines improve citation-grounded drafting when notes, metadata, and attachment text are merged before synthesis.
# Core Claims
- Combined metadata, notes, and full-text evidence produce more reliable citation suggestions than metadata-only indexing.
- Project-scoped reading reduces irrelevant retrieval during drafting.
# Methods
- Compares card-generation inputs across three corpus preparation strategies.
- Evaluates writing-support quality with citation recommendation tasks.
# Evidence
- Reports higher citation precision for the merged-input pipeline.
- Includes examples where note-only systems miss relevant methodological context.
# Quotable Passages
- p.12: "Project-scoped card retrieval improves citation precision during drafting."
- Note: Author stresses that project-bounded retrieval lowers topical drift.
# Writing Hints
- Useful for defining the problem
- Useful as a supporting citation for claims about scoped retrieval during drafting
The card format is intentionally designed for:
- direct human reading
- direct SKILL consumption
- predictable parsing by the Web service
Import Flow
Selected-Items Import
- The user selects multiple items in Zotero
- The Zotero bridge returns the selected
item_keyvalues - The service reads the local Zotero database and storage directory
- The reader collects metadata, notes, tags, collection paths, and attachment text
- The builder creates or updates source bundles
- The builder generates or refreshes Markdown cards and indexes
- The project manager adds the selected item keys to the target project
- The project view is rebuilt so Web and SKILL consumers can read the updated subset
Add / Remove Semantics
- Adding is
upsertbyitem_key - A Zotero item can belong to multiple projects
- Removing an item from a project only updates that project
- Removing an item from a project never deletes it from Zotero
- The canonical card remains in the global library unless explicit future garbage collection is added
LLM Configuration
Card generation is model-driven and configurable via API at the project level.
V1 requirements:
- support configurable provider and model fields
- support prompt templates for card generation
- keep the source bundle on disk before generation
- treat LLM output as structured content that is converted into canonical Markdown sections
The service should not hardcode a single provider design. It should expose a provider abstraction with one stable card-generation contract.
Modes
Writing Mode
Writing mode is project-scoped and exposes two independent actions, each implemented as a dedicated skill.
Skill 1: Citation Recommendation
Skill name: zotero-citation-recommender
Purpose: Given a writing intent or draft paragraph, recommend relevant literature from the project scope.
Input:
- writing intent text or draft paragraph
Output:
- structured list of recommended literature
- why each item is relevant
- card-derived claims or quotable spans
- suggested rhetorical role such as:
- definition
- supporting evidence
- contrast
- limitation
API endpoint: POST /api/projects/{project_id}/writing/recommend-citations
Skill 2: Citation Plan Generation
Skill name: zotero-citation-planner
Purpose: Generate a first-pass citation plan for a writing intent, suggesting paragraph structure, citation order, and the role of each cited work.
Input:
- writing intent text
- optional stance
- optional paragraph goal
- optional length expectation
Output:
- paragraph-level citation plan
- suggested sub-structure
- citation order
- recommended role for each cited work
- notes about which claims should be grounded by which items
This skill generates a plan, not polished final prose.
API endpoint: POST /api/projects/{project_id}/writing/generate-plan
Knowledge-Base Mode
Knowledge-base mode is reserved for future implementation and is not part of V1 delivery.
V1 only reserves:
- the global library storage model
- future-facing API stubs
- indexes suitable for full-library browsing and retrieval later
API Surface
The initial service should expose the following HTTP endpoints.
Project Endpoints
POST /api/projects- create a project
GET /api/projects- list projects
GET /api/projects/{project_id}- get project detail and counts
Import Endpoints
POST /api/projects/{project_id}/imports/selected-items- read selected item keys from the Zotero bridge
- ingest source data from the local Zotero directory
- generate or update cards
- add the imported items to the project
Project Item Endpoints
GET /api/projects/{project_id}/items- list project items
DELETE /api/projects/{project_id}/items/{item_key}- remove an item from the project
Card Endpoints
GET /api/projects/{project_id}/cards- list cards visible to the project
POST /api/projects/{project_id}/cards/rebuild- rebuild all cards in the project scope
POST /api/projects/{project_id}/cards/{item_key}/rebuild- rebuild one card
Writing Endpoints
POST /api/projects/{project_id}/writing/recommend-citations- return project-scoped citation candidates
POST /api/projects/{project_id}/writing/generate-plan- return a project-scoped first-pass citation plan
Future Knowledge-Base Endpoints
These routes are reserved in V1:
GET /api/library/cardsGET /api/library/collectionsPOST /api/library/search
Web UI
The Web UI is an operational console, not the primary agent interface.
Left Panel
- project switcher
- create project action
- import selected Zotero items action
- project collection tree
Center Panel
- project card list
- filters by title, author, tag, year, and keyword
- card detail viewer
- remove-from-project action
- rebuild-card action
Right Panel
- writing mode tabs
- recommendation
- citation plan
- structured input form
- structured output display
The UI should remain intentionally small. It exists to manage and inspect the workspace rather than replace the card files as the main knowledge surface.
Error Handling
The system should degrade gracefully and preserve intermediate assets where possible.
Bridge Failure
- if the Zotero bridge cannot read current selection, return an explicit import error
- do not affect existing projects or cards
Attachment Failure
- if attachment text extraction fails, still build a partial card from available metadata, abstract, notes, and tags
- mark
attachment_statusasmissingorpartial
LLM Failure
- keep the source bundle on disk
- mark the item as card generation failed
- allow single-item or project-level retry
Source Drift
- detect changes through
source_hash - mark cards stale when Zotero source content changes
- surface rebuild-needed state in project views
Testing Strategy
V1 should focus on deterministic tests around boundaries rather than broad UI snapshots.
SQLite Reader Tests
- parse items, creators, notes, tags, and collection relationships from fixture data
- resolve attachment mappings from the Zotero schema used by the local directory
Project Management Tests
- create projects
- add selected items
- deduplicate repeated imports
- remove project items without mutating the canonical library
Card Builder Tests
- normalize source bundles
- verify prompt input assembly
- stub model output
- verify Markdown card writing and JSON index updates
- verify stale detection using
source_hash
Writing Service Tests
- citation recommendation returns the expected structured shape
- citation plan generation returns the expected structured plan shape
- both actions remain project-scoped
API Tests
- validate project creation, import, removal, rebuild, and writing endpoints
SKILL Integration
The agent integration is file-based. Claude Code / Codex should read only the selected project's content.
Two writing-mode skills are provided:
skills/zotero-citation-recommender/SKILL.md— handles citation candidate recommendationskills/zotero-citation-planner/SKILL.md— handles first-pass citation plan generation
A project reader skill is also provided:
skills/zotero-project-reader/SKILL.md— core project and card reading
Responsibilities of each skill:
zotero-project-reader:
- read
projects/<project-id>/project.json - read
selected-items.jsonandproject-index.json - load only the canonical cards referenced by the project
zotero-citation-recommender:
- read project-scoped cards
- accept writing intent or draft paragraph as input
- return structured citation recommendations with relevance rationale, claims, and rhetorical roles
zotero-citation-planner:
- read project-scoped cards
- accept writing intent and optional stance/goal/length as input
- return paragraph-level citation plan with structure, order, and per-work roles
Each skill must not access the entire global library unless the target project explicitly references those items.
Implementation Priorities For V1
- Build the local workspace structure and project model
- Implement Zotero local-data reading from
zotero.sqliteandstorage/ - Implement the minimal Zotero selected-items bridge
- Build source bundle normalization and card generation
- Expose project import and item removal APIs
- Expose project-scoped writing endpoints
- Build the minimal Web console
- Add file-based SKILL integration
Open Decisions Resolved
The following decisions were explicitly settled during brainstorming:
- service entrypoint is
API + Web - Claude Code / Codex integration is file-based, not API-driven
- card output format is
Markdown cards + JSON index - writing mode contains two independent actions
- knowledge-base mode is deferred, but its storage model is reserved now
- the data source is the local Zotero data directory
- current-selection import uses a small Zotero bridge for selection keys only
- collection hierarchy must mirror Zotero
Delivery Boundary
V1 is complete when the user can:
- create a project
- select multiple items in Zotero
- import the selected items into that project
- generate canonical Markdown cards from local Zotero data and attachments
- remove items from the project without affecting Zotero
- request citation candidates for a project-scoped writing intent
- request a first-pass citation plan for a project-scoped writing intent
- point Claude Code / Codex at a project and have it read only that project's cards