research-lit

Search and analyze research papers, find related work, summarize key ideas. Use when user says "find papers", "related work", "literature review", "what does this paper say", or needs to understand academic papers.

By wanshuiyin · 502 installs

npx skills add wanshuiyin/auto-claude-code-research-in-sleep --skill research-lit

Source repository · Upstream listing

Research Literature Review Research topic: $ARGUMENTS Constants PAPER LIBRARY — Local directory containing user's paper collection (PDFs). Check these paths in order: 1. papers/ in the current project directory 2. literature/ in the current project directory 3. Custom path specified by user in CLAUDE.md under Paper Library MAX LOCAL PAPERS = 20 — Maximum number of local PDFs to scan (read first 3 pages each). If more are found, prioritize by filename relevance to the topic. SOURCES = all — Which literature sources to search. Options: zotero , obsidian , local , web , semantic scholar , deepxiv , exa , gemini , openalex , all . Full source table and selection rules: see Data Sources below. ARXIV DOWNLOAD = false — When true , download top 3 5 most relevant arXiv PDFs to PAPER LIBRARY after search. When false (default), only fetch metadata (title, abstract, authors) via arXiv API — no files are downloaded. ARXIV MAX DOWNLOAD = 5 — Maximum number of PDFs to download when ARXIV DOWNLOAD = true . 💡 Overrides: /research lit "topic" — paper library: ~/my papers/ — custom local PDF path /research lit "topic" — sources: zotero, local — only search Zotero + local PDFs /research lit "topic" — sources: web — only search the web (skip all local) /research lit "topic" — sources: web, semantic scholar — also search Semantic Scholar for published venue papers (IEEE, ACM, etc.) /research lit "topic" — sources: all, deepxiv — use default sources plus DeepXiv /research lit "topic" — arxiv download: true — download top relevant arXiv PDFs /research lit "topic" — arxiv download: true, max download: 10 — download up to 10 PDFs Data Sources This skill checks multiple sources in priority order . All are optional — if a source is not configured or not requested, skip it silently. Source Selection Parse $ARGUMENTS for a — sources: directive: If — sources: is specified : Only search the listed sources (comma separated). Valid values: zotero , obsidian , local , web , semantic scholar , deepxiv , exa , gemini , openalex , all . If not specified : Default to all — search every available source in priority order ( semantic scholar , deepxiv , exa , gemini , and openalex are excluded from all ; they must be explicitly listed). Examples: Source Table Priority Source ID How to detect What it provides 1 Zotero (via MCP) zotero Try calling any mcp zotero tool — if unavailable, skip Collections, tags, annotations, PDF highlights, BibTeX, semantic search 2 Obsidian (via MCP) obsidian Try calling any mcp obsidian vault tool — if unavailable, skip Research notes, paper summaries, tagged references, wikilinks 3 Local PDFs local Glob: papers/ / .pdf, literature/ / .pdf Raw PDF content (first 3 pages) 4 Web search web Always available (WebSearch) arXiv, Semantic Scholar, Google Scholar 5 Semantic Scholar API semantic scholar $S2 FETCHER resolves (canonical name semantic scholar fetch.py , per integration contract §2) Published venue papers (IEEE, ACM, Springer) with structured metadata: citation counts, venue info, TLDR. Only runs when explicitly requested via — sources: semantic scholar or — sources: web, semantic scholar 6 DeepXiv CLI deepxiv $DEEPXIV FETCHER resolves (canonical name deepxiv fetch.py , per integration contract §2) and deepxiv CLI present ( command v deepxiv ) Progressive paper retrieval: search, brief, head, section, trending, web search. Only runs when explicitly requested via — sources: deepxiv or — sources: all, deepxiv 7 Exa Search exa $EXA FETCHER resolves (canonical name exa search.py , per integration contract §2); fetcher handles exa py SDK + API key internally AI powered broad web search with content extraction (highlights, text, summaries). Covers blogs, docs, news, companies, and research papers beyond arXiv/S2. Only runs when explicitly requested via — sources: exa or — sources: all, exa 8 Gemini (MCP / CLI) gemini mcp gemini cli ask gemini tool available, or gemini CLI installed AI powered broad literature discovery — decomposes topics into sub problems, aliases, and variants for wider retrieval. Prefers MCP, falls back to CLI. Only runs when explicitly requested via — sources: gemini or — sources: all, gemini 9 OpenAlex openalex $OPENALEX FETCHER resolves (canonical name openalex fetch.py , per integration contract §2) and Python requests module importable Open citation graph with institutional affiliations, funding data, and comprehensive metadata across 250M+ works. Fully open API. Only runs when explicitly requested via — sources: openalex or — sources: all, openalex Graceful degradation : If no MCP servers are configured, the skill works exactly as before (local PDFs + web search). Zotero and Obsidian are pure additions. Workflow Step 0a: Search Zotero Library (if available) Skip this step entirely if Zotero MCP is not configured. Try calling a Zotero MCP tool (e.g., search). If it succeeds: 1. Search by topic : Use the Zotero search tool to find papers matching the research topic 2. Read collections : Check if the user has a relevant collection/folder for this topic 3. Extract annotations : For highly relevant papers, pull PDF highlights and notes — these represent what the user found important 4. Export BibTeX : Get citation data for relevant papers (useful for /paper write later) 5. Compile results : For each relevant Zotero entry, extract: Title, authors, year, venue User's annotations/highlights (if any) Tags the user assigned Which collection it belongs to 📚 Zotero annotations are gold — they show what the user personally highlighted as important, which is far more valuable than generic summaries. Step 0b: Search Obsidian Vault (if available) Skip this step entirely if Obsidian MCP is not configured. Try calling an Obsidian MCP tool (e.g., search). If it succeeds: 1. Search vault : Search for notes related to the research topic 2. Check tags : Look for notes tagged with relevant topics (e.g., diffusion models , paper review ) 3. Read research notes : For relevant notes, extract the user's own summaries and insights 4. Follow links : If notes link to other relevant notes (wikilinks), follow them for additional context 5. Compile results : For each relevant note: Note title and path User's summary/insights Links to other notes (research graph) Any frontmatter metadata (paper URL, status, rating) 📝 Obsidian notes represent the user's processed understanding — more valuable than raw paper content for understanding their perspective. Step 0c: Scan Local Paper Library Before searching online, check if the user already has relevant papers locally: 1. Locate library : Check PAPER LIBRARY paths for PDF files 2. De duplicate against Zotero : If Step 0a found papers, skip any local PDFs already covered by Zotero results (match by filename or title). 3. Filter by relevance : Match filenames and first page content against the research topic. Skip clearly unrelated papers. 4. Summarize relevant papers : For each relevant local PDF (up to MAX LOCAL PAPERS): Read first 3 pages (title, abstract, intro) Extract: title, authors, year, core contribution, relevance to topic Flag papers that are directly related vs tangentially related 5. Build local knowledge base : Compile summaries into a "papers you already have" section. This becomes the starting point — external search fills the gaps. 📚 If the user has a comprehensive local collection, the external search can be more targeted (focus on what's missing). ⚠️ If all three PAPER LIBRARY paths miss, say so before moving on — do not skip silently. A user whose PDFs live in a reference manager (Zotero, Mendeley, ...) otherwise assumes — sources: all covered them. Emit: WARN: local contributed nothing — no PDFs found in papers/, literature/, or a configured paper library. To include yours, add a " Paper Library" heading to CLAUDE.md followed by the directory path. Then continue to Step 1. Step 1: Search (external) Use WebSearch to find recent papers on the topic Check arXiv, Semantic Scholar, Google Scholar Focus on papers from last 2 years unless studying foundational work De duplicate : Skip papers already found in Zotero, Obsidian, or local library arXiv API search (runs when — sources: is unset, contains web or all ; no download by default — arXiv API is part of the Priority 4 Web tier, see Source Table above): Policy D2 tracking discipline (orchestrator managed) : the executor (you, the LLM) maintains an in context list of contributing sources. For helper backed bash sources (arxiv, semantic scholar, deepxiv, exa, openalex), a source contributes iff its bash block ran its helper successfully (helper resolved AND invocation exited 0; note: the helper exiting 0 with an empty result list still counts as "ran" — downstream relevance ranking is what decides whether the user actually sees content). For non helper sources (zotero / obsidian / local PDF / WebSearch / Gemini), the contribution rule is stated in the Step 1 finalization block below — these are tracked separately because they don't emit D2 contribution: log lines from bash. Sources that were not requested via — sources: do not count. At the end of Step 1 (before "Optional PDF download"), if zero sources contributed, surface a D2 empty aggregate error and stop. (See integration contract.md §2 Policy D2 — the in context tracking replaces a shared bash accumulator because SKILL bash blocks are executed as separate shells; state does not survive.) Resolve $ARXIV FETCHER via the canonical chain (Policy D2 — this source contributes to the multi source aggregate; warn and continue on failure, never abort the whole aggregate): Record keeping : track the D2 contribution: … lines emitted by each source's bash block. They form the contributing source list the orchestrator uses for the Step 1 finalization gate below. WebSearch (Priority 4) is treated as having contributed iff WebSearch was requested (no — sources: filter, or the list contains web or all ) AND was actually invoked; the orchestrator records that separately. (The finalization block below restates this rule canonically — both lines must stay in sync.) If $ARXIV FETCHER is empty (D2 graceful degradation), fall back to WebSearch for arXiv (same as before). The arXiv API returns structured metadata (title, abstract, full author list, categories, dates) — richer than WebSearch snippets. Merge these results with WebSearch findings and de duplicate. Semantic Scholar API search (only when semantic scholar is in sources): When the user explicitly requests — sources: semantic scholar (or — sources: web, semantic scholar ), search for published venue papers beyond arXiv: If $S2 FETCHER is empty (canonical chain exhausted), skip silently — D2 multi source aggregate continues with the remaining resolved sources. Why use Semantic Scholar? Many IEEE/ACM journal papers are NOT on arXiv. S2 fills the gap for published venue only papers with citation counts and venue metadata. De duplication between arXiv and S2 : Match by arXiv ID (S2 returns externalIds.ArXiv ): If a paper appears in both: check S2's venue / publicationVenue — if it has been published in a journal/conference (e.g. IEEE TWC, JSAC), use S2's metadata (venue, citationCount, DOI) as the authoritative version, since the published version supersedes the preprint. Keep the arXiv PDF link for download. If the S2 match has no venue (still just a preprint indexed by S2): keep the arXiv ver