tooluniverse-literature-deep-research

Deep literature review — PubMed, EuropePMC, bioRxiv preprints, citation networks, evidence synthesis. Disambiguates queries, runs collision-aware searches, grades evidence T1-T4, and produces structured reports. Use for systematic literature review, meta-analysis evidence collection, and detailed an

By mims-harvard · 646 installs

npx skills add mims-harvard/tooluniverse --skill tooluniverse-literature-deep-research

Source repository · Upstream listing

Literature Deep Research Systematic literature research: disambiguate, search with collision aware queries, grade evidence, produce structured reports. KEY PRINCIPLES : (1) Disambiguate first (2) Right size deliverable (3) Grade every claim T1 T4 (4) All sections mandatory even if "limited evidence" (5) Source attribution for every claim (6) English first queries, respond in user's language (7) Report = deliverable, not search log LOOK UP, DON'T GUESS Search PubMed/EuropePMC FIRST before reasoning. A published paper beats memory. Factoid search strategy: 1. Extract KEY TERMS (most specific nouns/verbs) 2. EuropePMC search articles(query="term1 term2 term3", limit=5) 3. No results BROADEN (remove most restrictive term) 4. Too many NARROW (add specific terms) 5. Answer usually in abstract of top results 6. Failed query try DIFFERENT TERMS/synonyms, don't repeat COMPUTE, DON'T DESCRIBE When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it. Workflow Phase 0: Mode Selection Mode When Deliverable Factoid Single concrete question 1 page fact check report + bibliography Mini review Narrow topic 1 3 page narrative Full Deep Research Comprehensive overview 15 section report + bibliography Factoid Mode (Fast Path) Domain Detection Pattern Domain Action Gene/protein symbol Biological target Full bio disambiguation Drug name Drug Drug disambiguation (1.5) Disease name Disease Disease disambiguation (1.6) CS/ML topic General academic Skip bio tools, literature only Cross domain Interdisciplinary Resolve each entity in its domain Cross Skill Delegation Gene/protein deep dive: tooluniverse target research Drug profile: tooluniverse drug research Disease profile: tooluniverse disease research Use this skill for literature synthesis . Use specialized skills for entity profiling . For max depth, run both. Phase 1: Subject Disambiguation + Profile 1.1 Biological Target Resolution 1.2 Naming Collision Detection Check first 20 results. If 20% off topic, build negative filter: NOT [collision1] NOT [collision2] . Gene family: "ADAR" NOT "ADAR2" NOT "ADARB1" . Cross domain: add context terms. 1.3 Baseline Profile (Bio Targets) GPCR targets: delegate to tooluniverse target research . 1.5 Drug Disambiguation Identity : OpenTargets get drug chembId by generic name , ChEMBL get drug , PubChem get CID by compound name , drugbank get drug basic info by drug name or id Targets : ChEMBL get drug mechanisms , OpenTargets get associated targets by drug chemblId , DGIdb get drug gene interactions Safety : OpenTargets get drug adverse events by chemblId , OpenTargets get drug indications by chemblId , search clinical trials 1.6 Disease Disambiguation 1.7 Compound Queries (e.g., "metformin in breast cancer") Resolve both entities, then cross reference via CTD get chemical gene interactions, CTD get chemical diseases, OpenTargets drug target/drug disease tools. Intersect shared targets/pathways. 1.8 General Academic / 1.9 Interdisciplinary Non bio: skip bio tools, use ArXiv/DBLP/OSF. Cross domain: resolve bio entities with 1.1 1.3, search CS/general in parallel, merge and cross reference. Phase 2: Literature Search Methodology stays internal. Report shows findings, not process. 2.1 Query Strategy Step 1: Seeds (15 30 core papers): domain specific title searches with date/sort filters. Step 2: Citation expansion : PubMed get cited by , EuropePMC get citations/references , PubMed get related , SemanticScholar get recommendations , OpenCitations get citations Step 3: Collision filtered broader queries : "[TERM]" AND ([context]) NOT [collision] 2.2 Literature Tools — core set + adaptive by domain Run the core multi field set on every review (catches what any single index misses), then add the domain rows that match the subject. Don't fire every source blindly — 6–10 well chosen indexes beat 20 noisy ones. ALWAYS run (core, all disciplines) : PubMed search articles , EuropePMC search articles , openalex search works (query param search / query ) or openalex literature search (query param search keywords ) — pick one and match its param; mixing them silently returns off topic results — and SemanticScholar search papers Then add by domain: Domain Add these Notes Biomedical / clinical PMC search papers (full text), PubTator3 LiteratureSearch (entity & relations: queries), PubMed Guidelines Search (clinical guidelines) PubTator normalizes gene/drug/disease entities Biology (ecology/evolution/plant) EuropePMC as PRIMARY + OpenAlex PubMed returns 0–1 for non clinical biology CS / ML / AI ArXiv search papers , DBLP search publications arXiv + CS bibliography Physics / HEP / astro InspireHEP search papers 1.6M+ particle/astro records Broad / hard to find / OA Crossref search works , CORE search papers , DOAJ search articles , Fatcat search scholar DOI registry + OA aggregators + Internet Archive Scholar Regional / EU funded OpenAIRE search publications , HAL search archive EU open science + French national archive Datasets / software / outputs Figshare search articles , Zenodo search records Citable DOIs for data & code Preprints (latest) EuropePMC search articles(source='PPR') , OSF search preprints , BioRxiv get preprint / MedRxiv get preprint (DOI lookup) bioRxiv/medRxiv/PsyArXiv etc. Multi source : advanced literature search agent (12+ DBs; needs Azure key fallback: query the core set individually). Citation impact : iCite search publications (RCR/APT), iCite get publications (by PMID), scite get tallies (support/contradict). PubMed only; for CS use SemanticScholar. A domain specific index returning 0 (e.g. ArXiv on a pure clinical topic) is normal — only worry if the whole core set is empty. 2.3 2.4 Full Text & PubMed Zero Result Fallback Full text: see FULLTEXT STRATEGY.md for three tier strategy. CRITICAL : PubMed returns 0 for ~30% of valid queries. Always retry with EuropePMC when PubMed returns empty. This is not optional. 2.5 Tool Failure / OA Handling Retry once fallback tool. Key fallbacks: PubMed get cited by EuropePMC get citations OpenCitations. OA: Unpaywall if configured, else Europe PMC/PMC/OpenAlex flags. Phase 3: Evidence Grading Tier Label Bio Example CS/ML Example T1 Mechanistic CRISPR KO + rescue, RCT Formal proof, controlled ablation T2 Functional siRNA knockdown phenotype Benchmark with baselines T3 Association GWAS, screen hit Observational, case study T4 Mention Review article Survey, workshop abstract Inline: Target X regulates Y [T1: PMID:12345678] . Per theme: summarize evidence distribution. Report Output File Mode [topic] report.md Full [topic] factcheck report.md Factoid [topic] bibliography.json + .csv All Progressive update : create report with all section headers immediately. Fill after each phase. Write Executive Summary LAST. Use 15 section template from REPORT TEMPLATE.md . Domain adaptations: bio (architecture/expression/GO/disease), drug (properties/MOA/PK/safety), disease (epi/patho/genes/treatments), general (history/theories/evidence/applications). Communication Brief progress updates only: "Resolving identifiers...", "Building paper set...", "Grading evidence..." Do NOT expose: raw tool outputs, dedup counts, search round details. References TOOL NAMES REFERENCE.md 123 tools with parameters REPORT TEMPLATE.md template, domain adaptations, bibliography, completeness checklist FULLTEXT STRATEGY.md three tier full text verification WORKFLOW.md compact cheat sheet EXAMPLES.md worked examples