tooluniverse-literature-deep-research
Deep literature review — PubMed, EuropePMC, bioRxiv preprints, citation networks, evidence synthesis. Disambiguates queries, runs collision-aware searches, grades evidence T1-T4, and produces structured reports. Use for systematic literature review, meta-analysis evidence collection, and detailed an
By mims-harvard · 646 installs
npx skills add mims-harvard/tooluniverse --skill tooluniverse-literature-deep-research
Source repository · Upstream listing
Literature Deep Research
Systematic literature research: disambiguate, search with collision aware queries, grade evidence, produce structured reports.
KEY PRINCIPLES : (1) Disambiguate first (2) Right size deliverable (3) Grade every claim T1 T4 (4) All sections mandatory even if "limited evidence" (5) Source attribution for every claim (6) English first queries, respond in user's language (7) Report = deliverable, not search log
LOOK UP, DON'T GUESS
Search PubMed/EuropePMC FIRST before reasoning. A published paper beats memory.
Factoid search strategy:
1. Extract KEY TERMS (most specific nouns/verbs)
2. EuropePMC search articles(query="term1 term2 term3", limit=5)
3. No results BROADEN (remove most restrictive term)
4. Too many NARROW (add specific terms)
5. Answer usually in abstract of top results
6. Failed query try DIFFERENT TERMS/synonyms, don't repeat
COMPUTE, DON'T DESCRIBE
When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
Workflow
Phase 0: Mode Selection
Mode When Deliverable
Factoid Single concrete question 1 page fact check report + bibliography
Mini review Narrow topic 1 3 page narrative
Full Deep Research Comprehensive overview 15 section report + bibliography
Factoid Mode (Fast Path)
Domain Detection
Pattern Domain Action
Gene/protein symbol Biological target Full bio disambiguation
Drug name Drug Drug disambiguation (1.5)
Disease name Disease Disease disambiguation (1.6)
CS/ML topic General academic Skip bio tools, literature only
Cross domain Interdisciplinary Resolve each entity in its domain
Cross Skill Delegation
Gene/protein deep dive: tooluniverse target research
Drug profile: tooluniverse drug research
Disease profile: tooluniverse disease research
Use this skill for literature synthesis . Use specialized skills for entity profiling . For max depth, run both.
Phase 1: Subject Disambiguation + Profile
1.1 Biological Target Resolution
1.2 Naming Collision Detection
Check first 20 results. If 20% off topic, build negative filter: NOT [collision1] NOT [collision2] .
Gene family: "ADAR" NOT "ADAR2" NOT "ADARB1" . Cross domain: add context terms.
1.3 Baseline Profile (Bio Targets)
GPCR targets: delegate to tooluniverse target research .
1.5 Drug Disambiguation
Identity : OpenTargets get drug chembId by generic name , ChEMBL get drug , PubChem get CID by compound name , drugbank get drug basic info by drug name or id
Targets : ChEMBL get drug mechanisms , OpenTargets get associated targets by drug chemblId , DGIdb get drug gene interactions
Safety : OpenTargets get drug adverse events by chemblId , OpenTargets get drug indications by chemblId , search clinical trials
1.6 Disease Disambiguation
1.7 Compound Queries (e.g., "metformin in breast cancer")
Resolve both entities, then cross reference via CTD get chemical gene interactions, CTD get chemical diseases, OpenTargets drug target/drug disease tools. Intersect shared targets/pathways.
1.8 General Academic / 1.9 Interdisciplinary
Non bio: skip bio tools, use ArXiv/DBLP/OSF. Cross domain: resolve bio entities with 1.1 1.3, search CS/general in parallel, merge and cross reference.
Phase 2: Literature Search
Methodology stays internal. Report shows findings, not process.
2.1 Query Strategy
Step 1: Seeds (15 30 core papers): domain specific title searches with date/sort filters.
Step 2: Citation expansion : PubMed get cited by , EuropePMC get citations/references , PubMed get related , SemanticScholar get recommendations , OpenCitations get citations
Step 3: Collision filtered broader queries : "[TERM]" AND ([context]) NOT [collision]
2.2 Literature Tools — core set + adaptive by domain
Run the core multi field set on every review (catches what any single index misses), then add the domain rows that match the subject. Don't fire every source blindly — 6–10 well chosen indexes beat 20 noisy ones.
ALWAYS run (core, all disciplines) : PubMed search articles , EuropePMC search articles , openalex search works (query param search / query ) or openalex literature search (query param search keywords ) — pick one and match its param; mixing them silently returns off topic results — and SemanticScholar search papers
Then add by domain:
Domain Add these Notes
Biomedical / clinical PMC search papers (full text), PubTator3 LiteratureSearch (entity & relations: queries), PubMed Guidelines Search (clinical guidelines) PubTator normalizes gene/drug/disease entities
Biology (ecology/evolution/plant) EuropePMC as PRIMARY + OpenAlex PubMed returns 0–1 for non clinical biology
CS / ML / AI ArXiv search papers , DBLP search publications arXiv + CS bibliography
Physics / HEP / astro InspireHEP search papers 1.6M+ particle/astro records
Broad / hard to find / OA Crossref search works , CORE search papers , DOAJ search articles , Fatcat search scholar DOI registry + OA aggregators + Internet Archive Scholar
Regional / EU funded OpenAIRE search publications , HAL search archive EU open science + French national archive
Datasets / software / outputs Figshare search articles , Zenodo search records Citable DOIs for data & code
Preprints (latest) EuropePMC search articles(source='PPR') , OSF search preprints , BioRxiv get preprint / MedRxiv get preprint (DOI lookup) bioRxiv/medRxiv/PsyArXiv etc.
Multi source : advanced literature search agent (12+ DBs; needs Azure key fallback: query the core set individually).
Citation impact : iCite search publications (RCR/APT), iCite get publications (by PMID), scite get tallies (support/contradict). PubMed only; for CS use SemanticScholar.
A domain specific index returning 0 (e.g. ArXiv on a pure clinical topic) is normal — only worry if the whole core set is empty.
2.3 2.4 Full Text & PubMed Zero Result Fallback
Full text: see FULLTEXT STRATEGY.md for three tier strategy.
CRITICAL : PubMed returns 0 for ~30% of valid queries. Always retry with EuropePMC when PubMed returns empty. This is not optional.
2.5 Tool Failure / OA Handling
Retry once fallback tool. Key fallbacks: PubMed get cited by EuropePMC get citations OpenCitations. OA: Unpaywall if configured, else Europe PMC/PMC/OpenAlex flags.
Phase 3: Evidence Grading
Tier Label Bio Example CS/ML Example
T1 Mechanistic CRISPR KO + rescue, RCT Formal proof, controlled ablation
T2 Functional siRNA knockdown phenotype Benchmark with baselines
T3 Association GWAS, screen hit Observational, case study
T4 Mention Review article Survey, workshop abstract
Inline: Target X regulates Y [T1: PMID:12345678] . Per theme: summarize evidence distribution.
Report Output
File Mode
[topic] report.md Full
[topic] factcheck report.md Factoid
[topic] bibliography.json + .csv All
Progressive update : create report with all section headers immediately. Fill after each phase. Write Executive Summary LAST.
Use 15 section template from REPORT TEMPLATE.md . Domain adaptations: bio (architecture/expression/GO/disease), drug (properties/MOA/PK/safety), disease (epi/patho/genes/treatments), general (history/theories/evidence/applications).
Communication
Brief progress updates only: "Resolving identifiers...", "Building paper set...", "Grading evidence..."
Do NOT expose: raw tool outputs, dedup counts, search round details.
References
TOOL NAMES REFERENCE.md 123 tools with parameters
REPORT TEMPLATE.md template, domain adaptations, bibliography, completeness checklist
FULLTEXT STRATEGY.md three tier full text verification
WORKFLOW.md compact cheat sheet
EXAMPLES.md worked examples