tooluniverse-disease-research
Generate comprehensive disease research reports covering genetics (causal genes, GWAS, OMIM), pathways (Reactome, KEGG), drugs (existing therapies, repurposing candidates), clinical trials, epidemiology (prevalence, incidence), and phenotypes (HPO). Use for full disease overviews, comprehensive dise
By mims-harvard · 422 installs
npx skills add mims-harvard/tooluniverse --skill tooluniverse-disease-research
Source repository · Upstream listing
ToolUniverse Disease Research
Generate a comprehensive disease research report with full source citations. The report is created as a markdown file and progressively updated during research.
IMPORTANT : Always use English disease names and search terms in tool calls. Respond in the user's language.
LOOK UP, DON'T GUESS
When asked about a disease, query Orphanet/OMIM/DisGeNET FIRST. Don't rely on memory for prevalence, genetics, or treatment — these change over time. When you're not sure about a fact, your first instinct should be to SEARCH for it using tools, not to reason harder from memory.
When to Use
User asks about any disease, syndrome, or medical condition
Needs comprehensive disease intelligence or a detailed research report
Asks "what do we know about [disease]?"
Core Workflow: Report First Approach
DO NOT show the search process to the user. Instead:
1. Create report file first Initialize {disease name} research report.md
2. Research each dimension Use all relevant tools
3. Update report progressively Write findings after each dimension
4. Include citations Every fact must reference its source tool
Disease Mechanism Reasoning
When synthesizing disease etiology, trace the full pathogenic cascade:
1. Genetic basis Which variants (rare or common) confer risk, and in which genes?
2. Molecular mechanism How do those variants alter protein function, expression, or regulation?
3. Cellular effect What downstream cellular processes are disrupted (signaling, metabolism, stress response)?
4. Tissue/organ manifestation How does cellular dysfunction present as organ level pathology?
This chain structures the Genetic & Molecular Basis (Section 3) and Biological Pathways (Section 5) sections.
10 Research Dimensions
Dim Section Key Tools
1 Identity & Classification OSL get efo id by disease name, ols search efo terms, ols get efo term, umls search concepts, icd search codes, snomed search concepts
2 Clinical Presentation OpenTargets phenotypes, HPO lookup, MedlinePlus
3 Genetic & Molecular Basis OpenTargets targets, ClinVar variants, GWAS associations, gnomAD
4 Treatment Landscape OpenTargets drugs, clinical trials, GtoPdb
5 Biological Pathways Reactome pathways, humanbase ppi analysis, GTEx expression, HPA
6 Epidemiology & Literature PubMed, OpenAlex, Europe PMC, Semantic Scholar
7 Similar Diseases OpenTargets similar entities
8 Cancer Specific (if applicable) CIViC genes/variants/therapies
9 Pharmacology GtoPdb targets/interactions/ligands
10 Drug Safety OpenTargets warnings, clinical trial AEs, FAERS
See: tool usage details.md for complete tool calls per section.
Normalizing free text to ontology IDs (Dimension 1)
When the input is messy free text (a sample attribute, a synonym, a tissue/organism label) rather than a clean disease name, use ZOOMA annotate text to map it to standardized ontology terms (EFO/MONDO/UBERON/etc.) before lookup. It returns each match as an ontology IRI with a confidence rating (HIGH/GOOD/MEDIUM/LOW), so you can keep only high confidence hits and feed the resolved ID into OLS / OpenTargets.
Each match also carries a ready to use curies field (e.g. MONDO:0004979 ) so you can feed the resolved ID straight into OLS / OpenTargets without parsing the IRI. ZOOMA is the live replacement for the retired OxO cross reference service; pair it with ols get efo term to expand the resolved IRI into labels, synonyms, and hierarchy.
Report Template
Create this file structure at the start:
Citation Format
Every piece of data MUST include its source:
In tables : Add a Source column with tool name
In lists : Finding [Source: tool name]
In prose : (Source: tool name, query: "...")
References section : Complete tool usage log with parameters
Progressive Update Pattern
Evidence Grading & Interpretation
Every finding in the report should be graded:
Grade Criteria Example
T1 (Strong) Replicated genetic evidence (GWAS, rare variants), FDA approved therapy BRCA1 → breast cancer; trastuzumab for HER2+
T2 (Moderate) Single genetic study, phase II+ trial data, strong biological evidence FOXO3 → longevity (centenarian studies)
T3 (Association) Observational data, gene expression changes, pathway membership IL 6 elevated in Alzheimer's CSF
T4 (Computational) Network proximity, text mining, predicted associations DisGeNET text mined gene disease link
Synthesis Questions (answer in Executive Summary)
After collecting data from all 10 dimensions, the report MUST answer:
1. What causes this disease? Summarize the genetic architecture (monogenic vs polygenic, key loci, penetrance)
2. What are the therapeutic options? Ranked by evidence level and approval status
3. What biomarkers exist? For diagnosis, prognosis, and treatment selection
4. What's the unmet need? What aspects lack effective treatment or understanding?
5. What are the active research frontiers? Based on clinical trials and recent publications
Interpreting Cross Database Concordance
When multiple databases provide different data for the same disease:
OpenTargets + DisGeNET + OMIM agree on a gene : T1 evidence — high confidence
Only OpenTargets reports an association : Check the datasource scores — genetic association literature animal model
DisGeNET score 0.5 but not in OpenTargets : May be text mined; verify with PubMed
Gene in GWAS but not OMIM : Likely a complex disease susceptibility locus, not Mendelian
Handling Conflicting Data
Conflict Resolution
Different prevalence estimates across sources Report range; note the most recent/largest study
Drug approved in one country but not another Note regulatory status per region
Gene disease association in one DB but absent in another Grade by evidence type; text mining alone is T4
Clinical trial results contradict label indications The trial result is newer evidence; note both
Final Report Quality Checklist
[ ] All 10 sections have content (or marked "No data available")
[ ] Every data point has a source citation
[ ] Executive summary reflects key findings
[ ] References section lists all tools used
[ ] Tables properly formatted
[ ] No placeholder text remains
Expected Output Scale
For a well studied disease (e.g., Alzheimer's), the final report should include:
5+ ontology IDs, 10+ synonyms, disease hierarchy
20+ phenotypes with HPO IDs
50+ genes, 30+ GWAS associations, 100+ ClinVar variants
20+ drugs, 50+ clinical trials
10+ pathways, PPI network, expression data
100+ publications
15+ similar diseases
Drug warnings and adverse events
Total: 500+ individual data points, each with source citation.
Cross Skill References
For rare disease differential diagnosis, run: python3 skills/tooluniverse rare disease diagnosis/scripts/clinical patterns.py type differential symptoms 'symptom1,symptom2'
Reference Files
[REPORT TEMPLATE.md](REPORT TEMPLATE.md) Full report markdown template and citation format guide
[RESEARCH PROTOCOL.md](RESEARCH PROTOCOL.md) Step by step code procedures, progressive update pattern, quality checklist
[tool usage details.md](tool usage details.md) Complete tool calls for each research dimension
[TOOLS REFERENCE.md](TOOLS REFERENCE.md) Complete tool documentation
[EXAMPLES.md](EXAMPLES.md) Sample disease research reports