tooluniverse-target-research

Comprehensive drug-target intelligence — tissue expression (GTEx, HPA), pathways, protein interactions (STRING), variant landscape (ClinVar, gnomAD), druggability (DGIdb, ChEMBL approved drugs). 9 parallel research paths with citations. Use for full target profile reports, target characterization fo

By mims-harvard · 396 installs

npx skills add mims-harvard/tooluniverse --skill tooluniverse-target-research

Source repository · Upstream listing

Comprehensive Target Intelligence Gatherer Gather complete target intelligence by exploring 9 parallel research paths. Supports targets identified by gene symbol, UniProt accession, Ensembl ID, or gene name. KEY PRINCIPLES : 1. Report first approach Create report file FIRST, then populate progressively 2. Tool parameter verification Verify params via get tool info before calling unfamiliar tools 3. Evidence grading Grade all claims by evidence strength (T1 T4) 4. Citation requirements Every fact must have inline source attribution 5. Mandatory completeness All sections must exist with data minimums or explicit "No data" notes 6. Disambiguation first Resolve all identifiers before research 7. Negative results documented "No drugs found" is data; empty sections are failures 8. Collision aware literature search Detect and filter naming collisions 9. English first queries Always use English terms in tool calls, even if the user writes in another language. Translate gene names, disease names, and search terms to English. Only try original language terms as a fallback if English returns no results. Respond in the user's language LOOK UP, DON'T GUESS When asked about a specific protein or gene target, look it up in UniProt/Ensembl/OpenTargets BEFORE reasoning about it. Verify the gene name, function, and disease associations from databases. When you're not sure about a fact, your first instinct should be to SEARCH for it using tools, not to reason harder from memory. When to Use This Skill Apply when users: Ask about a drug target, protein, or gene Need target validation or assessment Request druggability analysis Want comprehensive target profiling Ask "what do we know about [target]?" Need target disease associations Request safety profile for a target When NOT to use : Simple protein lookup, drug only queries, disease centric queries, sequence retrieval, structure download — use specialized skills instead. Target Evaluation Reasoning Framework Evaluating a drug target requires reasoning across four interconnected questions. Answer all four before forming a recommendation. 1. Is there genetic evidence linking this target to the disease? Genetic evidence is the strongest predictor of drug success — targets with human genetic support have approximately twice the clinical success rate as those without (Nelson et al. 2015). Ask: Are there GWAS associations connecting this gene to the disease? Do rare loss of function or gain of function variants cause or protect against the disease? Does the mouse knockout phenotype match the human disease (from OpenTargets mouse models)? OpenTargets assigns genetic evidence scores; a score 0.7 indicates strong support. ClinVar rare variant evidence and DisGeNET curated gene disease association scores add complementary layers. A target with no genetic link to the disease of interest carries a fundamental validation risk that cannot be resolved by downstream data. 2. Is the target druggable? Druggability has two components: structural accessibility and prior chemical matter. Structural accessibility means the target has a binding pocket where a small molecule or biologic can engage — surface exposed receptors, enzymes with well defined active sites, and protein protein interaction interfaces with hot spots are tractable. Intrinsically disordered proteins and transcription factors with flat, featureless binding surfaces are typically harder. Pharos TDL classification provides a tiered assessment: Tclin (approved drug), Tchem (known active compounds), Tbio (biological function known but no drugs), Tdark (poorly characterized). If ChEMBL or BindingDB have compounds with IC50 < 1μM, the target is chemically tractable. Chemical probes (from OpenTargets chemical probes endpoint) confirm a target can be modulated, which is distinct from drug like compounds. For GPCRs, check GPCRdb for curated agonists and antagonists. 3. Is the target safe to modulate? Safety concerns arise from two sources. First, on target effects: if the target is essential in normal tissues (mouse KO is lethal, or gnomAD pLI is high / LOEUF is low), full inhibition will produce toxicity — the question becomes whether a partial agonist or tissue targeted delivery can provide a therapeutic window. Second, off target effects: does the gene have family members that could be inadvertently hit? The OpenTargets safety profile aggregates known toxicity annotations, and DepMap essentiality scores tell you which cancer cell lines require this gene for survival (useful but not directly translatable to normal tissues). Expression specificity matters: a target expressed only in the disease relevant tissue is far safer than one expressed ubiquitously in critical organs (heart, kidney, brain). 4. What is the competitive landscape? A target with approved drugs may already be validated but competitive; a target with clinical stage programs from competitors establishes feasibility while creating IP barriers. An entirely novel target with no drug history requires more extensive internal validation. Assess: number of ChEMBL bioactivity records (chemical matter depth), approved drugs from OpenTargets drug associations, and literature activity trends (recent paper count and key research groups). A dark target (Tdark) with strong genetic evidence but no chemical matter is a high risk, high reward opportunity. Synthesizing the four dimensions : The ideal target has strong genetic evidence (GWAS + rare variant), a tractable binding site (Tclin or Tchem), acceptable safety profile (tissue specific expression, non lethal KO), and manageable competition. Gaps in any dimension represent validation tasks, not disqualifiers — but they must be acknowledged. A target with perfect druggability but no genetic link to disease is a tractability exercise, not a validated therapeutic hypothesis. Phase 0: Tool Parameter Verification (CRITICAL) BEFORE calling ANY tool for the first time , verify its parameters: Known parameter corrections: Reactome map uniprot to pathways : param is id (not uniprot id ) ensembl get xrefs : param is id (not gene id ) GTEx get median gene expression : requires gencode id + operation="median" ; try versioned Ensembl ID if empty OpenTargets : param is ensemblId (camelCase, not ensemblID ) STRING get protein interactions : takes protein ids (list) + species intact get interactions : takes identifier (UniProt accession, not gene symbol) Critical Workflow Requirements Report First (MANDATORY) : Create [TARGET] target report.md with all section headers and [Researching...] placeholders before starting research. Update progressively. Do not show raw tool outputs to the user. Evidence Grading (MANDATORY) : Grade every claim T1 T4. T1 = clinical/genetic data; T2 = curated databases or multiple studies; T3 = computational or single study; T4 = annotation or catalog entry. Core Strategy: 9 Research Paths For detailed code implementations of each path, see [IMPLEMENTATION.md](IMPLEMENTATION.md). Identifier Resolution (Phase 1) Resolve ALL identifiers before any research path. Required IDs: UniProt accession (for protein data, structure, interactions) Ensembl gene ID + versioned ID (for Open Targets, GTEx) Gene symbol (for DGIdb, gnomAD, literature) Entrez gene ID (for KEGG, MyGene) ChEMBL target ID (for bioactivity) Synonyms/full name (for collision aware literature search) After resolution, check if target is a GPCR via GPCRdb get protein . See [IMPLEMENTATION.md](IMPLEMENTATION.md) for resolution and GPCR detection code. PATH 0: Open Targets Foundation (ALWAYS FIRST) Run OpenTargets endpoints first to populate baseline data before specialized queries: OpenTargets get diseases phenotypes by target ensembl → disease associations (Section 8) OpenTargets get target tractability by ensemblID → druggability assessment (Section 9) OpenTargets get target safety profile by ensemblID → safety liabilities (Section 10) OpenTargets get target interactions by ensemblID → PPI network (Section 6) OpenTargets get target gene ontology by ensemblID → GO annotations (Section 5) OpenTargets get publications by target ensemblID → literature (Section 11) OpenTargets get biological mouse models by ensemblID → mouse KO phenotypes (Sections 8/10) OpenTargets get chemical probes by target ensemblID → chemical probes (Section 9) OpenTargets get associated drugs by target ensemblID → known drugs (Section 9) PATH 1: Core Identity Tools : UniProt get entry by accession , UniProt get function by accession , UniProt get recommended name by accession , UniProt get alternative names by accession , UniProt get subcellular location by accession , MyGene get gene annotation Populates : Sections 2 (Identifiers), 3 (Basic Information) PATH 2: Structure & Domains Use 3 step structure search chain (do NOT rely solely on PDB text search): 1. UniProt PDB cross references (most reliable) 2. Sequence based PDB search (catches missing annotations) 3. Domain based search (for multi domain proteins) 4. AlphaFold (always check) Tools : UniProt get entry by accession (PDB xrefs), RCSBData get entry , PDB search similar structures , alphafold get prediction , InterPro get protein domains , UniProt get ptm processing by accession GPCR targets : Also query GPCRdb get structures for active/inactive state data. Populates : Section 4 (Structural Biology) PATH 3: Function & Pathways Tools : GO get annotations for gene , Reactome map uniprot to pathways , kegg get gene info , WikiPathways search , enrichr gene enrichment analysis Populates : Section 5 (Function & Pathways) PATH 4: Protein Interactions Tools : STRING get protein interactions , intact get interactions , intact get complex details , BioGRID get interactions , HPA get protein interactions by gene Minimum : 20 interactors OR documented explanation. Populates : Section 6 (Protein Protein Interactions) PATH 5: Expression Profile GTEx with versioned ID fallback + HPA as backup. Tools : GTEx get median gene expression , HPA get rna expression by source , HPA get comprehensive gene details by ensembl id , HPA get subcellular location , HPA get cancer prognostics by gene , HPA get comparative expression by gene and cellline , CELLxGENE get expression data Reasoning : Expression specificity directly informs safety. Note whether expression is enriched in the disease relevant tissue vs. critical organs. Ubiquitous essential expression narrows the therapeutic window. Populates : Section 7 (Expression Profile) PATH 6: Variants & Disease Separate SNVs from CNVs in ClinVar results. Integrate DisGeNET for curated gene disease association scores. Tools : gnomad get gene constraints , ClinVar search variants , OpenTargets get diseases phenotypes by target ensembl , DisGeNET search gene , civic get variants by gene , cBioPortal get mutations Required constraint scores : pLI (probability of loss of function intolerance), LOEUF (loss of function observed/expected upper bound), missense Z score, pRec (recessive probability). High pLI ( 0.9) or low LOEUF (< 0.35) indicates the gene is intolerant to loss of function — a major safety flag for inhibitory therapeutic strategies. Populates : Section 8 (Genetic Variation & Disease) PATH 7: Druggability & Target Validation Tools : OpenTargets get target tractability by ensemblID , DGIdb get gene druggability , DGIdb get drug gene interactions , ChEMBL search targets , ChEMBL get target activities , Pharos get target , BindingDB get ligands by uniprot , PubChem search assays by target gene , DepMap get gene dependencies , OpenTargets get target safety profile by ensemblID , OpenTargets get biolo