scanpy
Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for expl
By k-dense-ai · 1,421 installs
npx skills add k-dense-ai/scientific-agent-skills --skill scanpy
Source repository · Upstream listing
Scanpy: Single Cell Analysis
Overview
Scanpy is a scalable Python toolkit for analyzing single cell RNA seq data, built on AnnData. Apply this skill for complete single cell workflows including quality control, normalization, dimensionality reduction, clustering, marker gene identification, visualization, and trajectory analysis. Current stable release: scanpy 1.12.x (January 2026).
Installation
Requires Python 3.12+ (scanpy 1.12 dropped Python ≤3.11) and anndata ≥0.10 .
The [leiden] extra installs python igraph and leidenalg , required for Leiden clustering. For reproducible environments, pin a version: uv pip install "scanpy[leiden]==1.12.1" .
For large or out of core datasets, many functions support [Dask](https://docs.dask.org/) arrays (experimental):
See the [Using dask with Scanpy](https://scanpy.scverse.org/en/stable/tutorials/experimental/dask.html) tutorial. For GPU accelerated scanpy like operations, use [rapids singlecell](https://rapids singlecell.readthedocs.io/) as a separate package.
If the input is an R native single cell object ( .rds , .RData , Seurat, or SingleCellExperiment), first convert it to .h5ad with R tooling, then load it with Scanpy. Read references/r interop.md for agent run installation and conversion instructions across macOS, Linux, and Windows.
For AnnData structure and I/O details, use the anndata skill. For probabilistic models and batch correction, use scvi tools .
When to Use This Skill
This skill should be used when:
Analyzing single cell RNA seq data (.h5ad, 10X, CSV formats)
Working with R friendly single cell datasets ( .rds , .RData , Seurat, SingleCellExperiment) that need conversion to .h5ad
Performing quality control on scRNA seq datasets
Creating UMAP, t SNE, or PCA visualizations
Identifying cell clusters and finding marker genes
Annotating cell types based on gene expression
Conducting trajectory inference or pseudotime analysis
Generating publication quality single cell plots
Script Toolkit (prefer these over writing code from scratch)
This skill bundles ready to run CLI scripts in scripts/ for every common step. Run these instead of hand writing scanpy code — they handle file loading by extension, figure setup, sensible defaults, raw count preservation, and progress logging. Each reads and writes .h5ad , so they chain together, and each has its own help . Only drop down to writing scanpy code when a task isn't covered by a script or needs unusual customization.
All scripts use a shared scripts/ common.py helper (loading, saving, figure config) — keep it alongside the others. Run from the skill directory or pass full paths; figures default to ./figures/ .
Script Purpose Typical call
run pipeline.py Full workflow in one command : load → QC → normalize → HVG → PCA → (batch) → UMAP → Leiden → markers python scripts/run pipeline.py raw.h5ad o processed.h5ad
inspect data.py Summarize an unknown dataset (shape, obs/var, layers, what's already computed, raw vs normalized) python scripts/inspect data.py data.h5ad
convert.py Load any format (10x dir/.h5, csv, loom, mtx) and write .h5ad python scripts/convert.py 10x dir/ o data.h5ad
qc analysis.py QC metrics, before/after plots, filtering, optional Scrublet doublets python scripts/qc analysis.py raw.h5ad o qc.h5ad scrublet
preprocess.py Normalize, log1p, HVG, optional scale/regress (keeps counts layer + raw ) python scripts/preprocess.py qc.h5ad o norm.h5ad
reduce dimensions.py PCA + variance plot, neighbors, UMAP, optional t SNE python scripts/reduce dimensions.py norm.h5ad o red.h5ad
batch correct.py Integration: harmony / bbknn / combat python scripts/batch correct.py red.h5ad o int.h5ad method harmony batch key sample
cluster.py Leiden (or louvain) at one or many resolutions python scripts/cluster.py red.h5ad o clu.h5ad resolution 0.3 0.6 1.0
find markers.py rank genes groups + per group CSVs + marker plots python scripts/find markers.py clu.h5ad groupby leiden o clu.h5ad
annotate.py Map clusters → cell types from JSON/CSV; optional marker reference dotplot python scripts/annotate.py clu.h5ad o ann.h5ad mapping map.json
score genes.py Score gene signatures (JSON) and/or cell cycle phase python scripts/score genes.py ann.h5ad o scored.h5ad gene sets sigs.json
pseudobulk.py Aggregate counts by sample × cell type → matrix for pydeseq2 python scripts/pseudobulk.py ann.h5ad by sample cell type out prefix pb
subset.py Subset by obs values or gene list (optionally clear stale embeddings) python scripts/subset.py ann.h5ad o tcells.h5ad obs cell type keep "T cells"
plot.py Generate umap/tsne/pca/violin/dotplot/heatmap/etc. from a processed object python scripts/plot.py ann.h5ad kind dotplot genes CD3D CD14 groupby cell type
One shot end to end run
Step by step chain (when you need to inspect/iterate between stages)
The sections below document the underlying scanpy calls each script performs — read them when customizing beyond the script flags.
Quick Start
Basic Import and Setup
Loading Data
For R native files, do not try to parse Seurat .rds directly in Python. Convert first:
Understanding AnnData Structure
The AnnData object is the core data structure in scanpy:
Standard Analysis Workflow
The seven steps, with code and the parameters that matter at each, are in
[references/analysis workflow.md](references/analysis workflow.md):
1. Quality control — filter cells and genes; inspect mitochondrial fraction and counts
before choosing thresholds rather than copying defaults.
2. Normalization and preprocessing — normalize, log transform, select highly variable
genes, and keep .raw for later plotting.
3. Dimensionality reduction — PCA, then the neighbour graph, then UMAP.
4. Clustering — Leiden at a resolution chosen for the question, not the default.
5. Marker gene identification — ranked genes per cluster.
6. Cell type annotation — mapping clusters to types from markers.
7. Save results — writing the annotated AnnData .
Common follow on tasks — publication plots, trajectory inference, pseudobulk differential
expression between conditions, gene set scoring, and batch correction — are in the same
file. See also [references/standard workflow.md](references/standard workflow.md) and
[references/plotting guide.md](references/plotting guide.md).
Key Parameters to Adjust
Quality Control
min genes : Minimum genes per cell (typically 200 500)
min cells : Minimum cells per gene (typically 3 10)
pct counts mt : Mitochondrial threshold (typically 5 20%)
Normalization
target sum : Target counts per cell (default 1e4)
Feature Selection
n top genes : Number of HVGs (typically 2000 3000)
min mean , max mean , min disp : HVG selection parameters
Dimensionality Reduction
n pcs : Number of principal components (check variance ratio plot)
n neighbors : Number of neighbors (typically 10 30)
Clustering
resolution : Clustering granularity (0.4 1.2, higher = more clusters)
Common Pitfalls and Best Practices
1. Always save raw counts : adata.raw = adata before filtering genes
2. Check QC plots carefully : Adjust thresholds based on dataset quality
3. Use Leiden clustering : sc.tl.louvain is deprecated in scanpy 1.12
4. Try multiple clustering resolutions : Find optimal granularity
5. Validate cell type annotations : Use multiple marker genes
6. Use use raw=True for gene expression plots : Shows normalized counts from .raw
7. Check PCA variance ratio : Determine optimal number of PCs
8. Save intermediate results : Long workflows can fail partway through
9. Pseudobulk for DE : Do not treat rank genes groups p values as rigorous DE between conditions
10. Save plots via settings : Use sc.settings.autosave instead of deprecated save= on plot functions
11. Convert R objects before Scanpy : Use R packages to convert Seurat or SingleCellExperiment .rds files to .h5ad , preserving counts, metadata, and gene identifiers
Bundled Resources
scripts/ (CLI toolkit)
A composable set of .h5ad in/ .h5ad out scripts covering the whole workflow plus a one command end to end pipeline. See the Script Toolkit section above for the full table and chaining examples. Each script has help . Files:
common.py — shared loading/saving/figure helpers imported by the others (not a CLI)
run pipeline.py — full pipeline in one command (flags or config JSON)
inspect data.py , convert.py — explore and load/convert any input format
qc analysis.py , preprocess.py , reduce dimensions.py , batch correct.py , cluster.py — pipeline steps
find markers.py , annotate.py , score genes.py , pseudobulk.py — markers, annotation, scoring, DE prep
subset.py , plot.py — subset by metadata/genes; generate any standard plot
Default to these scripts before writing scanpy code from scratch.
references/standard workflow.md
Complete step by step workflow with detailed explanations and code examples for:
Data loading and setup
Quality control with visualization
Normalization and scaling
Feature selection
Dimensionality reduction (PCA, UMAP, t SNE)
Clustering (Leiden)
Doublet detection (scrublet) and pseudobulk aggregation
Marker gene identification
Cell type annotation
Trajectory inference
Differential expression
Read this reference when performing a complete analysis from scratch.
references/api reference.md
Quick reference guide for scanpy functions organized by module:
Reading/writing data ( sc.read , adata.write )
Preprocessing ( sc.pp. )
Tools ( sc.tl. )
Plotting ( sc.pl. )
AnnData structure and manipulation
Settings and utilities
Use this for quick lookup of function signatures and common parameters.
references/plotting guide.md
Comprehensive visualization guide including:
Quality control plots
Dimensionality reduction visualizations
Clustering visualizations
Marker gene plots (heatmaps, dot plots, violin plots)
Trajectory and pseudotime plots
Publication quality customization
Multi panel figures
Color palettes and styling
Consult this when creating publication ready figures.
references/r interop.md
Agent runbook for installing R on macOS, Linux, and Windows, installing CRAN/Bioconductor conversion packages, inspecting .rds / .RData inputs, converting Seurat or SingleCellExperiment objects to .h5ad , and validating the result in Scanpy.
assets/analysis template.py
Complete analysis template providing a full workflow from data loading through cell type annotation. Copy and customize this template for new analyses:
The template includes all standard steps with configurable parameters and helpful comments.
assets/ JSON templates
Edit and pass templates so you don't author config/mappings from scratch:
assets/pipeline config.json — parameter set for run pipeline.py config
assets/celltype mapping.json — cluster → cell type map for annotate.py mapping
assets/gene signatures.json — gene set signatures for score genes.py gene sets
Additional Resources
Official scanpy documentation : https://scanpy.scverse.org/en/stable/
Scanpy tutorials : https://scanpy.scverse.org/en/stable/tutorials/index.html
Release notes : https://scanpy.scverse.org/en/stable/release notes/index.html
scverse ecosystem : https://scverse.org/ (related tools: squidpy, scvi tools, cellrank)
R interoperability : https://www.bioconductor.org/packages/release/bioc/html/zellkonverter.html and https://mojaveazure.github.io/seurat disk/
Best practices : Luecken & Theis (2019) "Current best practices in single cell RNA seq"
Tips for Effective Analysis
1. Start with the template : Use assets/analysis template.