Molecular featurization for ML (100+ featurizers). ECFP, MACCS, descriptors, pretrained models (ChemBERTa), convert SMILES to features, for QSAR and molecular ML.
WHAT YOU BECOME
Perfect for these scenarios
Remove technical variability across datasets while preserving biological signals
Automatically identify and label cell types using probabilistic deep learning
Combine scRNA-seq with ATAC or CITE-seq data for comprehensive analysis
Analyze gene expression patterns while preserving spatial context
MEASURED GAIN
Proven benefits and measurable impact
Reduce preprocessing and analysis time with automated probabilistic models
Achieve more accurate cell type identification and batch correction
Handle millions of cells efficiently with GPU-accelerated deep learning
WHAT YOU GET
Files, tags and the three-step install
Tip: Read the documentation and the code before first use, so you know what it does and which permissions it needs.
NEXT SCRIPTS
Recommended based on tags and category
Access AlphaFold's 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling large-scale biological datasets. Use when tasks involve AnnData objects, h5ad files, single-cell RNA-seq data, or integration with scanpy/scverse tools.
Comprehensive Python library for astronomy and astrophysics. This skill should be used when working with astronomical data including celestial coordinates, physical units, FITS files, cosmological calculations, time systems, tables, world coordinate systems (WCS), and astronomical data analysis. Use when tasks involve coordinate transformations, unit conversions, FITS file manipulation, cosmological distance calculations, time scale conversions, or astronomical data processing.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.