Access NIH Metabolomics Workbench via REST API (4,200+ studies). Query metabolites, RefMet nomenclature, MS/NMR data, m/z searches, study metadata, for metabolomics and biomarker discovery.
WHAT YOU BECOME
Perfect for these scenarios
Process billions of rows of CSV/HDF5 data without memory constraints.
Compute statistics and group-by operations on massive datasets instantly.
Create interactive plots and histograms for datasets larger than RAM.
Train machine learning models on datasets that don't fit in memory.
MEASURED GAIN
Proven benefits and measurable impact
Process large datasets up to 10x faster than in-memory solutions.
Handle datasets 100x larger than your available RAM with lazy evaluation.
Cut disk read time by 50% using memory-mapped files and streaming.
WHAT YOU GET
Files, tags and the three-step install
Tip: Read the documentation and the code before first use, so you know what it does and which permissions it needs.
NEXT SCRIPTS
Recommended based on tags and category
Access AlphaFold's 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Generate and maintain AGENTS.md files following the public agents.md convention. Use when creating documentation for AI agent workflows, onboarding guides, or when standardizing agent interaction patterns across projects.
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling large-scale biological datasets. Use when tasks involve AnnData objects, h5ad files, single-cell RNA-seq data, or integration with scanpy/scverse tools.