Skip to the page
Research notebook / A. Flores
folio 03·b
tmap 2.0

TMAP 2.0.

Interactive maps for datasets too large to understand as tables.

TMAP draws a 2D map of anything you can compare by similarity: molecules, proteins, images, cells. Similar objects end up next to each other, and because the map is a tree, you can still follow its overall structure with millions of points on it.

Preprint · ChemRxiv 2026

Plate I

Approved drugs · TMAP

An interactive map; drag to pan, scroll to zoom

Loading the interactive map…

open the map on its own page ↗

Plate I. MST over multi-domain embeddings: chemical, image, and protein space.

i.Why a map

A molecule, a protein or an image is usually stored as hundreds or thousands of numbers, and nobody can look at that directly. TMAP connects each point to its nearest neighbors, keeps a tree through them and lays the tree out flat. Clusters and outliers show up, and so do the branches that link one cluster to the next.

A medicinal chemist can use it to see families of related compounds, and a biologist to see neighborhoods of similar protein structures. It works the same way on image embeddings or single-cell data.

ii.What I rebuilt

The original TMAP worked, but it was one big C++ codebase with an old neighbor-search layer. TMAP 2.0 is Python and Numba, with a scikit-learn style API and a neighbor search you can swap out. You can use USearch HNSW (cosine, Euclidean or binary Jaccard), fall back to a Numba MinHash + LSH Forest to get the old behavior, or pass in your own kNN graph from MMseqs2, Foldseek or BLAST.

iii.What's new

On a benchmark of 1M points at d=128, recall@20 went from 49% with the old LSH path to about 99% with USearch. The Numba MinHash path is still there for parity, and it runs 2 to 3 times faster than the original C++. It also uses less memory.

You can now filter a map down to a subset, add new points to a map you already have, and use it from Jupyter with much less fuss.

iv.Where it's been used

Most of my testing was on chemistry, but TMAP 2.0 takes anything you can put in a vector. I've used it on 2.7M AlphaFold predicted structures (with a structure viewer inside the map), on a single-cell Arabidopsis atlas, and on collections of image embeddings. It's also being integrated into internal discovery pipelines at AbbVie and Roche.