spatialrefinery.xenium_to_spatialdata#
- spatialrefinery.xenium_to_spatialdata(dataset_path, output_path, output_name=None, include_aligned_image=True, create_spots=True, spot_sizes=None, overlap=0.0, values='transcripts', n_jobs=1, overwrite=False, max_spots_per_chunk=50000)#
Convert 10x Xenium raw data to SpatialData zarr format.
This is the main function for converting Xenium bundled outputs to a SpatialData object and saving it as a zarr file. It can optionally include aligned H&E images and create pseudo-spots at specified sizes.
- Parameters:
dataset_path (
str|Path) – Path to the Xenium dataset directory containing raw files.output_path (
str|Path) – Path to the directory where the zarr file will be saved.output_name (
str|None(default:None)) – Name for the output zarr file (without .zarr extension). If None, uses the dataset directory name. Default is None.include_aligned_image (
bool(default:True)) – Whether to include aligned H&E image if available. Default is True.create_spots (
bool(default:True)) – Whether to create pseudo-spots. Default is True.spot_sizes (
list|None(default:None)) – List of spot sizes in micrometers to create. Default is None. Spots are only created when bothcreate_spotsis True andspot_sizesis a non-empty list – the default combination (create_spots=True,spot_sizes=None) creates no spots.overlap (
float|None(default:0.0)) – Fractional overlap between adjacent hexagonal spots, forwarded tocreate_pseudo_spots. Default is 0.0 (no overlap).values (
str(default:'transcripts')) – Which SpatialData element to aggregate into spots:"transcripts"or"cell_boundaries". Default is"transcripts".n_jobs (
int(default:1)) – Number of workers for parallel processing. Default is 1.overwrite (
bool(default:False)) – Whether to overwrite existing zarr file. Default is False.max_spots_per_chunk (
int|None(default:50000)) – Retained for backwards compatibility and ignored; transcript aggregation is now streamed. Seecreate_pseudo_spots.
- Return type:
- Returns:
Path Path to the created zarr file.
- Raises:
FileNotFoundError – If
dataset_pathdoes not contain anexperiment.xeniumfile.
Notes
The cell table’s
varis the sample’s targeted gene panel and nothing else. The control and codeword feature types the cell-feature matrix also carries are dropped, their per-cell totals already being inobs. Panel sizes differ between samples, so this is a guarantee about whatvarholds, not how much.A protein sub-panel is handled without changing that shape. Each antibody’s per-cell measurement is kept in
table.obsm["protein_expression"], a cells x antibodies DataFrame columned by antibody name, withtable.uns["protein_expression"]recording theirnames,gene_idsandmetric. Samples with no antibodies simply get no such entry, so downstream code never has to branch on whether a sub-panel was run. Those values areMEAN_PER_CELL_STAINintensities rather than transcript counts and must not be normalised as counts; the pseudo-spot tables carry genes only, there being nothing intranscriptsto aggregate for an antibody stain. The private_split_feature_typeshelper does this and carries the full rationale.Examples
>>> from spatialrefinery import xenium_to_spatialdata >>> zarr_path = xenium_to_spatialdata( ... dataset_path="/path/to/xenium/data", ... output_path="/path/to/output", ... output_name="my_sample", ... create_spots=True, ... spot_sizes=[55, 100], ... )