Support homeAtera Instrument SoftwareAnalysis
Transcript Zarr Output Files

Transcript Zarr Output Files

The Atera Onboard Analysis pipeline generates two transcript data output files.

Back to cell and transcript output file landing page

The transcripts.zarr.zip output file contains data to evaluate transcript quality and localization. It has the following hierarchy of array data:

(root) ├── codeword_category ├── gene_category └── grid ├── 0,0 │ ├── cell_id │ ├── codeword_identity │ ├── gene_offset │ ├── location │ ├── overlaps_nucleus │ └── quality_score └── [X,Y]

Description for root group attributes:

FieldTypeDescription
namestrThe name of the dataset ("RnaDataset").
major_versionintMajor version for this file. This number is increased when breaking changes are made.
minor_versionintMinor version.
dataset_uuidstrUnique ID for this dataset.
data_formatintA field for internal pipeline use. Always set to 0.
number_rnasintThe total number of transcripts in the dataset.
spatial_unitsstrThe units of the stitched image space ("micron").
number_genesintThe number of genes in the dataset (summed over all RNA kits).
gene_nameslist[str]Names of the genes.
codeword_countintThe number of codewords (summed over all RNA kits).
codeword_gene_mappinglist[int]The index of the gene in gene_names specified by each codeword.
codeword_gene_nameslist[str]The name of the gene in gene_names specified by each codeword.
coordinate_spacestrFor internal pipeline use. Should have the value "refined-final_global_micron".
kit_infolist[map]One entry per panel in panel configuration, specified during instrument run set up.

Description for /grid array attributes:

FieldTypeDescription
grid_key_nameslist[str]The names of the grid keys used by the current grid (e.g., "grid_x_loc").
grid_sizefloatThe size of each of the X,Y tiles.
grid_keyslist[str]The grid keys (e.g., "0,0,0") for each of the X,Y tiles.
grid_number_objectslist[int]The number of transcripts in the X,Y tiles.
codeword_countslist[int]Number of codewords, e.g., codeword_counts[i] specifies the count of transcripts that decoded to codeword i.
gene_countslist[int]Number of genes, e.g., gene_counts[i] specifies the count of transcripts that decoded to gene i.

Description for root group arrays:

PathTypeDescription
/codeword_categoryboolA num_codewords x 9 boolean table that contains information about the categories that codewords belong to. column names and descriptions are contained in codeword_category/.zattrs.
/gene_categoryboolA num_genes x 9 boolean table that contains information about the categories that genes belong to. Column names and descriptions are contained in gene_category/.zattrs.
/gridn/aContains every tile group.

Description for /grid array:

PathTypeDescription
/cell_iduint32An array defining the cell each transcript was assigned to. Columns: cell_prefix, dataset_suffix (which matches the cells.zarr.zip); rows: number of transcripts in the tile.
/codeword_identityuint32The codeword index for each RNA. Codeword indices are zero-based and reference the codeword_gene_names attribute attached to the dataset.
/gene_offsetuint32An array defining the [start, end] row range of each gene within this tile, e.g., gene_offset[i] provides the range in the other arrays where data for gene i is located. The data is sorted by gene. Columns: start, end; rows: number of genes.
/locationfloat32The location of each transcript in physical coordinate space. Columns: x_position, y_position, and z_position of the transcript; rows: number of transcripts in the tile.
/overlaps_nucleusuint8An array defining whether the transcript falls inside its assigned cell's nucleus (1 where a cell is assigned).
/quality_scorefloat16The calibrated Q-Score for each transcript in the tile.

The binned_transcripts.zarr.zip file bins transcripts by tile for more efficient display of the spatial distribution of transcript density in a sample (e.g., in 10x Explorer). It has the following hierarchy of array data:

(root) └── gene/{GENE_IX}.zip └── grid/{X},{Y} ├── data # uint16, compressed sparse row data ├── indices # uint16 └── indptr # uint32

The root-level attributes for this file include: gene_counts, gene_names, grid_key_names, grid_keys, grid_size, major_version, and minor_version, name, and origin.

The definitions are the same as in the transcripts.zarr.zip file.

  • The name of the dataset is BinnedTranscripts.
  • The origin (dict[str,float]) is the origin of the grid {"x": min_x, "y": min_y}, in the same units as the point cloud.