The Atera Onboard Analysis pipeline generates two transcript data output files.
The transcripts.zarr.zip output file contains data to evaluate transcript quality and localization. It has the following hierarchy of array data:
(root)
├── codeword_category
├── gene_category
└── grid
├── 0,0
│ ├── cell_id
│ ├── codeword_identity
│ ├── gene_offset
│ ├── location
│ ├── overlaps_nucleus
│ └── quality_score
└── [X,Y]
Description for root group attributes:
| Field | Type | Description |
|---|---|---|
name | str | The name of the dataset ("RnaDataset"). |
major_version | int | Major version for this file. This number is increased when breaking changes are made. |
minor_version | int | Minor version. |
dataset_uuid | str | Unique ID for this dataset. |
data_format | int | A field for internal pipeline use. Always set to 0. |
number_rnas | int | The total number of transcripts in the dataset. |
spatial_units | str | The units of the stitched image space ("micron"). |
number_genes | int | The number of genes in the dataset (summed over all RNA kits). |
gene_names | list[str] | Names of the genes. |
codeword_count | int | The number of codewords (summed over all RNA kits). |
codeword_gene_mapping | list[int] | The index of the gene in gene_names specified by each codeword. |
codeword_gene_names | list[str] | The name of the gene in gene_names specified by each codeword. |
coordinate_space | str | For internal pipeline use. Should have the value "refined-final_global_micron". |
kit_info | list[map] | One entry per panel in panel configuration, specified during instrument run set up. |
Description for /grid array attributes:
| Field | Type | Description |
|---|---|---|
grid_key_names | list[str] | The names of the grid keys used by the current grid (e.g., "grid_x_loc"). |
grid_size | float | The size of each of the X,Y tiles. |
grid_keys | list[str] | The grid keys (e.g., "0,0,0") for each of the X,Y tiles. |
grid_number_objects | list[int] | The number of transcripts in the X,Y tiles. |
codeword_counts | list[int] | Number of codewords, e.g., codeword_counts[i] specifies the count of transcripts that decoded to codeword i. |
gene_counts | list[int] | Number of genes, e.g., gene_counts[i] specifies the count of transcripts that decoded to gene i. |
Description for root group arrays:
| Path | Type | Description |
|---|---|---|
/codeword_category | bool | A num_codewords x 9 boolean table that contains information about the categories that codewords belong to. column names and descriptions are contained in codeword_category/.zattrs. |
/gene_category | bool | A num_genes x 9 boolean table that contains information about the categories that genes belong to. Column names and descriptions are contained in gene_category/.zattrs. |
/grid | n/a | Contains every tile group. |
Description for /grid array:
| Path | Type | Description |
|---|---|---|
/cell_id | uint32 | An array defining the cell each transcript was assigned to. Columns: cell_prefix, dataset_suffix (which matches the cells.zarr.zip); rows: number of transcripts in the tile. |
/codeword_identity | uint32 | The codeword index for each RNA. Codeword indices are zero-based and reference the codeword_gene_names attribute attached to the dataset. |
/gene_offset | uint32 | An array defining the [start, end] row range of each gene within this tile, e.g., gene_offset[i] provides the range in the other arrays where data for gene i is located. The data is sorted by gene. Columns: start, end; rows: number of genes. |
/location | float32 | The location of each transcript in physical coordinate space. Columns: x_position, y_position, and z_position of the transcript; rows: number of transcripts in the tile. |
/overlaps_nucleus | uint8 | An array defining whether the transcript falls inside its assigned cell's nucleus (1 where a cell is assigned). |
/quality_score | float16 | The calibrated Q-Score for each transcript in the tile. |
The binned_transcripts.zarr.zip file bins transcripts by tile for more efficient display of the spatial distribution of transcript density in a sample (e.g., in 10x Explorer). It has the following hierarchy of array data:
(root)
└── gene/{GENE_IX}.zip
└── grid/{X},{Y}
├── data # uint16, compressed sparse row data
├── indices # uint16
└── indptr # uint32
The root-level attributes for this file include: gene_counts, gene_names, grid_key_names, grid_keys, grid_size, major_version, and minor_version, name, and origin.
The definitions are the same as in the transcripts.zarr.zip file.
- The
nameof the dataset isBinnedTranscripts. - The
origin(dict[str,float]) is the origin of the grid{"x": min_x, "y": min_y}, in the same units as the point cloud.