The Atera Onboard Analysis pipeline generates two cell-feature matrix data output files.
The cell_feature_matrix.zarr.zip output file contains a matrix of counts per cell and per feature (including gene and non-gene codewords), which have passed the default quality value (Q-Score) threshold of Q20, along with secondary analysis annotations:
- Contains compressed sparse row (CSR) cell-feature matrix
- Annotates the cells with cluster, Pan-human Azimuth cell types (for human WTA datasets), cell centroids, cell area
- Contains UMAP and PCA projection coordinates
- Contains differential expression results by cell clusters and cell types
It has the following hierarchy of array data:
(root)
├── X
│ ├── data
│ ├── indices
│ └── indptr
├── obs
│ ├── barcode
│ ├── cell_area
│ ├── cell_id
│ ├── centroid_column
│ ├── centroid_row
│ ├── control_codeword_counts
│ ├── control_probe_counts
│ ├── filtered
│ ├── leiden_res_1.0
│ ├── azimuth_broad
│ ├── azimuth_coarse
│ ├── azimuth_fine
│ ├── azimuth_full_hierarchy
│ ├── azimuth_score
│ ├── is_sketch
│ ├── nucleus_area
│ ├── total_counts
│ └── transcript_counts
├── obsm
│ ├── X_pca
│ └── X_umap
└── var
├── de_leiden_<parameters>_logfc
├── de_leiden_<parameters>_pval
├── de_leiden_<parameters>_pval_adj
├── de_leiden_<parameters>_rank
├── de_leiden_<parameters>_score
├── de_azimuth_<cell types>_logfc
├── de_azimuth_<cell types>_pval
├── de_azimuth_<cell types>_pval_adj
├── de_azimuth_<cell types>_rank
├── de_azimuth_<cell types>_score
├── feature_id
├── feature_name
├── feature_type
├── filtered
├── genome
└── highly_variable
This output information is stored in a zipped AnnData Zarr format. The structure of the AnnData format includes the primary cell-feature matrix data (X), along with attributes associated with the data.
Description of the cell (rows) by feature (columns; where features are genes) matrix, X/:
| Path | Type | Description |
|---|---|---|
/data | int32 | Compressed sparse row (CSR) cell-feature matrix of non-zero expression count values. |
/indices | int64 | Column (feature) index for each value in data. |
/indptr | int64 | Row (cell) pointer array marking each cell's slice of data/indices. |
Description of the per-cell observation metadata, obs/:
| Path | Type | Description |
|---|---|---|
/barcode | str | Unique per cell barcode. |
/cell_area | float32 | The two-dimensional area covered by the cell in µm2. |
/cell_id | int64 | Unique ID of the cell, consisting of a cell prefix and dataset suffix. |
/centroid_column | float32 | X location of the cell centroid in µm. |
/centroid_row | float32 | Y location of the cell centroid in µm. |
/control_codeword_counts | int64 | Count of negative control codewords. |
/control_probe_counts | int64 | Molecule count of negative control probes. |
/filtered | bool | Whether the cell was filtered (cell with fewer than 5 transcripts) prior to secondary analysis (true: cell is kept, false: cell is filtered). |
/leiden_res_1.0 | n/a | Contains /categories (string) with cluster labels and /codes (int8) with indices mapping the categories to each cell. |
/azimuth_broad | n/a | Contains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster. |
/azimuth_coarse | n/a | Contains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster. |
/azimuth_fine | n/a | Contains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster. |
/azimuth_full_hierarchy | n/a | Contains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster. |
/azimuth_score | n/a | Calibrated confidence score for cell type annotation. |
/is_sketch | bool | Whether the dataset was sketched (more than 1 million cells) prior to secondary analysis. |
/nucleus_area | float32 | The two-dimensional area covered by the nucleus in µm2. |
/total_counts | int64 | Sum total of transcript_counts, control_probe_counts, control_codeword_counts, genomic_control_counts, and unassigned_codeword_counts. |
/transcript_counts | int64 | Molecule count of gene features with Q-Score ≥ 20. |
Description of the multi-dimensional observation data, obsm/:
| Path | Type | Description |
|---|---|---|
/X_pca | float32 | Principal Components Analysis (PCA) embeddings. |
/X_umap | float32 | Uniform Manifold Approximation and Projection (UMAP) embeddings. |
Description of per-feature variable metadata, var/:
| Path | Type | Description |
|---|---|---|
/de_leiden_<parameters>_logfc | float32 | Log fold-change for unsupervised clusters from differential expression analysis. |
/de_leiden_<parameters>_pval | float32 | P-value. |
/de_leiden_<parameters>_pval_adj | float32 | Adjusted p-value. |
/de_leiden_<parameters>_rank | float32 | Rank of the gene within a cluster's differential expression results. |
/de_leiden_<parameters>_score | float32 | Differential expression statistic for the cluster. |
/de_azimuth_<cell types>_logfc | float32 | Log fold-change for cell type annotation clusters using Pan-human Azimuth model differential expression results. |
/de_azimuth_<cell types>_pval | float32 | P-value. |
/de_azimuth_<cell types>_pval_adj | float32 | Adjusted p-value. |
/de_azimuth_<cell types>_rank | float32 | Rank of the gene within a cluster's differential expression results. |
/de_azimuth_<cell types>_score | float32 | Differential expression statistic for the cluster. |
/feature_id | string | Gene ID. |
/feature_name | string | Gene name. |
/feature_type | categorical | Codeword category. |
/filtered | bool | Whether the gene was filtered. |
/genome | categorical | Reference genome label (/categories, string) and indices (/codes, int8) |
/highly_variable | bool | Whether the gene was labeled as highly variable in secondary analysis. |
The csc_cell_feature_matrix.zarr.zip contains the same cell-feature matrix as the cell_feature_matrix.zarr.zip, but is transposed from compressed sparse row (CSR) format to compressed sparse column (CSC) format. It is primarily used for more efficient visualization of the cell-feature matrix information (e.g., in 10x Explorer).
It has the following hierarchy:
├── X
│ ├── data # int32
│ ├── indices # int64
│ └── indptr # int64
├── obs
│ ├── barcode # StringDType()
│ └── cell_id # int64
└── var
├── feature_ids # StringDType()
├── feature_name # StringDType()
├── feature_types
│ ├── categories # StringDType()
│ └── codes # int8
└── genome
├── categories # StringDType()
└── codes # int8
The definitions are the same as in the cell_feature_matrix.zarr.zip file.