Support homeAtera Instrument SoftwareAnalysis
Cell-Feature Matrix Zarr Output Files

Cell-Feature Matrix Zarr Output Files

The Atera Onboard Analysis pipeline generates two cell-feature matrix data output files.

Back to cell and transcript output file landing page

The cell_feature_matrix.zarr.zip output file contains a matrix of counts per cell and per feature (including gene and non-gene codewords), which have passed the default quality value (Q-Score) threshold of Q20, along with secondary analysis annotations:

  • Contains compressed sparse row (CSR) cell-feature matrix
  • Annotates the cells with cluster, Pan-human Azimuth cell types (for human WTA datasets), cell centroids, cell area
  • Contains UMAP and PCA projection coordinates
  • Contains differential expression results by cell clusters and cell types

It has the following hierarchy of array data:

(root) ├── X │ ├── data │ ├── indices │ └── indptr ├── obs │ ├── barcode │ ├── cell_area │ ├── cell_id │ ├── centroid_column │ ├── centroid_row │ ├── control_codeword_counts │ ├── control_probe_counts │ ├── filtered │ ├── leiden_res_1.0 │ ├── azimuth_broad │ ├── azimuth_coarse │ ├── azimuth_fine │ ├── azimuth_full_hierarchy │ ├── azimuth_score │ ├── is_sketch │ ├── nucleus_area │ ├── total_counts │ └── transcript_counts ├── obsm │ ├── X_pca │ └── X_umap └── var ├── de_leiden_<parameters>_logfc ├── de_leiden_<parameters>_pval ├── de_leiden_<parameters>_pval_adj ├── de_leiden_<parameters>_rank ├── de_leiden_<parameters>_score ├── de_azimuth_<cell types>_logfc ├── de_azimuth_<cell types>_pval ├── de_azimuth_<cell types>_pval_adj ├── de_azimuth_<cell types>_rank ├── de_azimuth_<cell types>_score ├── feature_id ├── feature_name ├── feature_type ├── filtered ├── genome └── highly_variable

This output information is stored in a zipped AnnData Zarr format. The structure of the AnnData format includes the primary cell-feature matrix data (X), along with attributes associated with the data.

Description of the cell (rows) by feature (columns; where features are genes) matrix, X/:

PathTypeDescription
/dataint32Compressed sparse row (CSR) cell-feature matrix of non-zero expression count values.
/indicesint64Column (feature) index for each value in data.
/indptrint64Row (cell) pointer array marking each cell's slice of data/indices.

Description of the per-cell observation metadata, obs/:

PathTypeDescription
/barcodestrUnique per cell barcode.
/cell_areafloat32The two-dimensional area covered by the cell in µm2.
/cell_idint64Unique ID of the cell, consisting of a cell prefix and dataset suffix.
/centroid_columnfloat32X location of the cell centroid in µm.
/centroid_rowfloat32Y location of the cell centroid in µm.
/control_codeword_countsint64Count of negative control codewords.
/control_probe_countsint64Molecule count of negative control probes.
/filteredboolWhether the cell was filtered (cell with fewer than 5 transcripts) prior to secondary analysis (true: cell is kept, false: cell is filtered).
/leiden_res_1.0n/aContains /categories (string) with cluster labels and /codes (int8) with indices mapping the categories to each cell.
/azimuth_broadn/aContains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster.
/azimuth_coarsen/aContains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster.
/azimuth_finen/aContains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster.
/azimuth_full_hierarchyn/aContains /categories (string) with cluster labels, /codes (int8) with indices mapping the categories to each cell, and /color indicating a HEX color assigned to the cluster.
/azimuth_scoren/aCalibrated confidence score for cell type annotation.
/is_sketchboolWhether the dataset was sketched (more than 1 million cells) prior to secondary analysis.
/nucleus_areafloat32The two-dimensional area covered by the nucleus in µm2.
/total_countsint64Sum total of transcript_counts, control_probe_counts, control_codeword_counts, genomic_control_counts, and unassigned_codeword_counts.
/transcript_countsint64Molecule count of gene features with Q-Score ≥ 20.

Description of the multi-dimensional observation data, obsm/:

PathTypeDescription
/X_pcafloat32Principal Components Analysis (PCA) embeddings.
/X_umapfloat32Uniform Manifold Approximation and Projection (UMAP) embeddings.

Description of per-feature variable metadata, var/:

PathTypeDescription
/de_leiden_<parameters>_logfcfloat32Log fold-change for unsupervised clusters from differential expression analysis.
/de_leiden_<parameters>_pvalfloat32P-value.
/de_leiden_<parameters>_pval_adjfloat32Adjusted p-value.
/de_leiden_<parameters>_rankfloat32Rank of the gene within a cluster's differential expression results.
/de_leiden_<parameters>_scorefloat32Differential expression statistic for the cluster.
/de_azimuth_<cell types>_logfcfloat32Log fold-change for cell type annotation clusters using Pan-human Azimuth model differential expression results.
/de_azimuth_<cell types>_pvalfloat32P-value.
/de_azimuth_<cell types>_pval_adjfloat32Adjusted p-value.
/de_azimuth_<cell types>_rankfloat32Rank of the gene within a cluster's differential expression results.
/de_azimuth_<cell types>_scorefloat32Differential expression statistic for the cluster.
/feature_idstringGene ID.
/feature_namestringGene name.
/feature_typecategoricalCodeword category.
/filteredboolWhether the gene was filtered.
/genomecategoricalReference genome label (/categories, string) and indices (/codes, int8)
/highly_variableboolWhether the gene was labeled as highly variable in secondary analysis.

The csc_cell_feature_matrix.zarr.zip contains the same cell-feature matrix as the cell_feature_matrix.zarr.zip, but is transposed from compressed sparse row (CSR) format to compressed sparse column (CSC) format. It is primarily used for more efficient visualization of the cell-feature matrix information (e.g., in 10x Explorer).

It has the following hierarchy:

├── X │ ├── data # int32 │ ├── indices # int64 │ └── indptr # int64 ├── obs │ ├── barcode # StringDType() │ └── cell_id # int64 └── var ├── feature_ids # StringDType() ├── feature_name # StringDType() ├── feature_types │ ├── categories # StringDType() │ └── codes # int8 └── genome ├── categories # StringDType() └── codes # int8

The definitions are the same as in the cell_feature_matrix.zarr.zip file.