Support homeAtera Instrument SoftwareAnalysis
Archiving Atera Data

Archiving Atera Data

The Atera platform aims to support and embrace principles of data findability, accessibility, interoperability, and reusability (FAIR) so that it is easy to share newly generated Atera data for collaborative analysis and reproduce findings from published Atera data.

What Atera output should I keep for archival storage for reanalysis and grant funding requirements?

We recommend archiving Atera raw data and metadata outputs, which consist of:

  1. Decoded transcripts with assigned Phred-scaled Q-Scores
  2. High-resolution morphology images
  3. To replicate the exact cell-feature matrix counts, the cell segmentation mask is required from the cells.zarr.zip file.
  4. The analysis_summary.html is useful to archive for data reproducibility (record of metadata) and reanalysis.

Decoded transcripts are provided in Zarr format (transcripts.zarr.zip). Morphology images are provided in OME-TIFF format (morphology_3d/ and morphology_2d/). Cell segmentation masks are provided in Zarr format (cells.zarr.zip). The metadata is provided in the analysis_summary.html. These data should be archived to fulfill grant funding requirements and for reanalysis. They may be submitted to repositories such as GEO or BioImage Archive. All other Atera outputs are derived from these raw data in Atera Onboard Analysis, can be rederived after a Atera instrument run, and are not strictly necessary for long-term archival and reproducibility.

Additional detail on Atera raw data output:

  • An Atera Q-Score indicates the probability that the detected object exists and was correctly identified by the decoding algorithm. All decoded transcript Q-Scores are output in the transcripts file. The cells and cell-feature matrix output files in the Atera output bundle are filtered to Q-Score ≥ 20.
  • Atera morphology images (including DAPI and cell segmentation) will always be provided at the same resolution that our onboard segmentation algorithm uses as input. This ensures that you can benefit from improvements to our segmentation model as we add to its training over time, or run your own segmentation methods if you choose.
  • We will stand by these FAIR principles with future capabilities. High-resolution morphology images will continue to be included in the Atera output bundle for our onboard multimodal segmentation method.
  • Other outputs from Atera Onboard Analysis are derived data from these raw outputs, and the community can recapitulate them from Atera raw data.

Atera raw data reduces low-level internal sensor data as described at Overview of Atera Algorithms. It preserves details needed to assess decoded transcript quality, abstracting away low-level details of the instrumentation and assay that require calibration and specialized methods.

To achieve high-throughput, the Atera Instrument was designed to process the petabytes of internal sensor data acquired every run in memory, instead of writing to disk. In the spirit of scientific reproducibility, it is more useful to store the Atera decoded transcripts with assigned Phred-scaled Q-Scores and morphology images (typical output directory sizes) for reanalysis.

To add further transparency and to supplement existing methods to QC Atera data, downsampled RNA diagnostic images are available in the analysis_summary.html and qc_images.html outputs. These images are not needed for raw data archival, but should be useful in gaining confidence in the robustness of Atera's decoding algorithm and checking sample quality (e.g., for debris).

Each tissue region selected on the Atera Instrument produces a separate output directory with images, decoded transcripts, cell-feature count matrices, and more.

The file formats were carefully designed and chosen to balance compatibility, performance, and file size. There is no simple formula for calculating the output directory size from the Atera Instrument region area alone. Output size also depends on sample-specific factors like tissue shape, number of cells, number of decoded transcripts, and percent of high quality transcripts.

To help budget for data storage requirements, here are some examples based on estimations of datasets generated with Atera Onboard Analysis for FFPE samples.

Atera Human Whole Transcriptome Assay Panel dataset estimates:

RegionTypical (20 tx/µm2)Transcript-dense (35 tx/µm2)
1 cm280 GB120 GB
Typical slide (4 cm2)320 GB480 GB
Full slide (6 cm2)500 GB750 GB
Full run (4 slides x 6 cm2)2 TB3 TB

Atera Select Human Multi-Tissue Panel dataset estimates:

RegionTypical (2 tx/µm2)Transcript-dense (4 tx/µm2)
1 cm240 GB50 GB
Typical slide (4 cm2)160 GB200 GB
Full slide (6 cm2)240 GB300 GB
Full run (4 slides x 6 cm2)960 GB1.2 TB