# About

Power your multiomics journey with Illumina Multiomics Software.

## :microscope:Discover Illumina Multiomics

**Illumina Multiomics Software** offers a comprehensive range of solutions to streamline your journey from sample collection to multiomic insights. Discover assays designed for protein analysis, spatial data, single-cell RNA data, and more.

Our powerful study management tool helps organize your data from ingestion to analysis, while our advanced analysis platform generates actionable insights to drive your research forward.

Select your assay of interest below to learn about how to use it.

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><span data-gb-custom-inline data-tag="emoji" data-code="1f9ea">🧪</span> Assay that uses DRAGEN Protein Quantification for analysis</td><td><a href="/files/mdoFhKcoM612CHMr1gSt">/files/mdoFhKcoM612CHMr1gSt</a></td><td><a href="/spaces/nyQb4WG1K4VQZKovLv5v">/spaces/nyQb4WG1K4VQZKovLv5v</a></td></tr><tr><td><span data-gb-custom-inline data-tag="emoji" data-code="1f9ec">🧬</span> Assay that uses DRAGEN Single Cell RNA for analysis</td><td><a href="/files/g86ocFWKH0V37fqws48t">/files/g86ocFWKH0V37fqws48t</a></td><td><a href="/spaces/qVEYIKB8JFfdScsTocFN">/spaces/qVEYIKB8JFfdScsTocFN</a></td></tr><tr><td><span data-gb-custom-inline data-tag="emoji" data-code="1f465">👥</span> Software for study management, sample group creation, and running analyses</td><td><a href="/files/fHaNBqOzrm9gYwGN6K24">/files/fHaNBqOzrm9gYwGN6K24</a></td><td><a href="/spaces/WMxqQAMFOJtu98OBk9KN">/spaces/WMxqQAMFOJtu98OBk9KN</a></td></tr></tbody></table>

Explore our [**End-to-End Tutorials**](https://github.com/illumina-swi/icm-docs/blob/multiomics-prod/docs/broken-reference/README.md) for an introductory experience of Multiomics workflows.

{% hint style="info" %}
You can search our help documentation or ask questions with AI-generated answers using the search-box in the top of the page.

Navigate and explore using the left-panel.
{% endhint %}

{% embed url="<https://www.youtube.com/watch?v=Xsyp0qqKzkY>" %}
Multiomics
{% endembed %}


# Multiomics Workflows

Dive into our end-to-end tutorials for a guided, hands-on journey through powerful multiomics workflows.

### Proteomics Workflow

[Protein Quantification End-to-End Workflow with Connected Multiomics](/dragen-protein-quantification/after-counting-and-normalization/illumina-connected-multiomics-walkthrough)

{% embed url="<https://youtu.be/eSXYiB39sHk?si=DLw1d4xFa-O4NN2k>" %}
How to analyze data from an Illumina Protein Prep kit in Illumina Connected Multiomics
{% endembed %}

### Single Cell Workflow

[Single Cell End-to-End Workflow with Connected Multiomics](/dragen-single-cell-rna/tertiary-analysis/illumina-connected-multiomics)

{% embed url="<https://youtu.be/he2tEiDrjkI?si=iZ8Z0ebNxEt6tMRI>" %}
**How to analyze single cell transcriptomics data in Illumina Connected Multiomics**
{% endembed %}

### Spatial Workflow

{% embed url="<https://youtu.be/OEaEYgQJwHQ?si=3XF97uBGjL6lDMQ9>" %}
**How to analyze spatial data in Illumina Connected Multiomics**
{% endembed %}


# Bulk RNA

## Getting started

### [Logging into Connected Multiomics](https://help.multiomics.illumina.com/icm)

### [Creating a study and adding projects from Illumina Connected Analytics](https://help.multiomics.illumina.com/icm/studies/create-study)

### [Viewing results and navigating in Connected Multiomics](https://help.multiomics.illumina.com/icm/analyses/enter-analysis)

### Data / Task nodes and Performing tasks in Connected Multiomics

* Within a study, the *Analyses tab* contains two elements: task nodes (rectangles) and data nodes (circles) connected by lines and arrows. Collectively, they represent a data analysis pipeline.
* Clicking a data node brings up a context sensitive menu on the right. This menu changes depending on the type of data node. It will only present tasks which can be performed on that specific data type. Hover over the task to obtain additional information regarding each option.
* Select the task you wish to perform from the menu. When configuring task options, additional information regarding each option is available. Click **Finish** to perform the task.
* Depending on the task, a new data node may automatically be created and connected to the original data node. This contains the data resulting from the task. Tasks that do not produce new data types will not produce an additional data node.
* To view the results of a task, click the data node and choose the **Task report** option on the menu.

### Viewing and saving data

* All data contained in data nodes can be downloaded to the local machine by selecting the node and navigating to the bottom of the toolbox then choose **Download data**.
* The [Data Viewer](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/BEpNAu7pFMMuYvu7XKHz) can be used to plot, modify, and save data. In this walkthrough the PCA data node and Hierarchical clustering / heatmap node can be automatically opened in the Data viewer by double-clicking the data node or opening the Task report from the toolbox.
* To save an individual image within the Data Viewer to your machine, click **Plot** then **Export image** & select the format, size, and resolution then click **Save**. Use the plot-specific tools for this.
* All visualizations within a sheet in the Data Viewer can be exported as one image (e.g. use one image with all plots for a poster). Use the **Export** drop-down at the top of the data-viewer for this and select **Export image**.

## Input: secondary outputs from the DRAGEN analysis

You will noticed that there are two .sf file options to choose from in [secondary outputs of the DRAGEN analysis](https://help.dragen.illumina.com/product-guides/dragen-v4.3/dragen-rna-pipeline/gene-expression-quantification).

* `<outputPrefix>.quant.genes.sf` - Contains quantification results at the gene level. The results are produced by summing together all transcripts with the same geneID in the annotation file (GTF).
* `<outputPrefix>.quant.sf` - Contains quantification results at the transcript level.

## Import Data

Import data that has been processed through the DRAGEN RNA analysis pipeline in [BaseSpace](https://ilmn.basespace.illumina.com/apps/17802785/DRAGEN%20RNA), Illumina Connected Analytics, or the command line.

* Use the .sf file from the secondary outputs.

{% hint style="warning" %}
Remember the [genome reference and annotation file used during DRAGEN analysis](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)—you’ll need them for feature annotation. If you have used built in genome references, the following table shows the [default GTFs](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) being used.
{% endhint %}

| Annotation file (GTF) | Genome reference file                                                                                                                                                |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GENCODE v19           | <p>Homo sapiens \[UCSC] hg19 v5</p><p>Homo sapiens \[UCSC] hg19 v5 Pangenome</p><p>Homo sapiens \[NCBI] hs37d5 v5</p><p>Homo sapiens \[NCBI] hs37d5 v5 Pangenome</p> |
| GENCODE v44           | <p>Homo sapiens \[1000 Genomes] hg38 v5</p><p>Homo sapiens \[1000 Genomes] hg38 v5 Pangenome</p>                                                                     |
| GENCODE vM23          | Mus musculus \[UCSC] mm10                                                                                                                                            |
| ENSEMBL 98            | Rattus norvegicus \[UCSC] rn6                                                                                                                                        |

* After creating a study and adding data to the study, click **+ New Analysis**
* Give the Analysis a name, select the **Analysis Type > Custom: RNA**, select the sample groups to add to the analysis, and click **Run Analysis**

<figure><img src="/files/SWzg2CTRp5byXAidZCeA" alt=""><figcaption></figcaption></figure>

* The Status will show as Complete when ready to analyze

{% hint style="warning" %}
Click the Refresh button to see change the Status in real-time.
{% endhint %}

* Click the complete Analysis to open and customize the analysis pipeline

<figure><img src="/files/aea26puexyNMznSKTLl6" alt=""><figcaption></figcaption></figure>

* The Quantification node is created when the analysis completes.

{% hint style="warning" %}
Hover over nodes to see details about the data node. Below, the number of samples, features, and data size is shown in the Quantification node.
{% endhint %}

<figure><img src="/files/0xsVR0ZkVLNpVLPS0PaA" alt=""><figcaption><p>The Quantification node is the starting node for analysis</p></figcaption></figure>

## Annotate Features

Add gene-level annotations to the quantified data.

* Single-click the *Quantification* node.
* Select **Annotate features** under the *Pre-analysis tools* section in the toolbox on the right.
* Choose the **genome** and **annotation** files that match those used in DRAGEN then click **Finish**.
* Outcome:
  * Task node: *Annotate features*
  * Result node: *Annotated counts*

<figure><img src="/files/iytCrjAIhWplhSUcKaxW" alt=""><figcaption><p>Annotate feature task node &#x26; Annotated counts data node</p></figcaption></figure>

## [Normalization and Scaling](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/JgCD3TueS7Rk88X2nnV2)

Normalize the data to prepare for downstream analysis.

* Single-click the Annotated Counts node, then select the **Normalization** task from the *Normalization and Scaling* section.
* Click the **"Use Recommended"** button or select an alternative method. We recommend the widely used *Median ratio (DESeq2 only)* method.

<figure><img src="/files/YmVl4w4UUg6b0obvVpjD" alt=""><figcaption><p>Median ratio (DESeq2 only) is the recommended normalization method for bulk transcriptomic data</p></figcaption></figure>

* Outcome:
  * Task node: *Normalize counts*
  * Result node: *Normalized counts*

<figure><img src="/files/axIaD7nMfTHzwyQr56Ya" alt=""><figcaption><p>Normalize counts task node &#x26; Normalized counts data node</p></figcaption></figure>

## [Dimension Reduction (PCA)](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/DRlU8IxPKKSKAmYgbKZF)

Visualize sample clustering and variance.

* From the *Normalized Counts* node, select **PCA** under *Exploratory Analysis*.
* Outcome:
  * Task node: *PCA*
  * Result node: *PCA*

<figure><img src="/files/X7z5aESgLysq7O3wFWm5" alt=""><figcaption><p>PCA task node &#x26; PCA data node</p></figcaption></figure>

## [Differential Analysis](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/GIo34pINUdKUJJOJpFYY)

Compare gene expression across experimental groups.

* From the *Normalized counts* node, select **Differential Analysis** from the *Statistics* section.
* Choose your preferred model and set up the comparison. Note that we have chosen the [DESeq2 method](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/71qP0wdPqQaJ9aO0DnP2) and used the corresponding normalization prior.
* Outcome:
  * Task node: *Differential analysis* (labeled as model used)
  * Result node: *Differential results* (labeled as comparison made)

<figure><img src="/files/55GjPnrUN3oih4ms0ZRZ" alt=""><figcaption><p>Differential analysis task node &#x26; Differential analysis data node labeled as model &#x26; comparison made</p></figcaption></figure>

## Filter Feature List

Refine the list of genes/features based on criteria.

* Open the *Differential Results* node (double-click or single-click and select *Task report* from the toolbox).
* Use the **filter menu** to apply criteria relevant to your study.
* Click **Generate filtered node** once satisfied.
* Outcome:
  * Task node: *Filter list*
  * Result node: *Filtered feature list*

<figure><img src="/files/13CQm08eOd2DGg4jxd8q" alt=""><figcaption><p>Filter list task node &#x26; Filtered feature list data node</p></figcaption></figure>

## [Gene Set Enrichment](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/sLeiBmT7kMk2GhUW0XED)

Identify enriched biological pathways or gene sets.

* Select **Gene Set Enrichment** from the *Biological Interpretation* section.
* Choose between **KEGG Pathway Enrichment** or **Gene Set Ontology**.

{% hint style="warning" %}
The latest version of KEGG can be added in the **Settings > Library file management**
{% endhint %}

* Outcome:
  * Task node: *Gene set enrichment*
  * Result node: *Pathway enrichment*

<figure><img src="/files/XhLQnfZyc2HtjlkiqQ9R" alt=""><figcaption><p>Gene set enrichment task node &#x26; Pathway enrichment data node</p></figcaption></figure>

## [Hierarchical clustering / Heatmap](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/eJvtfsrz6cbICIjovtqB)

Visualize features in an informative way.

* Select **Hierarchical clustering / Heatmap** from the *Exploratory analysis* section.
* This task can be used for either a heatmap or bubble map. Choose the task options that best suite your needs.
* Double-click on the output node to visualize the results in the Data viewer.
* Outcome:
  * Node: *Hierarchical clustering / heatmap*

<figure><img src="/files/ZCwEl2PQdSSYOxVbGcJi" alt=""><figcaption><p>Hierarchical clustering / heatmap result</p></figcaption></figure>


# 5-base DNA

## Getting Started

[Logging into ICM](https://help.multiomics.illumina.com/icm)

[Creating a Study from a ICA Project](https://help.multiomics.illumina.com/icm/studies/create-study)

[Viewing Results and Navigating in ICM](https://help.multiomics.illumina.com/icm/analyses/enter-analysis)

## Demo Data

Demo data that can be used to follow along with this walkthrough is found in the Connected Multiomics Demo Data repository. The dataset can be found at /Multiomics-Demo-Data/Methylation/Illumina 5-base-solution. This demo data consists of 6 samples, from two pheonotype groups. In this walkthrough, we outline analysis steps that can be performed to explore the data, identify differentially methylated regions between the two sample groups, and find pathways overrepresented in the differential test result.

5 files per sample are required to analyze the DNA Methylation Prep data in the Connected Multiomics software. Add the following 5 files for each sample from the demo data folder to a study prior to starting an analysis:

* \<sample name>.CX\_report.txt.gz
* \<sample name>.methyl\_metrics.csv
* \<sample name>.mapping\_metrics.csv
* \<sample name>.wgs\_coverage\_metrics.csv
* \<sample name>.M-bias.txt

These files are generated from DRAGEN analysis. [CX\_report file](https://illumina.gitbook.io/dna-methylation-prep/8yf1CuwpRlzpAWMSt3ib/additional-information/key-output-files-and-metrics) is the key output file that contains methylation reads count at single nucleotide level. The metrics files and the M-bias file contain QC metrics for reads mapping quality and methylation calling, which will be used to generate visualizations in 5-base Methylation QC task in the Connected Multiomics.

## Custom 5-base Methylation Analysis

<figure><img src="/files/eM0kgv6mgS3NPX4Xz52s" alt=""><figcaption></figcaption></figure>

### Creating a Custom Analysis

After all 6 samples are added into a study, follow these steps to create a Custom analysis in the Connected Multiomics:

* Click on **+ New Analysis.**
* In the pop-up window, provide a name for the analysis, select **Custom: 5-base Methylation** as the Analysis Type, choose a sample group to be included in the analysis (all samples option is selected by default), and click on the **Run Analysis** button.

<figure><img src="/files/HaOtwcC5tiEAa3gn8uM2" alt=""><figcaption></figcaption></figure>

* Refresh the page to get the latest status of the analysis.
* When the Status is Complete, it indicates that launching the analysis has started, click on the analysis tile to enter the analysis module. You will see an ongoing **Import cohort** task that is importing the data into this analysis. After the **Import cohort** task is completed, the first data node called **5-base Methylation** is generated.
* To review the number of samples and features, hover over the **5-base Methylation** data node. Features refer to CpG sites.

<figure><img src="/files/o3osTYWV6PzHGKLKGJgp" alt=""><figcaption></figcaption></figure>

\
The 5-base Methylation data node contains raw methylated counts and unmethylated counts for CpG sites present in the CX reports. For sites with methylation calls on both strands in the CX report, the strands are collapsed such that poisiton on the postive strand is used and the methylation counts are summed. This data node also contains percent methylation levels which will be used in exploratory analysis such as Principal Component Analysis (PCA).

### Add Sample Metadata

We use **Metadata** tab within an analysis to manage sample metadata. Follow these steps to create a new sample attribute called sampleGroup, and assign attribute value to each sample:

* Click on **Metadata** tab. In **Sample attributes** menu on the left, click **Manage**.
* In the Manage sample attributes page, click **Add new attribute**. Type in **sampleGroup** in the **Name** text box, click **Add**.

<figure><img src="/files/POJTsE4CsuHTKpuTxIw5" alt=""><figcaption></figcaption></figure>

* Click **+** button to add two category values **A, B** to the sampleGroup attribute.

<figure><img src="/files/B6d8ViOn79bDciqk8Pto" alt=""><figcaption></figcaption></figure>

* Click **Back to metadata tab**.
* Click **Assign values** under **Sample attributes**. Use dropdown at each sample to assign a category value for the sampleGroup attribute. Assign value for each sample as screenshot below.

<figure><img src="/files/fFhmpW65lVso3g98Mi82" alt=""><figcaption></figcaption></figure>

* Click **Apply changes** to save the assigned values.

### 5-base Methylation QC

The **5-base methylation QC** task in the Connected Multiomics enables you to visualize sample-level QC metrics that describe reads mapping quality and CpG methylation calling. The QC metrics are extracted from the DRAGEN analysis metric files that were ingested into the study. To invoke the **5-base methylation QC** task:

* At **Analyses** page, click on the **5-base Methylation** node.
* Click **QA/QC** section in the context-sensitive task menu on the right.
* Click **5-base methylation QC**.

After the **5-base methylation QC** task is completed, double-click on the task node to open the QC report in a data viewer. The QC report consists of plots and tables organized in 2 sheets. Click sheet name at the bottom of the data viewer to navigate from one sheet to another.

<figure><img src="/files/2TpjKz1izD9h8a6MRGdJ" alt=""><figcaption></figcaption></figure>

* Sheet **Metrics** shows sample-level QC metrics plot. Each sample is a data point, they are randomly spead out on x-axis. The QC metric is represented by y-axis. Each plot is overlay with a violin plot to show distribution of the QC metrics.
  * Percent methylation in samples: Percentages of CpG methylation in samples.
  * Percent methylation in unmethylated control: Percentage of CpG methylation in the unmethylated control (lambda). Low value indicates good quality.
  * Percent methylation in methylated control: Percentage of CpG methylation in the methylated control (pUC19). High value indicates good quality.
  * Percent duplicate reads: Percentage of duplicate marked reads, as a result of PCR amplification.
  * Percent mapped reads: Percentage of mapped reads, indicate the alignment rate.
  * Average autosomal coverage: Mean autosomal coverage across the whole genome. Higher coverage indicates the counts of methylated/unmethylated more accurately reflects the true methylation amount at any particular site.
  * QC metrics table: Text representations of the QC metrics plots.

<figure><img src="/files/CrNBEOoCRxdaLU80UFuT" alt=""><figcaption></figcaption></figure>

* Sheet **M-bias** shows M-bias plots for methylation level and coverage across positions on read1 and read2. The M-bias should be consistent across all positions. It is common for the first/last 10 bases to have un-even methylation due to end-repair and sequencing artifacts.

### PCA

The principal components analysis (PCA) scatter plot allows us to visualize similarities and differences between the samples in a dataset. To invoke a PCA task:

* Click on the **5-base Methylation** node.
* Click **Exploratory analysis** section in the context-sensitive task menu.
* Click **PCA**.
* Set to use the top **100,000** features with the highest **variance** in calculation.
* Keep the rest of the parameters as default, and click **Finish**.

<figure><img src="/files/BtxKy0C2NP2eCG6tCIPx" alt=""><figcaption></figcaption></figure>

After the PCA task is completed, double click on the **PCA** node to view the PCA plot in a data viewer.

<figure><img src="/files/B4nDuLHFpDQyaHUlsrpD" alt=""><figcaption></figcaption></figure>

* The scatter plot shows the data distribution among the first three PCs. Each sample is a data point.
* The scree plot (top right panel) shows variance represented by each PC.
* The component loading table (bottom right panel) shows the correlation between CpG methylation sites and PC.
* For additional information on PCA, refer to the [PCA documentation](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/pca).

### Detect Differentially Methylated Regions (DMRs)

[DSS ](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/onVXlqYRra4JPW2pj7bS)(Dispersion Shrinkage for Sequencing data) enables the detection of differentially methylation regions using counts data at single nucleotide level. It uses beta-binomial distribution to model methylation counts at each CpG site and uses Wald test to identify differentially methylation loci (DML). Nearby DMLs are then merged into a region to form differentially methylated region (DMR). Set up a DSS task to identify DMRs between two sample groups:

* Click on the **5-base Methylation** node.
* Click **Statistics** section in the context-sensitive task menu.
* Click **Differential Methylation**.
* Select **DSS** as the Method to use for differential methylation analysis, click **Next.**
* Select **sampleGroup** as factor for analysis, click **Next.**

<figure><img src="/files/q1SklcQ9stlV8EUx9zLs" alt=""><figcaption></figcaption></figure>

* Drag **A** to the top right Numerator box and **B** to the bottom right Denominator box. Click on **Add comparison**.
* Keep the rest of the settings as default, then click **Finish**.

<figure><img src="/files/6HT6JRKtqvUJ2Uwi4i7f" alt=""><figcaption></figcaption></figure>

When the DSS task is completed, double click on the **A vs B (DMR)** node to open the DMR report. The DSS DMR task report lists regions on rows and the test statistics (areaStat, diff.Methy, etc.) on columns. Regions are listed in descending order by the abs(areaStat) so that the most significant DMR is listed first. diff.Methy statistics reports the difference in average methylation between the two groups, negative value indicates A is hypomethylated compared to B in the region, while positive value indicates A is hypermethylated compared to B in the region. Refer to [DSS documentation](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/onVXlqYRra4JPW2pj7bS) to learn more about the differential methylation report.

On the DMR report, click on the **volcano icon** ( ![](/files/IBVAyQSZVGSksLDXMJWw) ) next to the comparison name to open a differential methylation plot in a Data Viewer. Each data point in the plot is a region. The plot can be colored based on user-defined hypo- and hypermethylation thresholds:

* Click anywhere within the plot canvas on the top panel to select the plot.
* Click **Configure** icon on the left, click **Style**. In the Style dialog, set **Color by** option to **Significance**.
* Click **Configure** icon on the left, click **Statistics**. In the Statistics dialog, set **X threshold** to **-0.2** and **0.2**. Drag **Y threshold** sliding bar to maximum.

The regions are now colored in the volcano plot. Hypomethylated regions (diff.Methy < -0.2) are colored in blue, hypermethylated regions (diff.Methy > 0.2) are colored in red.

<figure><img src="/files/CfsmgFvHbClkHK5cQbIL" alt=""><figcaption></figcaption></figure>

### Filter DMRs

We recommend filtering DMRs by hypo- or hypermethylation status, using the diff.Methy statistics, to give the necessary context of which pathways are hypo- or hypermethylated from the differential comparison. To filter DMR results to hypermethylated DMRs,

* Click **A vs B (DMR)** node.
* Click **Filtering** section in the context-sensitive task menu.
* Click **Differential analysis filter**.
* Choose **Metadata** as Filter type.
* In Filter criteria section, set Filter features by **include A vs B: diff.Methy > 0.2**, then click **Finish**.

<figure><img src="/files/FR51CLXqPkBl0I9Yu8NI" alt=""><figcaption></figcaption></figure>

This generates a Filtered features list node that contains DMRs passing the filtering criteria. Same steps can be applied to generate a filtered list of hypomethylated DMRs, by setting filtering criteria to include regions with diff.Methy statistics < -0.2. The filtering threshold can be adjusted, more filtering criteria can be defined, based on your research questions.

### Annotate DMRs

Next, we are going to annotate the filtered DMRs list with genes information using an annotation model.

* Click **Filtered feature list** node.
* Click **Region analysis** section in the context-sensitive task menu.
* Click **Annotate regions**.
* Assembly for this demo dataset should be **Homo sapiens (human) - hg38**, choose **GENCODE Genes - release 44** as Annotation model, keep the remaining settings as default, click **Finish**.

<figure><img src="/files/ePKH0PCcCZYz9tMxIIZO" alt=""><figcaption></figcaption></figure>

When completed, double click **Annotated regions** node to open the annotation report. The annotation report shows a pie chart on gene section breakdown for the DMRs, and a table where each row is a DMR, columns are the annotated gene information.

* Click **Optional columns** on the top right of the table, tick **gene\_name** checkbox to display gene name in the table.

<figure><img src="/files/f9ST6GtY7Oec7pi9GxFP" alt=""><figcaption></figcaption></figure>

### Gene Set Enrichment <a href="#gene-set-enrichment" id="gene-set-enrichment"></a>

Gene set enrichment analysis identifies gene sets and pathways that are over-represented in a list of significant genes, providing clues to the biological meaning of your results.

* Click **Annotated regions** node.
* Click **Biological interpretation** section in the context-sensitive task menu.
* Click **Gene set enrichment**.
* Select **KEGG database** as **Database** for pathway enrichment analysis. Choose **Homo sapiens hsa\_v12\_25\_04\_07** from the **KEGG database** dropdown.
* At Feature identifier section, tick **Select feature identifier** checkbox, select **gene\_name**, then click **Finish**.

<figure><img src="/files/xYAydHw2DENBw7LAQkh2" alt=""><figcaption></figcaption></figure>

When completed, double click **Pathway enrichment** node to open the pathway enrichment report. Each row in the report is a pathway, with an enrichment score and p-value. It also lists how many genes in the pathway were in the input gene list and how many were not. Click on the pathway ID in the first column to view the pathway diagram. On the pathway diagram, click on a gene name links to KEGG page for additional details.

<figure><img src="/files/tGiNJii9qXGJx1tKxZZl" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/RYBgSxJtEm0bhDF7OClT" alt=""><figcaption></figcaption></figure>


# Introduction

End to End (E2E) [Illumina Protein Prep](https://www.illumina.com/products/by-type/sequencing-kits/library-prep-kits/protein-prep.html) workflow combines Illumina chemistry, SOMAmer technology, and DRAGEN data analysis for a comprehensive, automated NGS-based proteomics solution. This E2E solution provides the following:

* NGS readout of more than 9.5K unique protein targets for a single plasma or serum sample.
* From sample to processed results in under 2.5 days with just 4 hours of hands-on time.
* Integrated analysis with BaseSpace Sequence Hub or [BioInsight Platform Core](https://help.ica.illumina.com/) (formerly ICA) and [Illumina Connected Multiomics](https://help.multiomics.illumina.com/icm) (powered by [Partek](https://help.partek.illumina.com/)).
  * Includes both local and cloud solutions for planning a run and for processing data with DRAGEN Protein Quantification.

## End-to-End Overview

The E2E SomaSeq solution integrates automation steps to perform sample preparation, protein capture, sequencing, and bioinformatics analysis. Once sequencing is finished, the DRAGEN Protein Quantification application automatically initiates on the BioInsight Platform Core. The diagram below illustrates the E2E workflow.

<figure><img src="/files/sAFfw7pR2ZibDVJSSJMV" alt=""><figcaption><p>E2E Illumina Protein Prep (IPP) workflow.</p></figcaption></figure>

For documentation on the assay and automation components of Illumina Protein Prep, please refer to the [Illumina Protein Prep Product Documentation](https://support-docs.illumina.com/LP/IlluminaProteinPrep/Content/LP/IlluminaProteinPrep/Workflow.htm).

## Versioning

Unless otherwise specified, this documentation covers DRAGEN Protein Quantification v3.0.0.


# Prerequisites

Before setting up and running the Illumina Protein Prep End-to-End (E2E) solution, ensure that the necessary software, tools, and configurations are in place. These prerequisites can vary depending on the environment used to run the secondary analysis (e.g., via cloud or locally). Follow the steps below to configure instrument and software appropriately.

### Instrument Software Prerequisites

{% tabs %}
{% tab title="NovaSeq 6000" %}

* Control Software (v1.8.0 or later)
* Illumina Protein Prep custom recipe XML file installed on the sequencing instrument. Illumina provides the file.
  * Illumina Protein Prep NovaSeq 6000 v1.0.xml
* Illumina Protein Prep Automation System output files, depending on the sequencing set up.
  * NovaSeq 6000 (S4 flow cell)
    * Recommended: Two IPPAS output files / 192 reactions
      {% endtab %}

{% tab title="NovaSeq X Series" %}

* Control Software (v1.3.0 or later)
* Illumina Protein Prep custom recipe XML file installed on the sequencing instrument. Illumina provides the file.
  * Illumina Protein Prep NovaSeq X 10B v2.0.xml
  * Illumina Protein Prep NovaSeq X 25B 100 cycle v1.0.xml
  * Illumina Protein Prep NovaSeq X 25B 200 cycle v1.0.xml (upon request only)
  * Illumina Protein Prep NovaSeq X 25B 300 cycle v1.0.xml (upon request only)
* Illumina Protein Prep Automation System output files, depending on the sequencing set up.
  * NovaSeq X (10B flow cell)
    * Recommended: Two IPPAS output files / 192 reactions
  * NovaSeq X (25B flow cell)
    * Recommended: Four IPPAS output files / 384 reactions
      {% endtab %}
      {% endtabs %}

### Cloud Secondary Analysis Prerequisites

* BaseSpace Sequence Hub (BSSH) or BioInsight Platform Core (formerly ICA) account.
  * All BSSH and Platform Core subscriptions come with access to the DRAGEN Protein Quantification application.
  * \[Platform Core subscribers] To view results in Platform Core, select Platform Core Run Storage in BSSH Workgroup Settings.
  * For more information about which subscription is best for your use case, contact your FAS or TAM.
* For information on registering a BaseSpace Sequence Hub or Platform Core account, refer to the [account management documentation](https://help.connected.illumina.com/account-management/rg-registration) or connect with an Illumina Account Manager.

<mark style="color:orange;">**Note**</mark>: When performing analysis with more than one plate, each plate must have their own unique plate barcode value.

### Local Secondary Analysis Prerequisites

* Root privileges for install
* Willingness to install Docker (included in DRAGEN Application Manager installation)
* DRAGEN Phase 4 server with the minimum disk space:
  * NovaSeq6000, S4 flow cell: ≥ 10GB
  * NovaSeqX, 10B flow cell: ≥ 20 GB
  * NovaSeqX, 25B flow cell: ≥ 35 GB
* DRAGEN Phase 4 server with the minimum available RAM:
  * NovaSeq6000, S4 flow cell: ≥ 75GB
  * NovaSeqX, 10B flow cell: ≥ 80GB
  * NovaSeqX, 25B flow cell: ≥ 95GB
* External storage drive mounted to a DRAGEN server. This drive must be mounted through a network share and support NFS/CIFS/SMB protocols. Read and write permissions are required to use this network share. For more information, please refer to the [Illumina Protein Prep Product Documentation](https://support-docs.illumina.com/LP/IlluminaProteinPrep/Content/LP/IlluminaProteinPrep/Workflow.htm).


# Illumina Protein Prep Automation System Output Files

The Illumina Protein Prep Automation System Output File (IPPAS output file) is a .csv file that's produced after automated library prep is completed. It contains the following fields:

<table><thead><tr><th width="264">Column Header</th><th>Description</th></tr></thead><tbody><tr><td>Sample ID</td><td>Identifier per sample. Specified in Sample Manifest prior to automation.</td></tr><tr><td>Well position</td><td>Well position of the sample. Specified in Sample Manifest prior to automation.</td></tr><tr><td>Project</td><td>Optionally contains information about what project output a file should contain. See <a href="/pages/jSHi4FUYKmXei3MVYQD4">Lane Splitting and Multi-Analysis</a> by Project for more details.</td></tr><tr><td>PlateBarcode</td><td>Barcode associated with plate. Specified in Sample Manifest prior to automation.</td></tr><tr><td>BatchID</td><td>User-provided batch identifier</td></tr><tr><td>InputType</td><td>Sample input type, either plasma, serum, plamsa_calibrator, serum_calibrator, plasma_QC, serum_QC, or blank. Specified in Sample Manifest prior to automation.</td></tr><tr><td>MatrixTubeBarcode</td><td>Barcode associated with matrix tube. Specified in Sample Manifest prior to automation.</td></tr><tr><td>ControlID</td><td>Lot number for associated calibrator, QC, or blank sample.</td></tr><tr><td>ProbePlate</td><td>Proteomics Probe Plate barcode scanned during Proteomics Assay.</td></tr><tr><td>SOMAmerBeadPlate</td><td>SOMAmer-Bead Plate Dil 1 barcode.</td></tr><tr><td>Lanes</td><td>Optionally contains information about lane splitting. See <a href="/pages/jSHi4FUYKmXei3MVYQD4">Lane Splitting and Multi-Analysis</a> by Project for more details.</td></tr></tbody></table>


# Local Software Install

## <mark style="color:blue;">Moving application install file to the server:</mark>

### Via USB:

1. Run the following command:
   1. `cd /`
2. Run the following command to identify which USB ports are in use:
   1. `lsblk -I 8 –d`
3. Record the ports that are currently in use.
4. Run the following command to create the USB drive mount directory on the DRAGEN server:
   1. `mkdir /media/usb`
5. Connect the USB drive with the DRAGEN Protein Quantification Pipeline installer to the front of the DRAGEN server.
6. Run the following command to confirm the USB drive name and details:
   1. `lsblk`
7. INFO: The details include the name of the USB drive under the Name column (sda, sdb, sdc, or sdd). The partition name also displays under the drive name (for example, sdc1).
8. Compare the USB ports that display to the ports identified in step 4 and identify the new port that appears. This port is where the installation software is located.
9. Run the following command to mount the USB drive to the USB mount directory of the DRAGEN server:
   1. `mount /dev/<port> /media/usb/`
   2. For example: `mount /dev/sdc1 /media/usb`
10. Run the following command to find the SHA value in the installer file:
    1. `head -n25 /media/usb/install_DRAGEN_Protein_Quantification_v<version>.run | grep '^SHA'`
11. Review the following table for SHA values.

    <table><thead><tr><th width="121.111083984375">SW Version</th><th>SHA</th></tr></thead><tbody><tr><td>2.2.2</td><td>d1000ea10b7eb5483db776efbe4873e80533dd0e5f5d713339a89cf5e7ee1a2b</td></tr><tr><td>2.3.0</td><td>f2326f2a0df50f86af1aa1b2fbe514cefa67c8879e49891385a7e780b358264f</td></tr></tbody></table>

    1. <mark style="color:$danger;">WARNING</mark>: If the SHA values do not match, stop the installation and contact Illumina Technical Support.
12. Run the following command to make sure that the USB drive is mounted to the USB mount directory:
    1. `lsblk -I 8 –d`
13. Run the following command to confirm that install\_DRAGEN\_Protein\_Quantification\_v\<version>.run is in the USB drive mount directory:
    1. `ls /media/usb/`
14. Run the following command to copy install\_DRAGEN\_Protein\_Quantification\_v \<version>.run to the staging directory:
    1. `cp /media/usb/install_DRAGEN_Protein_Quantification_v<version>.run /staging/`
15. Unmount USB from mount directory:
    1. `umount /dev/<usb partition name>`

### Via Cloud Download

1. Navigate to staging:
   1. `cd /staging/`
2. Download the installer from its online location
   1. `curl <link>`

### Via Connected Server

1. Copy the installer from an attached server:
   1. `cp <external location of installer>/install_DRAGEN_Protein_Quantification_v <version>.run /staging/`

## <mark style="color:blue;">Installation of application:</mark>

1. Run the following command to change directories to staging:
   1. `cd /staging/`
2. If necessary, change the permissions on the .run file to that it is executable with the following command
   1. `chmod +x install_DRAGEN_Protein_Quantification_v<version>.run`
3. Run the following command to make a temporary directory and install the .run file:
   1. `sudo ./install_DRAGEN_Protein_Quantification_v<version>.run --target /staging/`
4. If DAM is not installed, run the following command:
   1. `sudo ./dragen-app-manager-<version> --target /staging/`
   2. `sudo ./install_DRAGEN_Protein_Quantification_v<version>.run --target /staging/`
5. Run the following command to show the contents of the /usr/local/bin/ directory:
   1. `ls /usr/local/bin/`
6. Make sure that the /usr/local/bin/ directory has the following scripts:
   1. check\_DRAGEN\_Protein\_Quantification\_v\<version>.sh
   2. uninstall\_DRAGEN\_Protein\_Quantification\_v\<version>.sh
   3. run\_DRAGEN\_Protein\_Quantification\_v\<version>.sh
7. While in /staging directory, run the following command to confirm that the installation is successful:
   1. check\_DRAGEN\_Protein\_Quantification\_v\<version>.sh
8. Run the following help command to display the command options:
   1. run\_DRAGEN\_Protein\_Quantification\_v\<version>.sh –h
9. Perform run mock test using the following command:
   1. `run_DRAGEN_Protein_Quantification_v<version>.sh --mock -r /staging/dragen-app-manager/applications/Illumina_DRAGEN_Protein_Quantification_<version>/resources/test-files/mock_run_folder/`
10. Confirm that the output files are present in the folder that was created during the mock test.
11. If an error occurs during installation, uninstall and reinstall DRAGEN Protein Quantification Pipeline as follows.
    1. Run the uninstall\_DRAGEN\_Protein\_Quantification\_\<version>.sh script.
    2. Make sure that the following message displays:
       1. Successfully uninstalled DRAGEN\_Protein\_Quantification scripts, test-data, and images.
    3. Confirm that install\_DRAGEN\_Protein\_Quantification\_v\<version>.run is present in staging:
       1. `ls /staging/`
    4. Repeat the steps in this section to reinstall DRAGEN Protein Quantification Pipeline.
    5. If the reinstallation fails, redo steps to put the application installer on the server again.
       1. Unmount and reinsert USB, proceeding with the Via USB subsection from the Move application installer to server section (unmount using step 17 from the above section of the guide).
    6. Repeat steps in this section to reinstall DRAGEN Protein Quantification Pipeline.


# Run Planning with the BSSH Run Planner Tool

To plan a successful sequencing run, a [sample sheet](https://help.connected.illumina.com/run-set-up/overview#what-is-a-sample-sheet) with details on run configuration (e.g., sequencer type, flowcell, and sample type) is required. Follow the instrument-specific steps below to create a sample sheet compatible with the Illumina Prep Kit and DRAGEN Protein Quantification.

{% tabs %}
{% tab title="BSSH Run Planner with NovaSeq 6000" %}

1. Log in to [BaseSpace Sequence Hub](https://login.illumina.com/) and select your workgroup.
2. In the Run Planning tool, configure the settings described in the following table. Some settings are instrument specific. When you select the DRAGEN Protein Quantification application, the library prep kit and index adapter kit populate automatically, along with additional instrument-specific settings.

This table describes the possible configuration settings and values.

<table><thead><tr><th width="201">Setting</th><th>NovaSeq 6000 Value</th></tr></thead><tbody><tr><td>Instrument Platform</td><td>NovaSeq 6000/6000Dx</td></tr><tr><td>Secondary Analysis</td><td><p>[Cloud analysis] BaseSpace/BioInsight Platform Core</p><p>[Local analysis] DRAGEN Server</p></td></tr><tr><td>Application</td><td>DRAGEN Protein Quantification (select the latest version)</td></tr><tr><td>Library Prep Kit</td><td>Illumina Protein Prep 9.5k (auto-populated)</td></tr><tr><td>Index Adapter Kit</td><td>Illumina DNA/RNA UD Indexes v3 (auto-populated)</td></tr><tr><td>[NovaSeq 6000]</td><td>These settings are configured automatically and are not editable.<br>- 2 indexes<br>- Single Read<br>- 15, 10, 10, 0</td></tr></tbody></table>

3. Lane Splitting is not supported for NovaSeq 6000 runs. Select "Repeat set of samples across all lanes."
4. Upload Illumina Protein Prep Automation System output file (\*.csv) to BSSH Run Planner.

   * Select Import samples, select the CSV file type, and upload the Illumina Protein Prep Automation System output file. The interface highlights invalid values immediately after the file is rendered.
   * \[Optional] To include additional plates in the run, repeat the import process and select Add to existing samples when prompted. The new samples are appended to the samples that were uploaded previously.
     * Two output files can be uploaded for a 192-sample sequencing run.
     * Four output files can be uploaded for a 384-sample sequencing run (25B).
     * Six output files can be uploaded for a 576-sample sequencing run (25B).

   <mark style="color:orange;">**WARNING**</mark> - Sample IDs and index sequences must be unique within a sequencing run. If combining libraries from multiple Illumina Protein Prep runs, avoid combining plates that contain the same sample IDs or index sequences.
5. \[Optional] Multi-project analysis: Users may add `Project` values in the Illumina Protein Prep output file (prior to uploading) or add values in the BSSH Run Planner. Multiple project values are permitted across samples within a plate. See [Lane Splitting and Multi-Analysis by Project](/dragen-protein-quantification/run-setup/lane-splitting-and-project-splitting) for more details.
6. \[Optional] Enter an appropriate output file prefix. This value is used as a part of the prefix for the secondary analysis output file names. The first character must be alphanumeric. For the remaining characters, alphanumeric, hyphens, underscores, and spaces are permitted.
7. Proceed to the Run Review page and save the planned run.
   * \[NovaSeq 6000, cloud analysis] Download the sample sheet and save it to a network location accessible to the sequencing instrument.
   * \[NovaSeq 6000 Local analysis] Select Export to download the sample sheet. Save the file to a network location accessible to the sequencing instrument.

For additional information on run planning, refer to the [BSSH Plan Runs](https://help.connected.illumina.com/basespace-sequence-hub/sequence/plan-runs) documentation.
{% endtab %}

{% tab title="BSSH Run Planner with NovaSeq X" %}

1. Log in to [BaseSpace Sequence Hub](https://login.illumina.com/) and select your workgroup.
2. In the Run Planning tool, configure the settings described in the following table. Some settings are instrument specific. When you select the DRAGEN Protein Quantification application, the library prep kit and index adapter kit populate automatically, along with additional instrument-specific settings.

This table describes the possible configuration settings and values.

<table><thead><tr><th width="201">Setting</th><th>NovaSeq X Value</th></tr></thead><tbody><tr><td>Instrument Platform</td><td>NovaSeq X Series</td></tr><tr><td>Secondary Analysis</td><td><p>[Cloud analysis] BaseSpace/BioInsight Platform Core</p><p>[Local analysis] DRAGEN Server</p></td></tr><tr><td><p>[Local analysis]</p><p>FASTQ file compression format</p></td><td><p>DRAGEN</p><p>This setting is required by default. The setting does not impact DRAGEN Protein Quantification as no FASTQ files are output.</p></td></tr><tr><td><p>[Local analysis]</p><p>Generate FastQC metrics</p></td><td>Yes<br>This setting is optional. The setting does not impact DRAGEN Protein Quantification as no FASTQ files are output.</td></tr><tr><td>Read Lengths</td><td>- Read 1: 15<br>- Index 1: 10<br>- Index 2: 10<br>- Read 2: 0</td></tr><tr><td>Application</td><td>DRAGEN Protein Quantification (select the latest version)</td></tr><tr><td>Library Prep Kit</td><td>Illumina Protein Prep 9.5k (auto-populated)</td></tr><tr><td>Index Adapter Kit</td><td>Illumina DNA/RNA UD Indexes v3 (auto-populated)</td></tr><tr><td>Override Cycles</td><td>These settings are configured automatically and should not be edited.<br>- Read 1: Y15<br>- Index 2: I10<br>- Index 2: I10</td></tr></tbody></table>

3. If lane splitting will be utilized with this sequencing run, it's recommended to add values to `Lane` column to Illumina Protein Prep Automation System output files locally. See [Lane Splitting and Multi-Analysis by Project](/dragen-protein-quantification/run-setup/lane-splitting-and-project-splitting) for more details.
   1. If lane splitting is not utilized, do not edit the `Lane` column in the Illumina Protein Prep Automation System output file.
4. Upload Illumina Protein Prep Automation System output file (\*.csv) to BSSH Run Planner.

   * Select Import samples, select the CSV file type, and upload the Illumina Protein Prep Automation System output file. The interface highlights invalid values immediately after the file is rendered.
   * \[Optional] To include additional plates in the run, repeat the import process and select "Add to existing samples" when prompted. The new samples are appended to the samples that were uploaded previously.
     * Two output files can be uploaded for a 192-sample sequencing run.
     * Four output files can be uploaded for a 384-sample sequencing run (25B).
     * Six output files can be uploaded for a 576-sample sequencing run (25B).
   * Barcode mismatch 1 and 2—No action is required. The default value is set to 1. Do not change this value.
   * If lane splitting will not be utilized, indicate that each sample is present in all lanes. For the first sample, click on the Lanes box and select the first checkbox. This will populate the cell with "1,2,3,4,5,6,7,8". Then, select the Lanes header and click "Fill down". This will add these lane values for all samples.

   ![](/files/ymxPnHJ5bXyxZiljqLNK) ![](/files/mFl0nK026iO3JTw9Laa9)

   <mark style="color:orange;">**WARNING**</mark> - Sample IDs and index sequences must be unique within a sequencing run. If combining libraries from multiple Illumina Protein Prep runs, avoid combining plates that contain the same sample IDs or index sequences. When combining libraries with non-unique indexes, ensure they are loaded into different flow cell lanes, and that lane splitting is enabled during sample sheet creation.
5. \[Optional] Multi-project analysis: Users may add `Project` values in the Illumina Protein Prep output file (prior to uploading) or add values in the BaseSpace Run Planner. Multiple project values are permitted across samples within a plate. See [Lane Splitting and Multi-Analysis by Project](/dragen-protein-quantification/run-setup/lane-splitting-and-project-splitting) for more details.
6. \[Optional] Enter an appropriate output file prefix. This value is used as a part of the prefix for the secondary analysis output file names. The first character must be alphanumeric. For the remaining characters, alphanumeric, hyphens, underscores, and spaces are permitted.
7. Proceed to the Run Review page and save the planned run.
   * \[NovaSeq X Series, cloud analysis] No action is required. The sample sheet is automatically uploaded to the instrument.
   * \[Local analysis] Select Export to download the sample sheet. Save the file to a network location accessible to the sequencing instrument.

For additional information on run planning, refer to the [BSSH Plan Runs](https://help.connected.illumina.com/basespace-sequence-hub/sequence/plan-runs) documentation.

**Additional Notes:**

* DRAGEN Protein Quantification does not support Multiple Analysis on NovaSeq X.
  {% endtab %}
  {% endtabs %}


# Sample Sheet Fields

A sample sheet is required to kick off secondary analysis. It can be made either using the BSSH Run Planner Tool (recommended), Excel Sample Sheet Generator, or manually. The following table describes the sample sheet fields and its values depending on the environment used to execute the DRAGEN Protein Quantification application. The pipeline can be executed via cloud using either [BioInsight Platform Core](https://help.connected.illumina.com/illumina-connected-analytics) or [BaseSpace Sequence Hub](https://help.connected.illumina.com/basespace-sequence-hub), or executed locally using a phase 4 [DRAGEN Server](https://help.connected.illumina.com/dragen).

## Sample Sheet Fields

This is a non comprehensive list of fields.

<table><thead><tr><th width="260">Section</th><th>Field</th><th>Value</th></tr></thead><tbody><tr><td>Header</td><td>FileFormatVersion</td><td>Must be "2"</td></tr><tr><td>Header</td><td>InstrumentPlatform</td><td>Must be "NovaSeqXSeries" or "NovaSeq"</td></tr><tr><td>Header</td><td>RunName</td><td>User-provided value</td></tr><tr><td>Header</td><td>RunDescription</td><td>User-provided value (optional)</td></tr><tr><td>Reads</td><td>Read1Cycle</td><td>15</td></tr><tr><td>Reads</td><td>Index1Cycle</td><td>10</td></tr><tr><td>Reads</td><td>Index2Cycle</td><td>10</td></tr><tr><td>Sequencing_Settings</td><td>LibraryPrepKits</td><td>Must be "IlluminaProteinPrep9.5k"</td></tr><tr><td>BCLConvert_Data</td><td>Sample_ID</td><td>Alphanumeric name up to 100 characters. Letters, numbers, dashes only (or any combination of letters, numbers, and dashes)</td></tr><tr><td>BCLConvert_Data</td><td>Index</td><td>i7 index sequence, including A, C, T or G letters, 10 nucleotides long</td></tr><tr><td>BCLConvert_Data</td><td>Index2</td><td>i5 index sequence, including A, C, T or G letters, 10 nucleotides long</td></tr><tr><td>BCLConvert_Data</td><td>Lane</td><td>Lane value (shall be a number between 1 and 8 inclusive) (optional if instrument is NovaSeq 6k, required if instrument is NovaSeq X)</td></tr><tr><td>Cloud_Proteomics_Settings (Cloud Analysis) or Proteomics_Settings (Local Analysis)</td><td>SoftwareVersion</td><td>Three-digit version of SW used in secondary analysis. For example, "2.3.0"</td></tr><tr><td>Cloud_Proteomics_Settings (Cloud Analysis) or Proteomics_Settings (Local Analysis)</td><td>StartsFromFastq</td><td>Must be "false"</td></tr><tr><td>Cloud_Proteomics_Settings (Cloud Analysis) or Proteomics_Settings (Local Analysis)</td><td>output_file_prefix</td><td>User-provided prefix</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>Sample_ID</td><td>Alphanumeric name up to 100 characters. Letters, numbers, dashes only (or any combination of letters, numbers, and dashes)</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>PlateBarcode</td><td>For each plate, associated plate barcode</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>MatrixTubeBarcode</td><td>For each plate, associated matrix tube barcode (optional)</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>BatchID</td><td>For each plate, user-provided batchID</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>InputType</td><td>For each sample, associated input type. Permitted values: Blank, Plasma, Serum, CSF, or MultiMatrix-&#x3C;tissuetype></td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>Control</td><td>For each sample, the control type. Permitted values: Blank, Calibrator, QC, or empty (for non-control samples)</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>ControlID</td><td>ID of the calibrator, QC, and blank lot (applied to controls only)</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>KitType</td><td>The kit type for the plate. Permitted values: Plasma, Serum, CSF, or MultiMatrix</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>ProbePlate</td><td>For each plate, probe plate barcode from library prep</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>SOMAmerBeadPlate</td><td>For each plate, SOMAmer Bead Plate barcode from library prep. SBP1 format for Plasma/Serum/CSF, MBP format for MultiMatrix.</td></tr><tr><td>Cloud_Proteomics_Data (Cloud Analysis) or Proteomics_Data (Local Analysis)</td><td>WellPosition</td><td>For each sample, well position with UDI set prefix (e.g., A-A01 through D-H12)</td></tr><tr><td>Cloud_Settings</td><td>GeneratedVersion</td><td>Software version of the BSSH Run Planner tool that generated the sample sheet</td></tr><tr><td>Cloud_Settings</td><td>Cloud_Proteomics_Pipeline (Cloud Only)</td><td><p>Platform Core Path to the proteomics pipeline.<br>For example:</p><pre class="language-shell"><code class="lang-shell">urn:ilmn:ica:pipeline:896aa1df-0412-45db-814e-4683b3215203#DRAGEN_Protein_Quantification_2-2-2
</code></pre></td></tr><tr><td>Cloud_Data</td><td>Sample_ID</td><td>Alphanumeric name up to 100 characters. Letters, numbers, dashes only (or any combination of letters, numbers, and dashes)</td></tr><tr><td>Cloud_Data</td><td>ProjectName</td><td>(optional) user-provided project. Multiple values are allowed across samples.</td></tr><tr><td>Cloud_Data</td><td>LibraryName</td><td>For each sample, must be &#x3C;Sample_ID>_&#x3C;index><em>_</em>&#x3C;index2></td></tr><tr><td>Cloud_Data</td><td>LibraryPrepKitName</td><td>Must be "IlluminaProteinPrep9.5k"</td></tr><tr><td>Cloud_Data</td><td>IndexAdapterKitName</td><td>Must be "IlluminaDNARNAUDISetABCDTagmentation_Proteomics"</td></tr></tbody></table>

## Samplesheet Examples

Examples of local and cloud sample sheets for NovaSeq 6000 and NovaSeq X are attached to this page.

### Local (Plasma/Serum)

### Cloud (Plasma/Serum)

### CSF

### MultiMatrix

For additional information, refer to the Illumina BioInsight Platform [Sample Sheet documentation](https://help.connected.illumina.com/run-set-up/overview).


# Local Sample Sheet Generation Tool

The Run Planning web interface, accessible via [BaseSpace Sequence Hub (BSSH)](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub.html), requires internet connection. The Proteomics Sample Sheet Generator tool (linked below) allows for offline creation of sample sheets compatible with the local DRAGEN Protein Quantification application.

## Prerequisites

1. Local Proteomics Sample Sheet Generator Excel Workbook.
2. One or more IPPAS output files (one per plate).
3. Access to a DRAGEN Server configured to the Illumina Protein Quantification Local Secondary Analysis (see prerequisites page for more information).

## Steps

1. Download and open the Local Proteomics Sample Sheet Generator excel file.
2. Navigate to the Start tab and follow the instructions described in the upper left portion of the sheet:
   1. Fill in the `RunName` and output\_file\_prefix.
   2. Change the `InstrumentPlatform` as needed.
   3. Copy and paste values from IPPAS output file(s) not including the headers into the specified fields.
   4. Add the Flow Cell lanes for each sample as a comma separated list with no spaces as needed.
   5. Ensure there are no errors displayed in the Row Check or Overall Check (a green cell is expected)
   6. Navigate to the "SaveAsCSV" and save it as CSV Comma-Delimited (filename must not contain special characters).
   7. Once the sample sheet is generated, update `LibraryPrepKits` under section `[Sequencing_Settings]` to "IlluminaProteinPrepKit9.5k". This change is required for compatibility with the latest local DRAGEN Protein Quantification.
3. (Recommended) Load this sample sheet onto the sequencer prior to sequencing.
   1. If the sample sheet is not included prior to sequencing, the user must manually reference the sample sheet when running DRAGEN Protein Quantification later.
4. Analyze the data locally using a DRAGEN Server with DRAGEN Protein Quantification Pipeline (see [Local Secondary Analysis page](/dragen-protein-quantification/counting-and-normalization/local-secondary-analysis) for more information).

### Proteomics Sample Sheet Generator

{% file src="/files/wwfyjWtULGh0qFXcvCP0" %}


# Lane Splitting and Multi-Analysis by Project

## Lane Splitting

The purpose of lane splitting is to enable the reuse of sample indexes (barcodes) on the same flow cell. To accomplish this, the samples indexed with the same barcodes must be physically separated by placing them on different lanes of the flow cell.

Currently, lane splitting functionality in Illumina Protein Prep is only supported on the NovaSeqX platform. To use lane splitting, use the sample sheet to indicate which lane(s) each sample is found in.\
Recommendation: Edit the "Lanes" column of the Illumina Protein Prep Automation System output file(s):

1. Find the `Lanes` column in the Illumina Protein Prep Automation System output file.
2. Add lane numbers in which the sample is located as a comma-separated value.\
   **Example:** '1,2,3,4' for a sample located in lanes 1, 2, 3, and 4.
3. Upload the modified Illumina Protein Prep Automation System output file to the [BSSH Run Planner](/dragen-protein-quantification/run-setup/page).

## Multi-Analysis by Project

The goal of multi-analysis by project is to improve flexibility by enabling multiple analysis outputs from a single sequencing run without needing to requeue the sample sheet. A single project must include at least one plate (with its controls) and can include multiple plates used in a sequencing run. Each project produces one set of output files (including normalized ADATs and a DRAGEN Report).

\
To use multi-analysis by project, include the project name in the `Project` column of the Illumina Protein Prep Automation System output files. If no project information is provided, DRAGEN Protein Quantification will assume all plates belong to the same project and will assign the project name based on the user provided 'run name'. The multi-analysis by project feature is available for both the NovaSeq 6000 and NovaSeq X platforms.

1. Find the `Project` column in the Illumina Protein Prep Automation System output file.
2. Edit `Project` column in an Illumina Protein Prep Automation System output file by assigning a project name to each sample. Project name string should contain only alphanumeric characters and underscores.
3. Upload modified Illumina Protein Prep Automation System output file to the Run Planner.

<mark style="color:orange;">**Note**</mark>: If you want all samples from a sequencing run to be analyzed together, DO NOT include any values in the "Project" column of the sample sheet.

<mark style="color:orange;">**Note**</mark>: If project splitting is enabled, all samples on one plate must be included in the same project. Samples from one plate CANNOT be split across projects.

## Lane Splitting and Multi-Analysis Examples

Here is an example of what the first lines of two individual IPPAS output files may look like after project and lanes values were added. The plate with Barcode "A" adds lanes "1,2,3,4" and project value "ProjectA". The plate with barcode "B" adds lanes "5,6,7,8" and project value "ProjectB".

<figure><img src="/files/iJmwwfITkBW8zKWBRKPr" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/zy52DVuyfPOF3oZobjjf" alt=""><figcaption></figcaption></figure>


# Cloud Autolaunch Secondary Analysis

## DRAGEN Protein Quantification Application

The DRAGEN Protein Quantification application is designed to perform counting and normalization for proteomics data from the Illumina Protein Prep pipeline. It converts data from the binary base call (BCL) files, generated by Illumina NovaSeq 6000 or NovaSeq X Series systems, into the normalized proteomic counts. Upon completion of sequencing, the application is automatically initiated for analysis on BaseSpace Sequence Hub (BSSH) or BioInsight Platform Core.

The sections below exemplify how to configure the instrument(s) for autolaunching the secondary analysis, manually initiate an analysis, and requeue an analysis.

## Setting Up Autolaunch on Sequencing Instrument

1. On your instrument, log into your workgroup and select the following run setup options:
   1. Workflow: NovaSeq Standard
   2. Read length: 15, 10, 10, 0
2. \[NovaSeq 6000] Upload the v2 sample sheet generated by the BSSH Run Planner tool to the sequencing instrument.
3. \[NovaSeq X Series] No action is required. The v2 sample sheet automatically displays on the instrument as a planned run.
4. Select the appropriate Illumina Protein Prep custom recipe for your sequencing instrument and flowcell. During the run, data is uploaded automatically to BSSH. Primary analysis and secondary analysis are completed automatically through autolaunch.
5. \[Optional] Use the BaseSpace Sequence Hub to monitor the run from start to completion.

For additional information on autolaunch, refer to the [Cloud Analysis Auto-launch](https://help.connected.illumina.com/analysis/analysis_autolaunch) page.

## Manually Kicking Off Secondary Analysis

If manual mode sequencing is performed, or autolaunch is unsuccessful, there is an option to manually upload a completed run folder to BSSH and kick off the autolaunch analysis. To use this method, the samplesheet must be in the uploaded run folder, and be named "SampleSheet.csv".

Follow the instructions from the [BSSH CLI documentation](https://developer.basespace.illumina.com/docs/content/documentation/cli/cli-overview) for [manually uploading the run folder](https://developer.basespace.illumina.com/docs/content/documentation/cli/cli-examples#Uploadingarunanamerun-uploada) to BSSH.

## Requeue

Follow the instructions described in [Requeue Analysis Page (Requeue Analysis in BaseSpace tab)](https://help.connected.illumina.com/analysis/analysis_autolaunch#requeue-analysis).


# Accessing Cloud Results

## Primary Metrics

To view primary metrics:

1. Go to the relevant BaseSpace Sequence Hub (BSSH) workgroup.
2. Navigate to the "Runs" tab and select the run.
3. The "Summary" tab gives an overview of the sequencing run quality (e.g., average %Q30, %PF Yield).
4. Navigate to "Metrics" for detailed per lane information on all sequencing metrics.

## Secondary Results

To view the analysis associated with a specific sequencing run:

1. Navigate to the "Summary" tab of the relevant run.
2. Click on the link below "Latest Analysis" (displaying the results from the latest analysis processed to the run data).\
   For re-queued/re-analyzed runs, the previously completed analyses can be found under the "Prior Analysis" section
3. Click on "Reports" and find the quality metrics associated with the secondary analysis on the run data.

<mark style="color:blue;">**Note**</mark>: Those who have access to a BioInsight Platform Core account and prefer to view results there may either click on "View Files in Platform Core" in the top right corner of your BSSH Analysis page or directly access the analysis in Platform Core. The secondary analysis results in Platform Core will be in a BSSH-managed project with the same name as the BSSH workgroup where the analysis was performed.

For further information on tracking and viewing run and analysis results in BaseSpace Sequence Hub, refer to the BSSH [View Data ](https://help.connected.illumina.com/basespace-sequence-hub/data/view-data)documentation.


# Local Secondary Analysis

DRAGEN Protein Quantification can be run locally after the installation of the local solution on a DRAGEN phase 4 server by an FAS.

To initiate a run:

```
run_DRAGEN_Protein_Quantification_v<version>.sh
    -r <full path to run folder>
    -s <full path to sample sheet>
    --analysisFolder <full path to output folder> 
```

The `--analysisFolder` parameter is optional. If no path is provided, output files will be put in a folder under the `/staging/` directory.

* <mark style="color:$warning;">WARNING</mark>: Currently, using a path off the DRAGEN server as an `--analysisFolder` (for example, to network attached storage) may cause an analysis failure. It's recommended to output the results to the DRAGEN server itself.

The `-s` parameter is optional if the sample sheet file is included in the run folder.

For details on the parameters used with the script, execute the following command.

```
run_DRAGEN_Protein_Quantification_v<version>.sh -h
```


# Counting and Normalization

## Counting

Protein counting is performed using DRAGEN BCL Convert. Sequencing produces barcoded reads for each sample that correspond to protein abundance. Barcoded reads are simultaneously demultiplexed and counted using DRAGEN BCL Convert. These barcode counts are stored as the Raw Counts ADAT.

## Normalization Summary

<figure><img src="/files/ZiSxrai0nFr6CEh1OdAg" alt=""><figcaption></figcaption></figure>

Normalization corrects for the sources of confounding variation, such as overall protein concentration differences, minor deviations in volume transfer during the assay, or efficiency of library preparation steps. It is performed sequentially and produces an individual ADAT file with counts for each of the three following steps:

* <mark style="color:blue;">**Readout Normalization:**</mark> This step uses **SOMAmer controls** to reduce technical variability introduced in the **NGS library prep**.
* <mark style="color:green;">**Plate Normalization**</mark><mark style="color:blue;">**:**</mark> This step uses positive controls (calibrators) to correct for biases between plates. An external reference provided by Illumina ensures all Illumina Protein Prep plates are comparable. This step can be broken down to five steps. See the diagram above and the section below for more information.
* <mark style="color:$warning;">**Sample Normalization**</mark><mark style="color:blue;">**:**</mark> This step uses **protein abundance** in each sample to reduce technical variability introduced during **protein quantification**, and **correct for differences in over all protein concentration**.

See [Metrics Appendix](/dragen-protein-quantification/counting-and-normalization/metrics-appendix) for a summary of all metrics.

## Normalization Steps in Detail

* <mark style="color:blue;">**Hybridization Normalization (also known as Readout Normalization):**</mark> This step corrects for biases that can occur during the hybridization and sequencing preparation stages of the assay. During hybridization, controls are spiked into each sample; during normalization, the counts for these controls are compared against an internal reference based on the non-blank controls on the plate. A scale factor is calculated for each sample, and if the scale factor is outside of the specification range (**0.4**–**2.5**), the sample will receive a **FLAG**.
* <mark style="color:blue;">**Internal Reference Median Normalization:**</mark> This step corrects for differences in the total protein abundance measurement of a sample. It is performed for each dilution group and runs separately for blank and calibrator control samples. A scale factor is calculated for each dilution group, by comparing the observed protein measurements to a reference of expected values for each protein.

  This reference is based on median protein counts across all samples of the same sample type on the same plate. If any of the scale factors for a sample are outside of the specification range (**0.4**–**2.5**), the sample will receive a **FLAG**.
* <mark style="color:blue;">**Plate Scaling:**</mark> This step corrects for possible changes in measured total protein counts between plates, when calibrator samples are present. The median of each SOMAmer measurement across the five calibrators is compared to an external calibration reference to calculate a plate scale factor.

  The first plate scaling step compares the calibrator medians to a reference derived from the sequencing instrument (NovaSeq 6000 or NovaSeq X Series) and adjusts the entire plate accordingly.

  The second step compares the scaled calibrator medians to a reference derived from the NovaSeq X Series (10B flow cell).\
  There is no specification range for plate scaling.

  * The references used in this step can be found in the SOMAmer metadata, under Ref.Bridging.\<CalibratorId>.\<Instrument>.\<Flowcell>.\<MasterMixLot#>. The reference used by cross-instrument plate scaling is Ref.Bridging.\<PlateBarcode>.NovaSeqX.10B.AA.
* <mark style="color:blue;">**Calibration:**</mark> This step corrects for batch effects that impact individual SOMAmers.

  The first calibration step (Platform Specific Calibration) compares the median of each SOMAmer measurement across the five Calibrator sample replicates to an external Calibration reference. This reference is derived from runs using the same instrument type, flow cell type, calibrator lot, and sample input type used in the run being analyzed. It then calculates a SOMAmer-specific scale factor and a plate-wide Calibrator metric (PlatformSpecificCalibrationTailPercent). PlatformSpecificCalibrationTailPercent corresponds to the percentage of SOMAmers with scale factors outside the specification range (**0.6**–**1.4**). If **15%** of SOMAmers fall outside of this specification, the plate receives a **WARNING** for the PlatformSpecificTailPercent\_PassFlag metric.

  The second calibration step (Cross Platform Calibration) compares the updated calibrator medians to a reference derived from the NovaSeq X Series (10B flow cell), using the same calibrator lot and sample input type as the run being analyzed. The scale factors used to align the median of the calibrators to the reference value are applied to all samples on the same plate.

  * The references used in this step can be found in the SOMAmer metadata, under Ref.Bridging.\<CalibratorId>.\<Instrument>.\<Flowcell>.\<MasterMixLot#>. The reference used by cross-instrument calibration is Ref.Bridging.\<CalibratorId>.NovaSeqX.10B.AA.
* <mark style="color:blue;">**External Reference Median Normalization (also known as Sample Normalization):**</mark> This step corrects for differences in the total protein signal for samples on each dilution plate. A scale factor is calculated by comparing the observed protein measurements to a reference of expected values for each protein. It is performed on plasma/serum samples and QC samples.

  If any of the scale factors for a sample are outside of the specification range (**0.4**–**2.5**), the sample will receive a **FLAG**.

  * The reference used in this step can be found in the SOMAmer metadata, under Ref.MedNormExt.Plasma or Ref.MedNormExt.Serum (dependent on the input type).

<figure><img src="/files/uIgxwC5HbDzM22lTJQi6" alt=""><figcaption><p>Outline of the protein capture assay and what parts of the process each normalization step impacts</p></figcaption></figure>


# QC Summary

## QC Checks

There are a number of quality control checks that are applied on a plate and sample level. See the metrics appendix page for a summary of all metrics.

* <mark style="color:blue;">**Minimum SOMAmer Read Counts:**</mark> **Non-blank samples** with less than **10 million** reads will receive a **FLAG** for SOMAmerReads\_PassFlag in the ADAT. These reads are counted in the raw counts step. Only human protein SOMAmers are part of this count, not controls. There is no specification for blank samples.
* <mark style="color:blue;">**Maximum SOMAmer Read Counts:**</mark>**&#x20;Blank samples** with more than **20 million normalized reads** will receive a **FLAG** for SOMAmerNormRead\_PassFlag in the ADAT. These reads are counted in the plate scale normalization step. Only human protein SOMAmers are part of this count, not controls. There is no specification for non-blank samples. A **plate** where **70% of the blanks** have a **FLAG** for this step will receive a **WARNING**.
* <mark style="color:blue;">**Reference Correlation:**</mark> This step produces a Spearman correlation coefficient describing how similar a sample is to a an external Plasma or Serum reference (see below). There is no pass flag for this step.
  * The reference used in this step can be found in the SOMAmer metadata, under Ref.MedNormExt.Plasma.QC or Ref.MedNormExt.Serum.QC (dependent on the input type).
* <mark style="color:blue;">**Empirical Hyb Temp:**</mark> This step uses 78 SOMAmer controls with a wide spectrum of melting temperature (Tm) from 28C to 72C, which represent the Tm of all probes used in the analysis. They were spiked to each sample at equal concentration; Tm controls with higher Tm have higher counts than those with lower Tm because their hybridization is more stable. The distribution of all Tm controls in a sample follows a logistic distribution; the inflection point is the EmpiricalHybTemp. This distribution is calculated using raw counts.
* <mark style="color:blue;">**QC Check:**</mark> This step compares the median of each SOMAmer measurement, across the three QC sample replicates, to an external QC reference. It then calculates a SOMAmer-specific scale factor and a QC metric (QCCheckTailPercent). QCCheckTailPercent corresponds to the percentage of SOMAmers with scale factors outside the specification range (**0.8**–**1.2**). If more than **15%** of the scale factors are outside of the specification range, the plate receives a **FAIL**.
  * The references used in this step can be found in the SOMAmer metadata, under Ref.QCCheck.Plasma or Ref.QCCheck.Serum (dependent on the input type).


# Interpretation of Results

## Sample Quality

The purpose of a flag is to highlight that a sample required a high degree of correction during the normalization process. This means a sample had sufficiently high or low signal, causing the normalization scale factors to be out of specification. The general recommendation is to exercise caution when using that sample in downstream analysis; normalization may not have been able to properly correct for the large changes in signal.

### Flag at SOMAmer Reads

A flag in a non-blank sample indicates low SOMAmer read depth. DRAGEN Protein Quantification has a minimum read depth to ensure measurement precision is achieved for each sample. Flagged samples may have a decrease in measurement precision.

### Flag at SOMAmer Normalized Reads

A flag in a blank sample may indicate increased background or contamination.

### Flag at Reference Correlation

A flag in a blank sample may indicate plasma or serum contamination. In internal studies, uncontaminated blanks generally have low RefCorr values (\~0.4) and blanks with greater than 2% plasma or serum contamination were observed to have a RefCorr values of at least 0.6.

### Flag at Empirical Hyb Temp

A flag in a sample indicates that there is an abnormal distribution of Tm controls in that well. These controls are added during the NGS portion of automation, so flags are the result of issues in this portion of the process. This can be caused by:

* An elevated hybridization temperature in the sample well caused the right shift in Tm controls distribution and EmpiricalHybTemp (>52.4)
  * This can be caused by evaporation or other factors
* The Tm control counts was so distorted by elevated hybridization temperature that their distribution does not follow a logistic distribution anymore. The EmpiricalHybTemp cannot be determined from the distribution and no value is provided
* Missing NGS reagents from that well

### Flag at Hybridization Normalization

A flag in a sample indicates that the hybridization controls in that sample had elevated or decreased signal compared to the plate plasma or serum controls. Flags here indicate significant differences between the sample and the non-blank controls. Potential reasons for this could be poor sample quality, protein loss during hybridization (e.g., due to elevated temperature) or other issues that arise during library prep. In this case, other failure modes will likely also be flagged, such as low Raw Counts, or high / low MedNormExt scale factors.

### Flag at Internal Reference Median Normalization

A flag in a calibrator indicates that the flagged calibrator replicate had a large discrepancy in signal when compared to the median of all calibrators within that plate. A single flagged calibrator will have limited impact on the plate due to use of median calibrator values in analysis. Multiple flagged calibrators could impact plate performance; reach out to Illumina Support for additional information.

### Flag at External Reference Median Normalization

A flag in a sample indicates that it had a large discrepancy in signal when compared to the external plasma or serum reference. Low signal may be caused by sample degradation or sample dilution; high signal may be caused by hemolysis. A difference in signal could also be caused by the diseased-state of the sample.

### Plate Fail

If a sample is on a plate that receives a FAIL, the sample receives a PLATE\_FAIL whether it was flagged or not at the sample level. PLATE\_FAIL indicates that the sample is likely not suitable for downstream analyses due to a plate-wide issue. See below for more information on plate failure.

## Plate Quality

QC Percent in Tails and Calibration Percent in Tails are used to determine plate quality by examining individual SOMAmers in each well. A QC Percent in Tails failure will likely require a repeat library preparation or re-sequencing for this plate; contact Illumina Support for additional information.

Review the following table for more details.

### Plate Quality Matrix

<table data-header-hidden data-full-width="true"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td><mark style="color:blue;"><strong>Calibration Percent in Tails</strong></mark></td><td><mark style="color:blue;"><strong>QC Percent in Tails</strong></mark></td><td><mark style="color:blue;"><strong>Interpretation</strong></mark></td></tr><tr><td><p>WARNING</p><p><em>The median of the calibrator replicates for that plate were sufficiently different from the expected.</em></p></td><td>FAIL<br><em>The median of the QC replicates for that plate were sufficiently different from the expected QC reference.</em></td><td><ul><li>Likely a failed run</li><li>Some samples may individually pass but should not be used for downstream analysis due to the plate failure</li><li>One or more normalization steps after Calibration were unable to rescue the run</li><li>Could be caused by a plate-wide issue, including sequencing run failure, automation failure, or reagent issue</li></ul></td></tr><tr><td><p>WARNING</p><p><em>The median of the calibrator replicates for that plate were sufficiently different from the expected.</em></p></td><td>PASS<br>The median of the <em>QC samples were sufficiently similar to expected values from QC reference.</em></td><td><ul><li>Successful Run</li><li>One or more normalization steps after the Calibration step have successfully removed the differences observed in calibrators</li></ul></td></tr><tr><td><p>PASS</p><p><em>The median of the calibrator samples were sufficiently similar to expected values from calibration reference.</em></p></td><td>FAIL<br><em>The median of the QC replicates for that plate were sufficiently different from the expected QC reference.</em></td><td><ul><li>Technically a failed run, but could be a false positive</li><li>Relatively rare combination</li><li>May indicate problem with only QC samples and not entire plate (e.g., edge effect)</li></ul></td></tr></tbody></table>

## Run Quality

There are two additional metrics on run quality that evaluates blank samples in a plate (see below). The general recommendation is that, if warnings are observed in one of these metrics, the plate should be used with caution in downstream analysis since there may have been sample contamination or elevated background. If issues with run quality are consistently seen, there may be issues with the assay set up. Contact Illumina Support for additional information or support.

### Warning for SOMAmer Normalized Reads

A warning on a plate may indicate plate-wide increased background or contamination.

### Warning at Reference Correlation

A warning on a plate may indicate plate-wide plasma or serum contamination.


# Common Failure Modes

Viscous Samples

Hemolyzed Samples

Un-Advised Tubes

Uneven distribution of reads

Poor sequencing quality/PCR failures

Temperate Failure

* likely won't impact controls


# Metrics Appendix

## Sample Level Metrics

Sample level metrics are found in the sample/row-level metadata in the ADAT.

<table><thead><tr><th width="136">Metric</th><th>ADAT Field</th><th>Applied To</th><th>Threshold</th><th>Impact to Sample</th></tr></thead><tbody><tr><td>Minimum SOMAmer Read Count</td><td>SOMAmerReads; SOMAmerReads_PassFlag</td><td>Non-Blank Samples</td><td>SOMAmerReads > 10 million</td><td>PASS/FLAG</td></tr><tr><td>Maximum SOMAmer Normalized Read Count</td><td>SOMAmerNormReads; SOMAmerNormReads_PassFlag</td><td>Blank Samples</td><td>SOMAmerNormReads &#x3C; 20 million reads</td><td>PASS/FLAG</td></tr><tr><td>Hybridization Normalization</td><td>HybNorm_1_ScaleFactor; HybNorm_PassFlag</td><td>Non-Blank Samples</td><td>0.4 ≤ HybNorm_1_ScaleFactor ≤ 2.5</td><td>PASS/FLAG</td></tr><tr><td>Internal Reference Median Normalization</td><td>MedNormInt_5e-05_ScaleFactor; MedNormInt_0.005_ScaleFactor; MedNormInt_0.2_ScaleFactor; MedNormInt_PassFlag</td><td>Individual Blank and Calibrator samples, split by dilution group</td><td>0.4 ≤ MedNormInt_5e-05_ScaleFactor ≤ 2.5<br><br>0.4 ≤ MedNormInt_0.005_ScaleFactor ≤ 2.5<br><br>0.4 ≤ MedNormInt_0.2_ScaleFactor ≤ 2.5</td><td>PASS/FLAG</td></tr><tr><td>External Reference Median Normalization</td><td>MedNormExt_5e-05_ScaleFactor; MedNormExt_0.005_ScaleFactor; MedNormExt_0.2_ScaleFactor; MedNormExt_PassFlag</td><td>Individual non-control and QC samples, split by dilution group</td><td>0.4 ≤ MedNormExt_5e-05_ScaleFactor ≤ 2.5<br><br>0.4 ≤ MedNormExt_0.005_ScaleFactor ≤ 2.5<br><br>0.4 ≤ MedNormExt_0.2_ScaleFactor ≤ 2.5<br><br></td><td>PASS/FLAG</td></tr><tr><td>TM Controls</td><td>EmpiricalHybTemp; EmpiricalHybTemp_PassFlag</td><td>All samples</td><td>EmpiricalHybTemp &#x3C; 52.4</td><td>PASS/FLAG</td></tr><tr><td>Row Check</td><td>RowCheck_PassFlag</td><td>All samples</td><td>PLATE_FAIL if the sample is on a plate that received a FAIL. Otherwise, PASS if all Pass Flags in this row are PASS and FLAG if not.</td><td>PASS/FLAG/PLATE_FAIL</td></tr></tbody></table>

## SOMAmer Level Metrics

SOMAmer level metrics are found in the SOMAmer/column-level metadata in the ADAT.

<table><thead><tr><th width="290.765625">Metric</th><th>ADAT Field</th></tr></thead><tbody><tr><td>Platform Specific Calibration</td><td>PlatformSpecificCalibrate_&#x3C;PlateBarcode>_ScaleFactor</td></tr><tr><td>Cross Platform Calibration</td><td>CrossPlatformCalibrate_&#x3C;PlateBarcode>_ScaleFactor</td></tr><tr><td>QC Check</td><td>QCCheck_&#x3C;PlateBarcode>_ScaleFactor</td></tr></tbody></table>

## Plate Level Metrics

Plate level metrics are found in the header metadata in the ADAT.

| Metric                        | ADAT Field                                                                           | Threshold                                          | Impact to Plate |
| ----------------------------- | ------------------------------------------------------------------------------------ | -------------------------------------------------- | --------------- |
| QC Percent in Tails           | QCCheckTailPercent; QCCheckTailPercent\_PassFlag                                     | QCCheckTailPercent < 15%                           | PASS/FAIL       |
| Calibration Percent in Tails  | PlatformSpecificCalibrateTailPercent; PlatformSpecificCalibrateTailPercent\_PassFlag | PlatformSpecificCalibrateTailPercent < 15%         | PASS/WARNING    |
| SOMAmer Normalized Reads      | PlateSOMAmerNormReads\_PassFlag                                                      | < 70% of blank samples flagged at SOMAmerNormReads | PASS/WARNING    |
| Platform Specific Plate Scale | PlatformSpecificPlateScale\_ScaleFactor                                              | N/A                                                | N/A             |
| Cross Platform Plate Scale    | CrossPlatformPlateScale\_ScaleFactor                                                 | N/A                                                | N/A             |
| Cross Platform Calibration    | CrossPlatformCalibrationTailPercent                                                  | N/A                                                | N/A             |


# Output Structure

DRAGEN Protein Quantification produces the following key output files in BaseSpace Sequence Hub:

* DRAGEN\_Protein\_Quantification\_\<SW Version>
  * \<project> (if no projects are provided, there shall be one folder titled with the RunName)
    * adat
      * \<output file prefix>\_Step3\_SampleNorm.adat: This file contains the final normalized counts for the samples in this project.
      * csv\_format: This folder contains the CSV version of the ADAT described above.
      * other\_normalization\_steps: This folder contains all raw count and intermediate normalized ADATs with their respective counts.
      * control\_only: This folder contains control-only ADATs that include counts from calibrator, QC, and blank samples only.
    * quality metrics
      * \<output file prefix>\_run\_qc\_stats.csv: This file provides a summary of per sample run qc statistics.
    * DRAGEN Report
* DRAGEN Report: Folder containing all projects DRAGEN report file/s
* Extra Files
* BioInsight Platform Core Logs: Folder containing information on Platform Core logs

## Example of Output Structure in BSSH

The example below contains an analysis configured with two Projects, "Project1" and "Project2". Analysis and outputs are processed for each project individually (check [Multi-Analysis section](/dragen-protein-quantification/run-setup/lane-splitting-and-project-splitting) for details).

<figure><img src="/files/ZJKwNxF793Djh8Bm3wqY" alt=""><figcaption><p>Each project folder (highlighted in yellow) includes "quality_metrics" and "adat" folders, along with a DRAGEN Report. The final normalized counts file (with suffix "_Step3_SampleNorm.adat") is available for each project and can be used in exploratory analysis.</p></figcaption></figure>


# DRAGEN Report

DRAGEN Reports is an HTML report that provides a quick overview of the quality of an E2E analysis. The report consists of three sections, which are displayed as tabs in the report. The following sections describe the tabs.

### **Plate QC**

This section is subdivided into the following subsections:

* **Reagent Lot Summary:** This table describes the reagents used per plate in the analysis.
* **Plate QC Summary:** This table provides the plate level metrics, including *Calibration % in Tails, QC % in Tails, Reference Correlation,* and *Blank Background* metrics.
* **Calibration Scale Factors:** This histogram illustrates the distribution of calibration scale factors for a given plate.
* **QC Scale Factors:** This histogram illustrates the distribution of QC scale factors for a given plate.

### **Sample QC**

This section of the report contains information on which samples passed or flagged specifications as well as SOMAmer count yield per sample. It is subdivided into the following subsections:

* **Sample QC Summary:** The table describes the percentage of samples (organized by sample type) that passed Quality Control.
* **QC Summary:** The heatmap shows the QC status of samples based on their position in the plate wells.
* **Flagged Samples:** The table identifies samples that failed one or more sequencing or normalization specifications.
* **SOMAmer Read Counts:** This graph represents the SOMAmer count for each sample included in the analysis.

### **Specification**

This report section details the QC metrics and normalization steps.


# ADAT Content - 9.5k

This page describes the output from the 9.5k Illumina Protein Prep assay, which differs slightly from the 6k assay. Use this for ADATs made on DRAGEN Protein Quantification v2.0 and higher.

The primary output file of DRAGEN Protein Quantification is the ADAT. One ADAT is produced for each normalization step performed. The final ADAT produced (\<output\_file\_prefix>\_Step3\_SampleNorm.adat) has the normalized counts at the end of the full normalization process.

There are four key components to the ADAT: the header, SOMAmer metadata, sample metadata, and counts.

<figure><img src="/files/lpvZL2ab4kx0dPMHakhs" alt=""><figcaption><p>ADAT Layout</p></figcaption></figure>

## ADAT Header

| Metric                                         | Description                                                                                                                                                                                                                                                                                                                                                                           |
| ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Version                                        | Version of the software used during analysis                                                                                                                                                                                                                                                                                                                                          |
| Title                                          | User reference for the anlaysis                                                                                                                                                                                                                                                                                                                                                       |
| SOMAmerReferenceSource                         | SOMAmer metadata version used during analysis                                                                                                                                                                                                                                                                                                                                         |
| AssayType                                      | Assay types used                                                                                                                                                                                                                                                                                                                                                                      |
| AssayVersion                                   | Version of the Illumina Protein Prep assay (for 9.5k product, value will be Illumina Protein Prep 9k)                                                                                                                                                                                                                                                                                 |
| AssayRobot                                     | Liquid handling robot                                                                                                                                                                                                                                                                                                                                                                 |
| RunId                                          | Unique identifier for the run (created by sequencing instrument)                                                                                                                                                                                                                                                                                                                      |
| Instrument Type                                | Instrument used for sequencing                                                                                                                                                                                                                                                                                                                                                        |
| Flowcell                                       | Flowcell used for sequencing                                                                                                                                                                                                                                                                                                                                                          |
| YieldDemux                                     | Total number of reads that are demultiplexed                                                                                                                                                                                                                                                                                                                                          |
| YieldQ30Demux                                  | Total number of reads that are demultiplexed with a passing QC score                                                                                                                                                                                                                                                                                                                  |
| Q30WeightedMean                                | Q30 primary sequencing metric, comuted as the Q30 weighted mean across lanes                                                                                                                                                                                                                                                                                                          |
| CreatedDate                                    | Date the run was performed                                                                                                                                                                                                                                                                                                                                                            |
| StudyOrganism                                  | Sample organism (human)                                                                                                                                                                                                                                                                                                                                                               |
| StudyMatrix                                    | Matrix used in study (serum or plasma)                                                                                                                                                                                                                                                                                                                                                |
| CalibratorId                                   | Lot number(s) of the calibrator samples                                                                                                                                                                                                                                                                                                                                               |
| ProcessSteps                                   | The process steps that occurred to produced the counts in the current ADAT                                                                                                                                                                                                                                                                                                            |
| PlateSOMAmerNormReads\_PassFlag                | PASS/FLAG for each plate. If greater than 70% of Blank Samples have SOMAmerNormReads\_PassFlag = FLAG, PlateSOMAmerNormReads\_PassFlag shall be WARNING. Otherwise, it shall have the value "PASS".                                                                                                                                                                                   |
| PlatformSpecificPlateScale\_ScaleFactor        | Scale factors for the PlateScale normalization step, when compared to a platform specific reference. The reference is based on data from the sequencing instrumented used for the run.                                                                                                                                                                                                |
| PlatformSpecificCalibrateTailPercent           | Percent of Platform Specific Calibration scale factors in tails (outside of the acceptable range of 0.6-1.4) for each plate when compared to a reference. The platform specific reference is based on data from the sequencing instrument used for the run.                                                                                                                           |
| PlatformSpecificCalibrateTailPercent\_PassFlag | PASS/WARNING for each plate. If PlatformSpecificCalibrateTailPercent for a plate is greater than .15, this value shall be "WARNING". If it's less than or equal to .15, it shall be "PASS".                                                                                                                                                                                           |
| CrossPlatformPlateScale\_ScaleFactor           | Scale factors for the PlateScale normalization step, when compared to a universal reference made with NovaSeq X data.                                                                                                                                                                                                                                                                 |
| CrossPlatformCalibrateTailPercent              | Percent of Cross Platform Calibration scale factors in tails (outside of the acceptable range of 0.6-1.4) for each plate when compared to a reference. The cross platform reference is based on data from NovaSeq X runs.                                                                                                                                                             |
| QCCheckTailPercent                             | Percent of QC scale factors in tails for each plate.                                                                                                                                                                                                                                                                                                                                  |
| QCCheckTailPercent\_PassFlag                   | PASS/FAIL for each plate. If QCCheckTailPercent for a plate is greater than .15, this value shall be "FAIL". If it's less than or equal to .15, it shall be "PASS".                                                                                                                                                                                                                   |
| IntraPlateMedianCV                             | <p>Median Percent CV metric for each plate, computed Sample Normalization, providing:</p><ul><li>(sample Standard Deviation of calibrator samples/Mean of calibrator samples)\*100</li><li>(sample Standard Deviation of QC samples/Mean of QC samples)\*100</li></ul>                                                                                                                |
| IntraPlateQ75CV                                | Same as IntraPlateMedianCV (above) but showing 75th percentile of Percent CV instead of median                                                                                                                                                                                                                                                                                        |
| IntraPlateQ90CV                                | Same as IntraPlateMedianCV (above) but showing 90th percentile of Percent CV instead of median                                                                                                                                                                                                                                                                                        |
| InterPlateMedianCV                             | <p>Median Percent CV metric across plates, providing:</p><ul><li>(sample Standard Deviation of plate mean count in calibrator samples/Mean of plate mean count in calibrator samples)\*100</li><li>(sample Standard Deviation of plate mean count in QC samples/Mean of plate mean count QC samples)\*100</li></ul><p>Provided only if there are 4 or more plates in the analysis</p> |
| InterPlateQ75CV                                | Same as InterPlateMedianCV (above) but showing 75th percentile of Percent CV instead of median                                                                                                                                                                                                                                                                                        |
| InterPlateQ90CV                                | Same as InterPlateMedianCV (above) but showing 90th percentile of Percent CV instead of median                                                                                                                                                                                                                                                                                        |
| MedianS2B                                      | <p>Median Signal to Background metric for each plate based on human somamers after Plate Normalization, providing:</p><ul><li>Median Calibrator/Median Blank</li><li>Median QC/Median Blank</li></ul>                                                                                                                                                                                 |
| GeneratedBy                                    | Version of SomaData parser used to write ADAT (ex. SomaData\_1.0.0)                                                                                                                                                                                                                                                                                                                   |

## Sample Metadata

<table><thead><tr><th width="324">Metric</th><th>Description</th></tr></thead><tbody><tr><td>SampleID</td><td>Sample identifier</td></tr><tr><td>PlateId</td><td>Unique plate identifier</td></tr><tr><td>MatrixTubeBarcode</td><td>Matrix tube barcode scanned during library prep</td></tr><tr><td>BatchID</td><td>User-provided batch identifier</td></tr><tr><td>InputType</td><td><p>Sample type specified in manifest file.</p><p>Examples: Plasma_Calibrator, Plasma_QC, Plasma, Serum_Calibrator, Serum_ QC, Serum, Blank</p></td></tr><tr><td>MatrixTube</td><td><p>Matrix used.</p><p>Examples: Plasma, Serum</p></td></tr><tr><td>SampleType</td><td><p>Type of sample processed.</p><p>Examples: Plasma, Serum, Blank, QC, Calibrator</p></td></tr><tr><td>ControlID</td><td>ID of the calibrator, QC, or blank lot (applied to controls only).</td></tr><tr><td>ProbePlate</td><td>Probe plate lot number, from library prep</td></tr><tr><td>SOMAmerBeadPlate</td><td>SOMAmer bead plate lot number. Last two digits indicate the master mix lot number.</td></tr><tr><td>WellPosition</td><td>Location of the sample on the 96 well plate (A1-H12)</td></tr><tr><td>Project</td><td>(Optional) User-provided project identifier</td></tr><tr><td>SOMAmerReads</td><td>Number of raw counts human SOMAmer reads (excluding control reads)</td></tr><tr><td>SOMAmerReads_PassFlag</td><td><p>This flag indicates whether a non-blank sample meets specifications for number of SOMAmer reads.<br><br>For non-blank samples:</p><ul><li>SOMAmerReads ≥ 10 million = PASS</li><li>SOMAmerReads &#x3C; 10 million = FLAG</li></ul><p>Not applied to blank samples.</p></td></tr><tr><td>SOMAmerNormReads</td><td>Number of normalized counts human SOMAmer reads at plate scale normalization step (excluding control reads)</td></tr><tr><td>SOMAmerNormReads_PassFlag</td><td><p>This flag indicates whether a blank sample meets the specifications for number of normalized SOMAmer reads.<br><br>For blank samples:</p><ul><li>SOMAmerNormReads ≤ 20 million = PASS</li><li>SOMAmerNormReads > 20 million = FLAG</li></ul><p>Not applied to non-blank samples<strong>.</strong></p></td></tr><tr><td>RefCorr</td><td>Spearman correlation value for comparing samples to the plasma or serum reference.</td></tr><tr><td>EmpericalHybTemp</td><td>This indicates the actual hybridization temperature of a sample, obtained empirically from a set of 78 temperature controls.</td></tr><tr><td>EmpiricalHybTemp_PassFlag</td><td><p>This flag indicates if a samples meets the specification for hybridization temperate.<br></p><ul><li>EmpiricalHybTemp &#x3C; 54.2 = PASS</li><li>EmpiricalHybTemp > 54.2 = FLAG</li></ul></td></tr><tr><td>HybNorm_1_ScaleFactor</td><td>The hybridization control scale factor</td></tr><tr><td>HybNorm_PassFlag</td><td>This flag indicates whether the HybNorm_1_ScaleFactor is within the specified acceptance criteria range of 0.4–2.5.</td></tr><tr><td>MedNormInt_5e-05_ScaleFactor</td><td>The MedNormInt scale factor for the 0.005% dilution group</td></tr><tr><td>MedNormInt_0.005_ScaleFactor</td><td>The MedNormInt scale factor for the 0.5% dilution group</td></tr><tr><td>MedNormInt_0.2_ScaleFactor</td><td>The MedNormInt scale factor for the 20% dilution group</td></tr><tr><td>MedNormInt_PassFlag</td><td>This flag indicates whether all dilution group scale factors are within the specified acceptance criteria range of 0.4–2.5 for the MedNormInt step.<br>Applies to blank and calibrator samples.</td></tr><tr><td>MedNormExt_5e-05_ScaleFactor</td><td>The MedNormExt scale factor for the 0.005% dilution group</td></tr><tr><td>MedNormExt_0.005_ScaleFactor</td><td>The MedNormExt scale factor for the 0.5% dilution group</td></tr><tr><td>MedNormExt_0.2_ScaleFactor</td><td>The MedNormExt scale factor for the 20% dilution group</td></tr><tr><td>MedNormExt_PassFlag</td><td>This flag indicates whether all dilution group scale factors are within the specified acceptance criteria range of 0.4–2.5 for the MedNormExt step.<br>Applied to QC and serum/plasma samples.</td></tr><tr><td>RowCheck_PassFlag</td><td>This flag indicates whether the sample's plate failed and if not whether all row scale factors are within the specified acceptance criteria range</td></tr></tbody></table>

## SOMAmer Metadata

| Metric                                             | Description                                                                                                                                                                                                                               |
| -------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SeqId                                              | Unique sequence identifier for the SOMAmer                                                                                                                                                                                                |
| SeqIdVersion                                       | SOMAmer sequence version                                                                                                                                                                                                                  |
| SomaId                                             | Somalogic-provided identifier                                                                                                                                                                                                             |
| Target                                             | Protein target identifier                                                                                                                                                                                                                 |
| Target Full Name                                   | Protein target full name                                                                                                                                                                                                                  |
| Type                                               | Target type (protein)                                                                                                                                                                                                                     |
| UniProt ID                                         | UniProt identifier                                                                                                                                                                                                                        |
| Entrez Gene ID                                     | Entrez Gene identifier                                                                                                                                                                                                                    |
| Entrez Gene Symbol                                 | Entrez Gene symbol                                                                                                                                                                                                                        |
| HybControl                                         | True is the SOMAmer is a hyb control, else False                                                                                                                                                                                          |
| LoD.Plasma                                         | Limit of detection for plasma SOMAmers                                                                                                                                                                                                    |
| LoD.Serum                                          | Limit of detection for serum SOMAmers                                                                                                                                                                                                     |
| Organism                                           | Organism (e.g. Human, Mouse) of the SOMAmer                                                                                                                                                                                               |
| PlatformSpecificCalibrate\_\<PlateId>\_ScaleFactor | Scale Factor for Platform Specific Calibration                                                                                                                                                                                            |
| CrossPlatformCalibrate\_\<PlateId>\_ScaleFactor    | Scale Factor for Cross Platform Calibration                                                                                                                                                                                               |
| QCCheck\_\<PlateId>\_ScaleFactor                   | Scale Factor for QC Check                                                                                                                                                                                                                 |
| Dilution                                           | Dilution group classification for the SOMAmer                                                                                                                                                                                             |
| DRC\_Level                                         | Compression of SOMAmer range that occurred during the assay.                                                                                                                                                                              |
| References                                         | References used in the analysis.                                                                                                                                                                                                          |
| Units                                              | Units of matrix content                                                                                                                                                                                                                   |
| BlockList                                          | Whether the somamer is excluded from normalization and downstream analysis due to known poor NGS performance. This metric only appears in raw count ADATs. Other ADATs contain somamers with BlockList set to FALSE and omit this metric. |


# ADAT Content - 6k

This page describes the output from the 6k Illumina Protein Prep assay, which differs slightly from the 9.5k assay. Do not reference this page unless your analysis was created on DRAGEN Protein Quantification v1.8 or earlier.

The primary output file of DRAGEN Protein Quantification is the ADAT. One ADAT is produced for each normalization step performed. The final ADAT produced (\<output\_file\_prefix>\_Step7\_FinalNormStep\_MedNormExt.adat) has the normalized counts at the end of the full normalization process.

There are four key components to the ADAT: the header, SOMAmer metadata, sample metadata, and counts.

<figure><img src="/files/lpvZL2ab4kx0dPMHakhs" alt=""><figcaption><p>ADAT Layout</p></figcaption></figure>

## ADAT Header

| Metric                                  | Description                                                                                                                                                                                                                                                 |
| --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Version                                 | Version of the software used during analysis                                                                                                                                                                                                                |
| Title                                   | User reference for the anlaysis                                                                                                                                                                                                                             |
| OutputDirectory                         | Not applicable                                                                                                                                                                                                                                              |
| SOMAmerReferenceSource                  | SOMAmer metadata version used during analysis                                                                                                                                                                                                               |
| AssayType                               | Assay types used                                                                                                                                                                                                                                            |
| SiteId                                  | Not applicable                                                                                                                                                                                                                                              |
| AssayVersion                            | Version of the Illumina Protein Prep assay (for 6k product, value will be Illumina Protein Kit 1.0)                                                                                                                                                         |
| AssayRobot                              | Liquid handling robot                                                                                                                                                                                                                                       |
| RunId                                   | Unique identifier for the run (created by sequencing instrument)                                                                                                                                                                                            |
| Instrument Type                         | Instrument used for sequencing                                                                                                                                                                                                                              |
| NGSLot                                  | Internal reference for control samples                                                                                                                                                                                                                      |
| Flowcell                                | Flowcell used for sequencing                                                                                                                                                                                                                                |
| YieldDemux                              | Total number of reads that are demultiplexed                                                                                                                                                                                                                |
| YieldQ30Demux                           | Total number of reads that are demultiplexed with a passing QC score                                                                                                                                                                                        |
| Q30WeightedMean                         | Q30 primary sequencing metric, comuted as the Q30 weighted mean across lanes                                                                                                                                                                                |
| CreatedBy                               | Not applicable                                                                                                                                                                                                                                              |
| EnteredBy                               | Not applicable                                                                                                                                                                                                                                              |
| ExpDate                                 | Not applicable                                                                                                                                                                                                                                              |
| CreatedDate                             | Date the run was performed                                                                                                                                                                                                                                  |
| Notes                                   | Not applicable                                                                                                                                                                                                                                              |
| StudyOrganism                           | Sample organism (human)                                                                                                                                                                                                                                     |
| StudyMatrix                             | Matrix used in study (serum or plasma)                                                                                                                                                                                                                      |
| CalibratorId                            | Lot number(s) of the calibrator samples                                                                                                                                                                                                                     |
| ReportConfig                            | Configuration used for normalization                                                                                                                                                                                                                        |
| ProcessSteps                            | The process steps that occurred to produced the counts in the current ADAT                                                                                                                                                                                  |
| PlatformSpecificPlateScale\_ScaleFactor | Scale factors for the PlateScale normalization step, when compared to a platform specific reference. The reference is based on data from the sequencing instrumented used for the run.                                                                      |
| PlatformSpecificCalibrateTailPercent    | Percent of Platform Specific Calibration scale factors in tails (outside of the acceptable range of 0.6-1.4) for each plate when compared to a reference. The platform specific reference is based on data from the sequencing instrument used for the run. |
| CrossPlatformPlateScale\_ScaleFactor    | Scale factors for the PlateScale normalization step, when compared to a universal reference made with NovaSeq X data.                                                                                                                                       |
| CrossPlatformCalibrateTailPercent       | Percent of Cross Platform Calibration scale factors in tails (outside of the acceptable range of 0.6-1.4) for each plate when compared to a reference. The cross platform reference is based on data from NovaSeq X runs.                                   |
| QCCheckTailPercent                      | Percent of QC scale factors in tails for each plate.                                                                                                                                                                                                        |
| QCCheckTailPercent\_PassFlag            | PASS/FAIL for each plate. If QCCheckTailPercent for a plate is greater than .15, this value shall be "FAIL". If it's less than or equal to .15, it shall be "PASS".                                                                                         |
| GeneratedBy                             | Version of SomaData parser used to write ADAT (ex. SomaData\_1.0.0)                                                                                                                                                                                         |

## Sample Metadata

<table><thead><tr><th width="324">Metric</th><th>Description</th></tr></thead><tbody><tr><td>SampleID</td><td>Sample identifier</td></tr><tr><td>PlateId</td><td>Unique plate identifier</td></tr><tr><td>InputType</td><td><p>Sample type specified in manifest file.</p><p>Examples: Plasma_Calibrator, Plasma_QC, Plasma, Serum_Calibrator, Serum_ QC, Serum, Blank</p></td></tr><tr><td>MatrixTube</td><td><p>Matrix used.</p><p>Examples: Plasma, Serum</p></td></tr><tr><td>SampleType</td><td><p>Type of sample processed.</p><p>Examples: Plasma, Serum, Blank, QC, Calibrator</p></td></tr><tr><td>CalibratorId</td><td>Calibrator lot/ID.</td></tr><tr><td>WellPosition</td><td>Location of the sample on the 96 well plate (A1-H12)</td></tr><tr><td>SOMAmer Reads</td><td>Number of raw counts human SOMAmer reads (excluding control reads)</td></tr><tr><td>SOMAmerReads_PassFlag</td><td><p>This flag indicates whether a non-blank sample meets specifications for number of SOMAmer reads.<br><br>For non-blank samples:</p><ul><li>SOMAmerReads ≥ 10 million = PASS</li><li>SOMAmerReads &#x3C; 10 million = FLAG</li></ul><p>Not applied to blank samples.</p></td></tr><tr><td>HybNorm_1_ScaleFactor</td><td>The hybridization control scale factor for hyb plate 1</td></tr><tr><td>HybNorm_2_ScaleFactor</td><td>The hybridization control scale factor for hyb plate 2</td></tr><tr><td>HybNorm_PassFlag</td><td>This flag indicates whether the HybNorm_1_ScaleFactor is within the specified acceptance criteria range of 0.4–2.5.</td></tr><tr><td>MedNormInt_5e-05_ScaleFactor</td><td>The MedNormInt scale factor for the 0.005% dilution group</td></tr><tr><td>MedNormInt_0.005_ScaleFactor</td><td>The MedNormInt scale factor for the 0.5% dilution group</td></tr><tr><td>MedNormInt_0.2_ScaleFactor</td><td>The MedNormInt scale factor for the 20% dilution group</td></tr><tr><td>MedNormInt_PassFlag</td><td>This flag indicates whether all dilution group scale factors are within the specified acceptance criteria range of 0.4–2.5 for the MedNormInt step.<br>Applies to blank and calibrator samples.</td></tr><tr><td>MedNormExt_5e-05_ScaleFactor</td><td>The MedNormExt scale factor for the 0.005% dilution group</td></tr><tr><td>MedNormExt_0.005_ScaleFactor</td><td>The MedNormExt scale factor for the 0.5% dilution group</td></tr><tr><td>MedNormExt_0.2_ScaleFactor</td><td>The MedNormExt scale factor for the 20% dilution group</td></tr><tr><td>MedNormExt_PassFlag</td><td>This flag indicates whether all dilution group scale factors are within the specified acceptance criteria range of 0.4–2.5 for the MedNormExt step.<br>Applied to QC and serum/plasma samples.</td></tr><tr><td>ANML_5e-05_ScaleFactor</td><td>The ANML scale factor for the 0.005% dilution group. Note: ANML norm is not output.</td></tr><tr><td>ANML_0.005_ScaleFactor</td><td>The ANML scale factor for the 0.5% dilution group. Note: ANML norm is not output.</td></tr><tr><td>ANML_0.2_ScaleFactor</td><td>The ANML scale factor for the 20% dilution group. Note: ANML norm is not output.</td></tr><tr><td>ANML_5e-05_fraction_used</td><td>Fraction of probes in 0.005% dilution group used to compute corresponding ANML scale factor.</td></tr><tr><td>ANML_0.005_fraction_used</td><td>Fraction of probes in 0.5% dilution group used to compute corresponding ANML scale factor.</td></tr><tr><td>ANML_0.2_fraction_used</td><td>Fraction of probes in 20% dilution group used to compute corresponding ANML scale factor.</td></tr><tr><td>RowCheck_PassFlag</td><td>This flag indicates whether all row scale factors are within the specified acceptance criteria range.</td></tr><tr><td>PlateRunDate</td><td>Not applicable</td></tr><tr><td>SampleNotes</td><td>Not applicable</td></tr><tr><td>Barcode</td><td>Not applicable</td></tr></tbody></table>

## SOMAmer Metadata

| Metric                                             | Description                                                                                                                                                                                                                                                                             |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SeqId                                              | Unique sequence identifier for the SOMAmer                                                                                                                                                                                                                                              |
| SeqIdVersion                                       | SOMAmer sequence version                                                                                                                                                                                                                                                                |
| SomaId                                             | Somalogic-provided identifier                                                                                                                                                                                                                                                           |
| Target                                             | Protein target identifier                                                                                                                                                                                                                                                               |
| Target Full Name                                   | Protein target full name                                                                                                                                                                                                                                                                |
| Type                                               | Target type (protein)                                                                                                                                                                                                                                                                   |
| UniProt                                            | UniProt identifier                                                                                                                                                                                                                                                                      |
| EntrezGeneID                                       | Entrez Gene identifier                                                                                                                                                                                                                                                                  |
| EntrezGeneSymbol                                   | Entrez Gene symbol                                                                                                                                                                                                                                                                      |
| HybControl                                         | True if the SOMAmer is a hyb control, else False                                                                                                                                                                                                                                        |
| MedNormControl                                     | True if the SOMAmer is used during MedNormExt normalization, else False                                                                                                                                                                                                                 |
| LoD.Plasma                                         | Limit of detection for plasma SOMAmers                                                                                                                                                                                                                                                  |
| LoD.Serum                                          | Limit of detection for serum SOMAmers                                                                                                                                                                                                                                                   |
| Organism                                           | Organism (e.g. Human, Mouse) of the SOMAmer                                                                                                                                                                                                                                             |
| PlatformSpecificCalibrate\_\<PlateId>\_ScaleFactor | Scale Factor for Platform Specific Calibration                                                                                                                                                                                                                                          |
| PlatformSpecificCalibrate\_\<PlateId>\_PassFlag    | Flag to indicates whether the PlatformSpecificCalibrate scale factor for this SeqId was in the acceptance criteria range of 0.6-1.4. Note: this flag is **only** used for evaluating the PlatformSpecificCalibrateTailPercent and should not be used to exclude SOMAmers from analysis. |
| CrossPlatformCalibrate\_\<PlateId>\_ScaleFactor    | Scale Factor for Cross Platform Calibration                                                                                                                                                                                                                                             |
| CrossPlatformCalibrate\_\<PlateId>\_PassFlag       | Flag to indicates whether the CrossPlatformCalibrate scale factor for this SeqId was in the acceptance criteria range of 0.6-1.4. Note: this flag is **only** used for evaluating the CrossPlatformCalibrateTailPercent and should not be used to exclude SOMAmers from analysis.       |
| QCCheck\_\<PlateId>\_ScaleFactor                   | Scale Factor for QC Check                                                                                                                                                                                                                                                               |
| QCCheck\_\<PlateId>\_PassFlag                      | Flag to indicates whether the QCCheck scale factor for this SeqId was in the acceptance criteria range of 0.8-1.2. Note: this flag is **only** used for evaluating the QCCheckTailPercent and should not be used to exclude SOMAmers from analysis.                                     |
| ColCheck\_\<PlateId>\_PassFlag                     | This flag indicates whether all SOMAmer scale factors are within the specified acceptance criteria range. Note: this flag should not be used to exclude SOMAmers from analysis.                                                                                                         |
| Dilution                                           | Dilution group classification for the SOMAmer                                                                                                                                                                                                                                           |
| DRC                                                | Compression of SOMAmer range that occurred during the assay.                                                                                                                                                                                                                            |
| HCG                                                | Hybridization Control Group. Identifies the hybridization plate of the SOMAmer's NGS reporter.                                                                                                                                                                                          |
| BlackList                                          | Identifies SOMAmers with low performing NGS reporters in product testing. SOMAmers with the value TRUE should be excluded from analysis.                                                                                                                                                |
| References (Ref.\*)                                | References available to the SW version used for analysis.                                                                                                                                                                                                                               |
| Units                                              | Units of matrix content                                                                                                                                                                                                                                                                 |


# Using the ADAT

ADATs can be analyzed in R or Python using parsers created by Somalogic.

DRAGEN Protein Quantification v2.0.0 is compatible with SomaData v1.0.0 (Python - formerly called Canopy) and SomaDataIO v6.1.0 (R).

* Python: <https://github.com/SomaLogic/Canopy>
* R: <https://github.com/SomaLogic/SomaDataIO>

Somadata creates an ADAT object, which is an extension of a Pandas DataFrame.

* rows correspond to samples
* columns correspond to SOMAmers
* values are normalized counts

Below are examples on parsing an ADAT in Python and R.

### Parsing ADAT in Python

```python
import somadata

# read the adat
my_adat = somadata.read_adat('/path/to/my/file1.adat')

# retrieve sample metadata
sample_meta = my_adat.index.to_frame(index=False)

# retrieve SOMAmer metadata
soma_meta = my_adat.columns.to_frame(index=False)

# retrieve the scale factors for all plates
plate_scale_factors_dict = my_adat.header_metadata['PlatformSpecificPlateScale_ScaleFactor']

### additional optional manipulation 

# concatenate multiple adats into a single adat
my_adat2 = somadata.read_adat('/path/to/my/file2.adat')
my_merged_adat = soma.smart_adat_concatenation([my_adat, my_adat2])

# write it to a file 
my_merged_adat = my_merged_adat.to_adat('/path/to/merged/adat')
```

### Parsing ADAT in R

```r
library(SomaDataIO)

# check all package functions
ls("package:SomaDataIO")

#read the adat
my_adat <- read_adat('/path/to/my/file1.adat')

# retrieve sample metadata
sample_meta <- my_adat[getMeta(my_adat)]

# retrieve SOMAmer metadata
soma_meta <- my_adat[getAnalytes(my_adat)]

# create a function that parses the header
parse_adat_header <- function(adat_file, max_lines = 500) {
  header_lines <- readLines(adat_file, n = max_lines)
  table_row <- grep("TABLE_BEGIN", header_lines)
  header_lines[1:(table_row-1)] %>% 
    strsplit("\t") %>% 
    {
      values <- map(., 2) # Get values from key value pairs
      keys <- map_chr(., 1) # Get keys...
      setNames(values, keys)
    }
}

my_adat_header <- parse_adat_header('/path/to/my/file1.adat')

# retrieve the scale factors for all plates
plate_scale_factors_dict = my_adat_header$PlatformSpecificPlateScale_ScaleFactor

### additional optional manipulation 
my_adat2 = read_adat('/path/to/my/file2.adat')
my_merged_adat = rbind(my_adat, my_adat2)

# write it to a file 
my_merged_adat = write_adat(my_merged_adat, '/path/to/merged/adat')

```


# Compatibility with Excel

While using the R and Python parsers is the recommended way to manipulate ADAT files, it's also possible to open and view them in Excel.

CSV version of the ADAT with final normalized counts is also provided under `adat/csv_format/` output directory.

### Opening an ADAT in Excel (Windows)

1. Open a blank excel workbook and browse for a file.
2. Confirm that "All Files (\*.\*)" are searchable.

<div align="left"><figure><img src="/files/lzrYxdTPRuSlPknyi4zw" alt="" width="375"><figcaption></figcaption></figure></div>

3. Accept default parameters from Excel.

<div align="left"><figure><img src="/files/IG4p8Pj9UaK4zonYessa" alt="" width="304"><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/QDMDLUValCG7WykgmSbv" alt="" width="305"><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/SrfbiCMzUYxzbQQuK2xA" alt="" width="304"><figcaption></figcaption></figure></div>

4. View ADAT in Excel!

###

### Opening an ADAT in Excel (Mac)

1. Open a blank excel workbook and select File --> Import.

![](/files/yeYg8OephLXsh2CxGaZV)

2. Select the previously downloaded .adat file.

<div align="left"><figure><img src="/files/MrVwRVaCg20MzzH7kF7m" alt="" width="375"><figcaption></figcaption></figure></div>

3. After Click on "Get Data", select "Text file" and click on "Import".

<div align="left"><figure><img src="/files/dhOqtDdFq72qi1obUefb" alt="" width="375"><figcaption></figcaption></figure></div>

4. Select Delimited (default) and click on "Next >".

<div align="left"><figure><img src="/files/4ObLBeOQfyfXBxrdbcpK" alt="" width="375"><figcaption></figcaption></figure></div>

5. Select "Tab" (default) and click on "Next >".

<div align="left"><figure><img src="/files/Z95EypY8oF8riKocxiHN" alt="" width="375"><figcaption></figcaption></figure></div>

6. Select "General" (default) and click on "Next >".

<div align="left"><figure><img src="/files/anNoVQwmaREv2ilDXBIO" alt="" width="375"><figcaption></figcaption></figure></div>

7. Select the sheet and click on "Import".

![](/files/k89DvpYzY8xwC5bWFdbf)

7. View the ADAT in Excel!


# Illumina Connected Multiomics Walkthrough

Illumina Connected Multiomics provides interactive visualizations and powerful statistics. This is a walkthrough of an analysis that could be done in Connected Multiomics with an example proteomic data set, produced by DRAGEN Protein Quantification. It covers the following features:

* Creating a default analysis
* Creating a custom analysis
* Managing sample metadata
* Filtering samples
* Filtering features
* Data Transformation
* PCA
* Differential expression
* Hierarchical clustering and creating heatmaps
* Gene set enrichment analysis

For information on the Connected Multiomics Platform, including how to log in, please reference the following documentation: <https://help.multiomics.illumina.com/icm>

### Demo Data <a href="#create-an-advanced-analysis" id="create-an-advanced-analysis"></a>

Demo data that can be used to follow along with this walkthrough is found in the Connected Multiomics Demo Data repository. To add this dataset to a study, perform the following steps:

<div align="left"><figure><img src="/files/57vMG5cmEDx2v7HbWb50" alt=""><figcaption></figcaption></figure></div>

<figure><img src="/files/48mLDMJLh60xNVRXvZdR" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/gXt51R1O5uwH9FnyeFiZ" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/LBt84IdUrlLEX25WUphE" alt=""><figcaption></figcaption></figure>

After clicking "+ Add Demo Data", the data used in this walkthrough can be found at /Multiomics-Demo-Data/Proteomics/NovaSeq 6k-S4 Cancer-Normal. For this study, both SampleType (CRC/Control) and TimePoint (T1...T8) are used. This data must be ingested prior to starting the analysis. Add both the ADAT (counts) and TSV (metadata) to the study.

### Creating a Default Analysis <a href="#create-an-advanced-analysis" id="create-an-advanced-analysis"></a>

* Click on '+ New Analysis'.
* In the pop-up window, provide a name for the analysis, select **‘Default: Illumina Proteomics**’ as the Analysis Type, choose the sample group to be included in the analysis ('All 9k Illumina protein Prep Samples' will be selected by default), and click on the ‘Run Analysis’ button.
* Exploring the PCA plot:

  * By default, the plot is colored by BatchID. To change this, in the left hand bar, select Configure > Style > Color by.

  <figure><img src="/files/EsjOuEqBsdZFuvjjl6HQ" alt=""><figcaption></figcaption></figure>

  * You may also want to explore other principal components. To do this, on the left side bar for the PCA plot, choose configure, axes, and update the data for each axis.

<figure><img src="/files/aOW4iUG4qDPQK3PvF8F1" alt=""><figcaption></figcaption></figure>

### Creating a Custom Analysis <a href="#create-an-advanced-analysis" id="create-an-advanced-analysis"></a>

<figure><img src="/files/Kx0uvi2G8P56b4n0R50T" alt=""><figcaption><p>Custom Analysis Example</p></figcaption></figure>

* Click on ‘+ New Analysis’.
* In the pop-up window, provide a name for the analysis, select ‘**Custom: Illumina Proteomics**’ as the Analysis Type, choose the sample group to be included in the analysis ('All 9k Illumina protein Prep Samples' will be selected by default), and click on the ‘Run Analysis’ button.

​​![](/files/v9RgqgFcK85WSQuZPa4x)

* Note: make sure there are no duplicated Sample IDs in the analysis groups.
* A pop-up message will show up if the analysis creation is successful.

​![](/files/hhcgrXGAQ20BmYbNrJyk)

* Refresh the page to get the latest status of the analysis.
* When the Status is ‘Complete’, click on the analysis tile to enter the analysis module.

​![](/files/ekot1R7L2z2rWeQuVxWu)

* There is no default initiated analysis for the custom proteomic data. To review the number of samples and features, hover over the data node.

<div align="left"><figure><img src="/files/Fw1SjarZy75x2HwfdHcT" alt="" width="563"><figcaption></figcaption></figure></div>

* Throughout the below analyses, rectangles/task nodes will produce circles/data nodes. Double clicking on the task node will describe the task as it occurred, and double clicking on the data node will take you to the results of the analysis. Nodes will be greyed out while the analysis is still in progress.

### Managing sample metadata

* Click on 'Metadata' tab to view and add sample metadata.

<div align="left"><figure><img src="/files/98KZIGEyEnKHhn2hohIk" alt="" width="563"><figcaption></figcaption></figure></div>

* Click on '**Manage**' under 'Sample attributes' to reorder the metadata. Drag 'SampleStatus' and 'TimePoint' boxes to the front since they are the features that need to be colored for the downstream analysis. You can also add/remove/reorder other metadata or add new category to the current metadata in this page.

<div align="left"><figure><img src="/files/AFxkO9uNPnTjsGUrrcYH" alt="" width="563"><figcaption></figcaption></figure></div>

### Filtering samples

* Return to the analysis page, click on the 'Quantification' node, choose 'Filtering' > 'Filter samples' from the right hand tool box.

<div align="left"><figure><img src="/files/F9QA834buNignd1CJe5z" alt="" width="310"><figcaption></figcaption></figure></div>

* Select the samples with TimePoint T1 and T2.

<div align="left"><figure><img src="/files/KOGjR6jEr7X7E2Vrt2C2" alt="" width="375"><figcaption></figcaption></figure></div>

### Filtering features

* Click on 'Finish' and return to the analysis page. Click on the 'TimePoint in T1,T2' node, choose 'Filtering' > 'Filter features' from the right hand tool box.

<div align="left"><figure><img src="/files/3l04YzZQMvnCDlTi5OoN" alt="" width="300"><figcaption></figcaption></figure></div>

* Filter only include human protein by choosing **Metadata** option and specify filter criteria as inlucde **Organism** in **Human**, click **Finish**.

<div align="left"><figure><img src="/files/2M1QBDUVLMhkTfe1Pdnn" alt="" width="563"><figcaption></figcaption></figure></div>

Double click the feature filter output data node to open the report, check the distribution of the data. If the min is 0, perform the **Normalization** task to add 1. If the min is not 0, skip the normalization step since the data is already normalized.

### Normalization

* Click on the 'Filtered counts' node, choose 'Normalization and Scaling' > 'Normalization' from the right hand tool box.

<div align="left"><figure><img src="/files/N3Wt7xUZwBSURjig3JRJ" alt="" width="310"><figcaption></figcaption></figure></div>

* Choose 'Add' and drag it to the right-hand box to avoid 0 counts. This prevents any 0 count values which could impact Limma-trend differential analysis, which assumes continuous data. Then click on 'Finish' to return to the analysis dashboard.

<figure><img src="/files/EtM6b7zhN0nmjBscZJp9" alt=""><figcaption></figcaption></figure>

### PCA

* Click on the 'Normalized counts' node, select '**Exploratory analysis' > 'PCA**' from the right hand tool box.

<div align="left"><figure><img src="/files/ruubjqZMD4CuPh663dPR" alt="" width="307"><figcaption></figcaption></figure></div>

* Use the default setting and click 'Finish'.

<div align="left"><figure><img src="/files/VH0899rEEkE4A1ZnrVgl" alt="" width="375"><figcaption></figcaption></figure></div>

* Double click on the 'PCA' node to view the PCA report.
  * The scatter plot shows the data distribution (colored by SampleStatus) among the first three PCs.
  * The scree plot (top right panel) shows the variant represented by each PC.
  * The component loading table (bottom right panel) shows the correlation between every protein/SOMAmer and each PC. The variable in this table represents the SOMAmer's SeqID.
  * For additional information on PCA, review the following documentation: <https://help.partek.illumina.com/partek-flow/user-manual/task-menu/exploratory-analysis/pca>

<div align="left"><figure><img src="/files/bChlR03E4hc0aFwb3fbF" alt="" width="563"><figcaption></figcaption></figure></div>

### Differential expression

* Click on the 'Normalized counts' node, select '**Statistics' > 'Differential analysis**' from the right hand menu.

<div align="left"><figure><img src="/files/K1DRKgr0IusBavR44T6k" alt="" width="312"><figcaption></figcaption></figure></div>

* Select 'Limma-trend' (default) method and click 'Next'.

<figure><img src="/files/NKcSjFHZT6INk0BaKBFe" alt=""><figcaption></figcaption></figure>

NOTE: Limma-trend is a robust model that fits the assumptions for small sample sizes of normalized protein counts. The Limma-trend model is also flexible with categorical and quantitative variables. For other datasets or experimental designs, consider other methods.

* Select 'SampleStatus', 'TimePoint', 'DonorID' then click 'Add factors'. Select 'SampleStatus' and 'TimePoint' then click on 'Add interaction' to add the factors. Click on 'Next' to set up comparisons.

<div align="left"><figure><img src="/files/0R6gU0nZCuPrBKYFPxmp" alt="" width="375"><figcaption></figcaption></figure></div>

* Drag 'CRC' to the top right box and 'Control' to the bottom right box. Click on 'Add comparison'. Then Select 'SampleStatus\*TimePoint' from the Factor dropdown menu. Add T1 and T2 comparison between CRC and Control. Keep "Combine" selected for each of these comparisons.

<div align="left"><figure><img src="/files/T7h62fKsILLfNhz45onk" alt="" width="375"><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/qskFnFkAOrmGQIxMuFtJ" alt="" width="375"><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/bFfqgmY1A5EQmIgtPMVB" alt=""><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/CiwHjLStTrKhIOONVqht" alt=""><figcaption></figcaption></figure></div>

* No need to perform anymore filter or normalization, so keep the *Low value filte*r unchecked and choose **None** for *count normalization*.

<div align="left"><figure><img src="/files/7byc2daXTRb3rnUltqVd" alt="" width="563"><figcaption></figcaption></figure></div>

* Click on '**Finish**' bottom at the bottom.
* When the task is done, double click on the output node to view the report. On the left hand menu
  * select 'FDR', choose 'Per contrast' and specify 0.05 for CRC vs Control comparison
  * select 'Fold change', choose 'Per contranst' and specify -2 to 2 for CRC vs Control comparison.
  * click on 'Generate Filtered Node'
  * repeat this process on CRC T1 vs Control T1 comparison and CRC T2 vs Control T2 comparison.

<div align="left"><figure><img src="/files/09MEgAAEZZLwwAbzqCbO" alt="" width="563"><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="/files/D6QpUPprfb6opElnVqYB" alt="" width="126"><figcaption></figcaption></figure></div>

* Return to the analyses dashboard and there will be 3 filtered feature list nodes added to the pipeline; right click on the 'Filtered feature list' node and click on 'Rename data node' to rename the node as 'T vs N'; apply the same procedure to the other two filtered feature lists and rename them as 'T vs N Time 1' and 'T vs N Time 2' respectively.

<div align="left"><figure><img src="/files/46tXVumOLLgWLiLaCkJN" alt="" width="375"><figcaption></figcaption></figure></div>

* To compare the filtered feature lists, click on the 'Venn diagram' on the bottom menu and tick on the filtered lists ('T vs N', 'T vs N Time 1' and 'T vs N Time 2'); then click on 'Display selection' button on the bottom to visualize the Venn diagram.

<figure><img src="/files/s45GNMIcoHNhz3TAKez1" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/8JXxBuB77zqNTm7uPHCU" alt=""><figcaption></figcaption></figure>

### Hierarchical clustering and creating heatmaps

* Click on 'T vs N' data node, select '**Exploratory analysis' > 'Hierarchical clustering / heatmap**' from the right hand tool box.

<div align="left"><figure><img src="/files/TzqAdrSr8d1x0jxGghUo" alt="" width="289"><figcaption></figcaption></figure></div>

* Choose 'Heatmap' and select the feature order and sample order.
  * Choose 'Cluster' (default) as feature order
  * Choose 'Assign order' and select 'SampleStatus' from the dropdown menu.
  * Click on 'Finish' button at the bottom of the page.

<figure><img src="/files/6SpBZ7jNwBKnvhqFi8wG" alt=""><figcaption></figcaption></figure>

* Double click on the 'Hierarchical clustering / heatmap' node to view it.

<figure><img src="/files/nZZdMzhQP8tYShSR0fYs" alt=""><figcaption></figcaption></figure>

* For additional information on hierarchical clustering, view the following documentation: <https://help.partek.illumina.com/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering>

### Correlation Engine

* Click on 'T vs N' data node, select '**Biological interpretation > Correlation Engine pathway**' from the right hand tool box.

<div align="left"><figure><img src="/files/5SjJf2cLkUCPt326MUfw" alt="" width="248"><figcaption></figcaption></figure></div>

* Choose **Homo sapiens** as O*rganism*, **Protein Expression** as *Data type*, specify a name of a *Correlation Engine projec*t and *study name.*

<div align="left"><figure><img src="/files/KKp8xx4GDqCxdNXyn4ue" alt="" width="563"><figcaption></figcaption></figure></div>

* Click **Next**
* Specify Contrasts, multiple contrasts can be selected, use gene symbol as identifier. Click on '**Finish**' button at the bottom of the page.

<div align="left"><figure><img src="/files/BfbxrgvPQ5ZeQL7JPC4W" alt="" width="563"><figcaption></figcaption></figure></div>

* Double click on the **Correlation Engine** node to view the report.

<div align="left"><figure><img src="/files/0sx47R9eg6L4XQYgzAgg" alt="" width="563"><figcaption></figcaption></figure></div>

* Click on **Open Data Viewer auto session** to view the top 30 gene sets in *Data viewer*.

<div align="left"><figure><img src="/files/gMaovbrLafOd3zYTMuCI" alt="" width="563"><figcaption></figcaption></figure></div>

### GSEA

* To detect differential pathways between diseased and control samples, click on the 'Normalized counts' node (or filtered count data node) and select '**Biological interpretation > GSEA**' from the right hand tool box.

<div align="left"><figure><img src="/files/ULaMvdMtT2ClMtDGXqjQ" alt="" width="280"><figcaption></figcaption></figure></div>

* Select 'KEGG database' (default) and click on '**Next**' button at the bottom of the page.

<div align="left"><figure><img src="/files/ukpwPYLSxMrbjYeZMxAf" alt="" width="563"><figcaption></figcaption></figure></div>

* Select 'SampleStatus' and click on '**Next**'.

<div align="left"><figure><img src="/files/keLJBOfhpUnFGBiJxqpw" alt="" width="385"><figcaption></figcaption></figure></div>

* Drag 'CRC' to the top right box and 'Control' to the bottom right box. Keep "Combine" selected for this comparison. Click on '**Add comparison**'.

<div align="left"><figure><img src="/files/zvLIz6v5lqfElmw8Qc51" alt="" width="375"><figcaption></figcaption></figure></div>

* No need to perform filter or normalization, click **Finish** on the bottom of the page.
* Double click on the 'GSEA' node to view the results.

<div align="left"><figure><img src="/files/g7scX21oiWBD8mSiO7wm" alt="" width="563"><figcaption></figcaption></figure></div>

* Click on the enrichment plot icon after each row index to visualize the enrichment score of the corresponding pathway.

<figure><img src="/files/7pFnDjlbALSdSEFlwuTO" alt=""><figcaption></figcaption></figure>

<mark style="color:blue;">**NOTE**</mark>. For clarify on the differences between Gene Set Enrichment Analysis and GSEA, please view this documentation: <https://help.partek.illumina.com/partek-flow/frequently-asked-questions#what-is-the-difference-between-gsea-and-gene-set-enrichment>


# FAQs

## Process FAQs

* How can I share data with a collaborator in BSSH?
  * See the following documentation: <https://help.connected.illumina.com/basespace-sequence-hub/collaborate/share-with-collaborators>
* I sequenced in manual mode/need to kick of autolaunch after sequencing is complete. How can I do this?
  * Use the BSSH CLI to upload the run folder to BSSH. Make sure the samplesheet is named "SampleSheet.csv" and you are using an up-to-date version of the tool - at least v1.6.1.
  * For additional information, see the following documentation:
    * <https://help.connected.illumina.com/basespace-sequence-hub/cmd-line-interfaces/basespace-cli>
* How do I download/export the samplesheet?

  * Go to the BSSH Run Planner:
    * Go to BSSH, select the intended workgroup on the top right, and then select the Runs tab
    * Additional information on the BSSH Run Planner: <https://help.basespace.illumina.com/sequence/plan-runs>

  <figure><img src="/files/A5FHnH0l2jWDCklrDsp1" alt=""><figcaption></figcaption></figure>

  * For a new run:

    * Select the "New Run" button on the right and plan a new run with the desired pipeline.
    * After ingesting the IPPAS output files in the BSSH Run Planner, the user will reach the following Run Review page:

    <figure><img src="/files/wZ4PI1k2bZymZmdSc3an" alt=""><figcaption></figcaption></figure>

    * Select the "Export" option on the bottom right to download the samplesheet created
    * Select the "Save as Draft" or "Save as Planned" options to save this run as a draft or planned run respectively<br>
  * For an "Active" or "Planned" run:

    * Select the intended tab mentioned under the Runs header.
    * Select the required run from the list of runs:

    <figure><img src="/files/Te9TCxw8oRY76p6Ey5VR" alt=""><figcaption></figcaption></figure>

    * Next, click File > Download > Sample Sheet:

    <figure><img src="/files/Sexw4bJBiAmEdEUccYY4" alt=""><figcaption></figcaption></figure>
* How can I requeue a run on BSSH with a new samplesheet/new version of the pipeline?

  * Go to the BSSH Run Planner tool and plan a new run with the desired pipeline.
  * Save the exported samplesheet (see the above step for additional information).
  * Go to the original run, and click Status > Requeue > Planned Run.

  <figure><img src="/files/CfjDgpB0kmuAdJKZXcTd" alt="" width="553"><figcaption></figcaption></figure>

  * Select "Use a new Sample Sheet".
  * Upload the new sample sheet.
  * On the next page, click "requeue".
* Can multi-analysis by project split samples on the same plate?
  * No. All samples on the same plate must have the same project.

## Metrics and Bioinformatics FAQs

* What is DRC\_Level and why is this used?
  * DRC stands for Dynamic Range Compression. It is used to even out the SOMAmer concentrations from a 5-log dynamic range to 2-log. Without DRC, the most abundant SOMAmers would occupy the majority of sequencing readout capacity. Each sample would require extremely high sequencing depth to cover the unabundant SOMAmers and accurately quantify them. In using DRC, probes are grouped based on their observed abundances under no DRC condition, and each group is compressed by a different set ratio. The compression level of each SOMAmer is included in the `DRC_Level` row of SOMAmer metadata.
* How can I compare the outputs of DRAGEN Protein Quantification to SomaScan Array, and how do their metrics compare?
  * Unfortunately, it's not possible to directly compare the outputs of these two products due to the differing DRC strategies used.
  * Metrics and Normalization steps with different names between Somalogic's tool and Illumina's DRAGEN Protein Quantification. See the table below for the approximate mapping.

    | Somalogic Normalization Steps | Illumina Protein Prep Normalization Steps | Notes                                        |
    | ----------------------------- | ----------------------------------------- | -------------------------------------------- |
    | Raw RFU                       | Raw                                       |                                              |
    | Hyb Normalization             | HybNorm                                   |                                              |
    | medNormInt                    | MedNormInt                                |                                              |
    | plateScale                    | PlatformSpecificPlateScale                |                                              |
    | Calibration                   | PlatformSpecificCalibrate                 |                                              |
    | -                             | CrossPlatformPlateScale                   | Applied to normalize data across instruments |
    | -                             | CrossPlatformCalibrate                    | Applied to normalize data across instruments |
    | -                             | MedNormExt                                |                                              |
    | anmlQC                        | -                                         |                                              |
    | qcCheck                       | QCCheck                                   |                                              |
    | anmlSMP                       | -                                         |                                              |
    | Filtered                      | -                                         |                                              |
* What is Dilution and why is it used?
  * Each plate is split into multiple dilution groups, and SOMAmers are added to the plate based on that group. This is necessary because SOMAmers need to be more concentrated than proteins to facilitate binding. SOMAmers on catch0 beads have a physical concentration cap requiring the proteins to be diluted at different levels to achieve the concentration gaps to SOMAmers. The dilution group the SOMAmer is added to can be identified in the SOMAmer metadata.
* Do the counts in the ADAT represent the original absolute quantification of the sample?
  * No. Counts in the ADAT cover the relative SOMAmer quantification of the sample. Factors like DRC, Dilution, and PCR mean that counts are compressed by differing factors; these factors are not uncompressed in the final ADAT.
* How should signal and background be compared?
  * To identify background, blank samples on the plate (negative controls) can be used to obtain a per SOMAmer background.
    * Additionally, a global Limit of Detection (LoD) is computed for each SOMAmer for both plasma and serum using internal data. This is in the SOMAmer metadata section of the ADAT.
  * To identify signal, QC samples on the plate (positive controls) can be used to obtain a per SOMAmer signal.
  * When comparing signal to background, it's important to do so on a per SOMAmer level as background can vary from SOMAmer to SOMAmer.
* Where are the FASTQs?
  * DRAGEN Protein Quantification does not produce FASTQs. By utilizing DRAGEN Counting, the software uses BCL Convert to demultiplex *and* count proteins per sample at the same time. This reduces analysis time but also removes FASTQ as a pipeline output.
* What if the PF or Q30 values are lower than expected?
  * First, check secondary metrics. If secondary metrics are passing, it's acceptable to proceed with further analysis. If secondary metrics have warnings or failures, consider re-pooling or repeating the dilute and denature and sequencing of the saved pool. If the data still looks poor, reach out to tech support for further guidance.
* Why are there counts for blank samples, and non-human SOMAmer counts for human samples?
  * This is the background of the Illumina Protein Prep assay, which is due to non-specific binding of the SOMAmers. During library prep, all samples will experience SOMAMers sticking to beads or the side and bottom of the plate. These SOMAmers will then be brought through to sequencing. This is one of the reasons this is a relative abundance assay - the difference in SOMAmer counts is more important than the absolute value of the counts due to this background.
* Why are there SOMAmer counts in blank samples?
  * This is the background of the Illumina Protein Prep assay, which is due to non-specific binding of the SOMAmers. During library prep, all samples will experience SOMAMers sticking to beads or the side and bottom of the plate. These SOMAmers will then be brought through to sequencing.
* There are some proteins that multiple SOMAmers map to. Why is this and how does it impact analysis?
  * There are a handful of SOMAmers that map to the same protein. Some were created as Somalogic improved the SELEX process and a new SOMAmer’s affinity to the protein improved. Some may target different domains, isoforms, or cleavage products of a protein. They may or may not compete for the same epitope. Given the multiple reasons for multiple SOMAmers per protein, Illumina doesn't recommend any general methods for analysis that apply to all proteins.
* What organisms does the SOMAmer metadata cover in the 9.5k assay?

| Organism                    | # of SOMAmers in the 9.5 Assay |
| --------------------------- | ------------------------------ |
| Human                       | 10326                          |
| Mouse                       | 230                            |
| Gila monster                | 3                              |
| Hornet                      | 3                              |
| Jellyfish                   | 3                              |
| African clawed frog         | 3                              |
| Thermus thermophilus        | 3                              |
| European elder              | 2                              |
| Common eastern firefly      | 2                              |
| E. coli                     | 1                              |
| HIV-1                       | 1                              |
| HIV-2                       | 1                              |
| Bacillus stearothermophilus | 1                              |
| Ensifer meliloti            | 1                              |
| Red alga                    | 1                              |


# Known Limitations

**Known Limitations/Issues with DRAGEN Protein Quantification v2.3.0 Software:**

* Only supports 96 samples per PlateBarcode for a single analysis.
* SampleIDs must be unique prior to uploading IPPAS output files to BaseSpace Sequence Hub run planner tool.
* Unable to support AA probe pools runs (as noted by the last two values in the Probe Plate barcode. ie PP201013012345678-AA).
* DRAGEN Reports yield graph always displays 96 samples if possible; if there are fewer than 96 on a single plate, it will pull in samples from the next plate.

**Known Limitations/Issues with Local Sample Sheet Generation Tool v1.0.0:**

* Sample sheet must be saved as "CSV Comma-Delimited" and not "CSV UTF-8".
* Once saved, `LibraryPrepKits` under section `[Sequencing_Settings]` must be set to "IlluminaProteinPrepKit9.5k" for compatibility with the latest SW.


# Demo Data

Demo data for DRAGEN Protein Quantification can be found in the BSSH "Demo Data" tab under "Illumina Protein Prep". Import the Run to view a successful cloud analysis.


# Acronym Glossary

<table><thead><tr><th></th><th></th><th data-hidden></th></tr></thead><tbody><tr><td>IPP</td><td>Illumina Protein Prep</td><td></td></tr><tr><td>IPPAS</td><td>Illumina Protein Prep Automation System</td><td></td></tr><tr><td>ADAT</td><td>Counts file suffix</td><td></td></tr><tr><td>ICA</td><td>Illumina Connected Analytics</td><td></td></tr><tr><td>BSSH</td><td>BaseSpace Sequence Hub</td><td></td></tr><tr><td>ICM</td><td>Illumina Connected Multiomics</td><td></td></tr><tr><td>SOMAmer</td><td>Slow Off-rate Modified Aptamer</td><td></td></tr></tbody></table>


# Documentation Revision History

<table><thead><tr><th width="146.22216796875">Date</th><th>Changes Made</th></tr></thead><tbody><tr><td>March 2025</td><td>Initial Release</td></tr><tr><td>May 2025</td><td><ul><li>Multiomic Walkthrough Updated</li><li>ADAT Compatibility with Excel Added</li></ul></td></tr><tr><td>June 2025</td><td><p>Release with DRAGEN Protein Quantification v2.1.0</p><ul><li><p>Changes to ADAT Content</p><ul><li>Header: Addition of Median, Q75, Q90 Intra-Plate and Inter-Plate CVs</li><li>Header: Removal of PlateRefCorr_PassFlag</li><li>SOMAmer metadata clean up</li><li>SOMAmer metadata: Removal of PassFlags for Calibration and QC: not used for analysis</li><li>Sample metadata: EmpHybTemp_PassFlag added</li><li>Sample metadata: RefCorr_PassFlag removed</li><li>Sample metadata: MedNormInt_PassFlags for serum, plasma and QC samples, and MedNormExt_PassFlags from blank and calibrator samples removed</li></ul></li><li>Updated hybridization normalization to only use calibrator and QC samples</li><li>DRAGEN Application Manager support</li></ul></td></tr><tr><td>September 2025</td><td><ul><li>Updated 6k and 9.5k SeqID lists in multiomics walkthrough</li><li>Updated SW version to v2.2.2</li></ul></td></tr><tr><td>November 2025</td><td><ul><li>Added Known Limitations page</li></ul></td></tr><tr><td>February 2026</td><td><p>Release with DRAGEN Protein Quantification v2.3.0</p><ul><li>Changes to ADAT output directory structure and filenames</li><li>Introduction of PLATE_FAIL for sample-level RowCheck_Flag</li></ul></td></tr></tbody></table>


# Introduction

The DRAGEN Single Cell RNA application is a secondary analysis tool that can process multiplexed single cell RNA-Seq data in binary base call (BCL) files produced by NovaSeq 6000/6000Dx, NextSeq 1000/2000, and NovaSeq X Series sequencing systems to a cell-by-gene expression matrix.

You can perform secondary analysis in the cloud via BaseSpace Sequence Hub or Illumina BioInsight Platform Core (formerly ICA). When performing secondary analysis in the cloud, the analysis application launches automatically in BaseSpace Sequence Hub or Platform Core after the sequencing workflow completes.


# Prerequisites

* NovaSeq 6000/6000Dx, NextSeq 2000, or NovaSeq X Series
* Illumina Single Cell 3' RNA Prep Kit
  * T2: 1 library per sample, 1 FASTQ pair
  * T10: 1 library per sample, 1 FASTQ pair
  * T20: 1 library per sample, 1 FASTQ pair
  * T100: 4 libraries per sample, 1 FASTQ pair
  * M1: 8 libraries per sample, 8 FASTQ pairs
  * For more information about the library preparation kit, refer to the [Illumina Single Cell Prep Support Site](https://support.illumina.com/sequencing/sequencing_kits/illumina-single-cell-prep.html).
* A cloud account with a valid subscription. For information on registering your BaseSpace Sequence Hub or Bioinsight Platform Core account, refer to [Software Registration page](https://help.connected.illumina.com/account-management/rg-registration).


# Run Planning in BaseSpace Sequence Hub

The BaseSpace Sequence Hub Run Planning tool is used to generate a valid sample sheet in v2 format for use on a supported sequencer. Filling out the form on the user interface will produce a sample sheet with the required fields filled in that can be used to auto-launch a DRAGEN Single Cell RNA analysis. To create a sample sheet manually, see [Sample Sheet Requirements](/dragen-single-cell-rna/run-set-up-in-bssh/run-planning/sample-sheet-requirements) for details. Refer to [Cloud Analysis Auto-launch](https://help.connected.illumina.com/analysis/analysis_autolaunch) for more information about the Run Planning workflow and auto-launch.

{% hint style="info" %}
For NextSeq 2000, the samplesheet created by the Run Planning tool needs to be exported and uploaded to the instrument. When using the NovaSeq X, the samplesheet will automatically be on the instrument.
{% endhint %}

Use the steps below to create a DRAGEN Single Cell RNA run with the BaseSpace Run Planning tool. To get to the Run Planning tool, open BaseSpace Sequence Hub and navigate to the **Runs** page by using the navigation bar or by opening the menu on the left-hand side. From the **New Run** dropdown menu select **Run Planning**.

<figure><img src="/files/ekZspeTqYPHOVPCLfwRb" alt=""><figcaption></figcaption></figure>

## Step 1: Run Settings

<table><thead><tr><th width="206">Parameter Name</th><th width="123">Required?</th><th>Description</th></tr></thead><tbody><tr><td>Run Name</td><td>Required</td><td>Run Name can contain 255 alphanumeric characters, dashes, underscores, periods, and spaces; and must start with an alphanumeric, a dash or an underscore.</td></tr><tr><td>Run Description</td><td>Optional</td><td>Run Description can contain 8192 characters except square brackets, asterisks, and commas.</td></tr><tr><td>Instrument Platform</td><td>Required</td><td><p>Choose from DRAGEN Single Cell RNA software supported instruments:</p><ul><li>NovaSeq X Series</li><li>NovaSeq 6000/6000Dx</li><li>NextSeq 1000/2000</li></ul></td></tr><tr><td>Secondary Analysis</td><td>Required</td><td>Select BaseSpace / BioInsight Platform Core.</td></tr><tr><td>Read 1</td><td>Required only for Instrument Platform NovaSeq X Series</td><td>45 for DRAGEN Single Cell RNA analysis. May be different if running multiple applications in a single run.</td></tr><tr><td>Index 1</td><td>Required only for Instrument Platform NovaSeq X Series</td><td>10 for DRAGEN Single Cell RNA analysis. May be different if running multiple applications in a single run.</td></tr><tr><td>Index 2</td><td>Required only for Instrument Platform NovaSeq X Series</td><td>10 for DRAGEN Single Cell RNA analysis. May be different if running multiple applications in a single run.</td></tr><tr><td>Read 2</td><td>Required on this step only for runs on NovaSeq X Series</td><td>72 for DRAGEN Single Cell RNA analysis. May be different if running multiple applications in a single run.</td></tr><tr><td>Sample Container ID</td><td>Optional</td><td>Unique identifier for the container that holds the sample.</td></tr></tbody></table>

## Step 2: Configuration

{% hint style="info" %}
On NovaSeq X Series, this page is called "Configuration 1". The top right-hand corner of the UI displays the Read 1, Index 1, Index 2 and Read 2 entered on the previous run settings screen.
{% endhint %}

<table><thead><tr><th width="193">Parameter Name</th><th width="129">Required?</th><th>Description</th></tr></thead><tbody><tr><td>Application</td><td>Required</td><td>Select "DRAGEN Single Cell RNA - 4.5.0" or the latest version available</td></tr><tr><td>Description</td><td>Optional</td><td>Optional Text Field</td></tr><tr><td>Library Prep Kit</td><td>Required</td><td>Select Illumina Single Cell 3’ RNA Prep or Illumina Single Cell CRISPR Prep</td></tr><tr><td>Index Adapter Kit</td><td>Required</td><td><p>Select a supported index adapter kit:</p><ul><li>Illumina Single Cell UD 8 Indexes</li><li>Illumina Single Cell UD Indexes Set A</li></ul></td></tr><tr><td>Reference Genome</td><td>Required on this step only for runs on NovaSeq X Series</td><td><p>Select the appropriate genome reference for the sample type.</p><p>When selecting human, it is recommended to use linear references for RNA analysis. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support">DRAGEN Reference Support</a> for more information.</p></td></tr></tbody></table>

## Step 3: Run Configuration and Analysis Settings

The run configuration and analysis settings options differ slightly depending on the instrument platform chosen for the run in step 1. The below subsections describe the fields available for each option.

{% tabs %}
{% tab title="NovaSeq X" %}

<table><thead><tr><th width="202">Parameter Name</th><th width="122">Required?</th><th>Description</th></tr></thead><tbody><tr><td>Description</td><td>Optional</td><td>Optional Text Field</td></tr><tr><td>Library Prep Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Index Adapter Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Reference Genome</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>RNA Annotation File</td><td>Optional</td><td><p>For custom references, use this field to select the corresponding GTF file to use for annotation.</p><p>For built in references, use this field to override default annotations. The following list shows the default GTFs being used for annotation.</p><ul><li><p>GENCODE v19</p><ul><li>Homo sapiens [UCSC] hg19 v5</li><li>Homo sapiens [UCSC] hg19 v5 Pangenome</li><li>Homo sapiens [NCBI] hs37d5 v5</li><li>Homo sapiens [NCBI] hs37d5 v5 Pangenome</li></ul></li><li><p>GENCODE v44</p><ul><li>Homo sapiens [1000 Genomes] hg38 v5</li><li>Homo sapiens [1000 Genomes] hg38 v5 Pangenome</li></ul></li><li><p>GENCODE vM23</p><ul><li>Mus musculus [UCSC] mm10</li></ul></li><li><p>ENSEMBL 98</p><ul><li>Rattus norvegicus [UCSC] rn6</li></ul></li></ul></td></tr><tr><td>Feature Barcode Reference</td><td>Required for Illumina Single Cell CRISPR Library Prep</td><td>Specify a CSV feature reference file that contains feature barcode information as specified in the <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#inputs">DRAGEN documentation</a>.</td></tr><tr><td>Custom Adapters</td><td>Optional</td><td>Select if custom adapter reads will be specified instead of those in the index adapter kit.</td></tr><tr><td>Adapter Read 1</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Adapter Read 2</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Override Cycles</td><td>Required</td><td>Defaults to U45;I10;I10;Y72. May be different if running multiple applications in a single run.</td></tr><tr><td>Lane Usage</td><td>Optional</td><td>Select the checkbox if samples are loaded in all lanes. If selected, the generated sample sheet will not contain the Lane column.</td></tr><tr><td>Sample Table</td><td>Required</td><td><p>The sample table should be filled out based on how the sample will be prepared based on the library preparation kit used. See <a href="https://support.illumina.com/sequencing/sequencing_kits/illumina-single-cell-prep.html">Illumina Single Cell 3' RNA Prep Documentation</a> for more information. The following fields are included in the table:</p><ul><li>Sample Name</li><li>Expression Lanes</li><li>Expression Index ID</li><li>Expression Barcode Mismatches Index 1 - the allowed number of index read 1 mismatches. The default is 1.</li><li>Expression Barcode Mismatches Index 2 - the allowed number of index read 2 mismatches. The default is 1.</li><li>Feature Lanes*</li><li>Feature Index ID*</li><li>Feature Barcode Mismatches Index 1* - the allowed number of index read 1 mismatches. The default is 1.</li><li>Feature Barcode Mismatches Index 2* - the allowed number of index read 1 mismatches. The default is 1.</li><li>Thresholding Method - specify the method for determining the count threshold value. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other#cell-filtering">DRAGEN documentation</a> for more details.</li><li>Expected Number of Cells</li><li>Project - used to specify the associated BaseSpace Project to output data to. If left empty, Project will default to the Project name derived from the Experiment/Run name.</li></ul><p>*Feature fields are required for Illumina Single Cell CRISPR Library Prep</p></td></tr><tr><td>Configuration Type</td><td>Optional</td><td>Specify either Illumina Single Cell 3’ RNA or Illumina Single Cell CRISPR</td></tr><tr><td>Barcode Read</td><td>Required</td><td>Defaults to Read 1</td></tr><tr><td>RNA Library Type</td><td>Required</td><td>Defaults to Stranded Forward</td></tr><tr><td>Barcode Sequence File</td><td>Optional</td><td>Not required for Illumina Single Cell Prep Kits. Specify a file containing valid cell barcode sequences. Maps to --single-cell-barcode-sequence-whitelist in command line arguments.</td></tr><tr><td>Barcode Position</td><td>Required</td><td>Defaults to 0_7+11_16+20_25+31_38</td></tr><tr><td>UMI/BI Position</td><td>Required</td><td>Defaults to 39_41</td></tr></tbody></table>
{% endtab %}

{% tab title="NovaSeq 6000" %}

<table><thead><tr><th width="202">Parameter Name</th><th width="122">Required?</th><th>Description</th></tr></thead><tbody><tr><td>Description</td><td>Optional</td><td>Optional Text Field</td></tr><tr><td>Library Prep Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Index Adapter Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Reference Genome</td><td>Required</td><td><p>Select the appropriate genome reference for the sample type.</p><p>When selecting human, it is recommended to use linear references for RNA analysis. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support">DRAGEN Reference Support</a> for more information.</p></td></tr><tr><td>RNA Annotation File</td><td>Optional</td><td><p>For custom references, use this field to select the corresponding GTF file to use for annotation.</p><p>For built in references, use this field to override default annotations. The following list shows the default GTFs being used for annotation.</p><ul><li><p>GENCODE v19</p><ul><li>Homo sapiens [UCSC] hg19 v5</li><li>Homo sapiens [UCSC] hg19 v5 Pangenome</li><li>Homo sapiens [NCBI] hs37d5 v5</li><li>Homo sapiens [NCBI] hs37d5 v5 Pangenome</li></ul></li><li><p>GENCODE v44</p><ul><li>Homo sapiens [1000 Genomes] hg38 v5</li><li>Homo sapiens [1000 Genomes] hg38 v5 Pangenome</li></ul></li><li><p>GENCODE vM23</p><ul><li>Mus musculus [UCSC] mm10</li></ul></li><li><p>ENSEMBL 98</p><ul><li>Rattus norvegicus [UCSC] rn6</li></ul></li></ul></td></tr><tr><td>Feature Barcode Reference</td><td>Required for Illumina Single Cell CRISPR Library Prep</td><td>Specify a CSV feature reference file that contains feature barcode information as specified in the <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#inputs">DRAGEN documentation</a>.</td></tr><tr><td>Custom Adapters</td><td>Optional</td><td>Select if custom adapter reads will be specified instead of those in the index adapter kit.</td></tr><tr><td>Adapter Read 1</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Adapter Read 2</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Index Reads</td><td>Required</td><td>Defaults to 2 indexes</td></tr><tr><td>Read Type</td><td>Required</td><td>Defaults to Paired End</td></tr><tr><td>Read Lengths</td><td>Required</td><td><p>Defaults to 45:10:10:72. May be different if running multiple applications in a single run. The default is compatible with 150 cycle SBS kits. If using a larger kit, the Read 2 cycle information can be increased.</p><p>There are diminishing returns for increased read lengths as the insert will read through the cDNA sequence into the poly-A region with longer read lengths.</p><p>Read 1 should not be updated as it contains the cell barcode and binning index. Longer read lengths will need to be trimmed. Shorter read lengths will impact cell barcode identification</p></td></tr><tr><td>Lane Usage</td><td>Optional</td><td>Select the checkbox if samples are loaded in all lanes. If selected, the generated sample sheet will not contain the Lane column.</td></tr><tr><td>Sample Table</td><td>Required</td><td><p>The sample table should be filled out based on how the sample will be prepared based on the library preparation kit used. See <a href="https://support.illumina.com/sequencing/sequencing_kits/illumina-single-cell-prep.html">Illumina Single Cell 3' RNA Prep Documentation</a> for more information. The following fields are included in the table:</p><ul><li>Sample Name</li><li>Expression Lanes</li><li>Expression Index ID</li><li>Feature Lanes*</li><li>Feature Index ID*</li><li>Thresholding Method - specify the method for determining the count threshold value. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other#cell-filtering">DRAGEN documentation</a> for more details.</li><li>Expected Number of Cells</li><li>Project - used to specify the associated BaseSpace Project to output data to. If left empty, Project will default to the Project name derived from the Experiment/Run name.</li></ul><p>*Feature fields are required for Illumina Single Cell CRISPR Library Prep</p></td></tr><tr><td>Barcode Mismatches Index 1</td><td>Required</td><td>The allowed number of index read 1 mismatches. The default is 1 and the maximum value is 2.</td></tr><tr><td>Barcode Mismatches Index 2</td><td>Required</td><td>The allowed number of index read 2 mismatches. The default is 1 and the maximum value is 2.</td></tr><tr><td>Override Cycles</td><td>Required</td><td>Defaults to U45;I10;I10;Y72. May be different if running multiple applications in a single run.</td></tr><tr><td>Configuration Type</td><td>Optional</td><td>Specify either Illumina Single Cell 3’ RNA or Illumina Single Cell CRISPR</td></tr><tr><td>Barcode Read</td><td>Required</td><td>Defaults to Read 1</td></tr><tr><td>RNA Library Type</td><td>Required</td><td>Defaults to Stranded Forward</td></tr><tr><td>Barcode Sequence File</td><td>Optional</td><td>Specify a file containing valid cell barcode sequences. Maps to --single-cell-barcode-sequence-whitelist in command line arguments. Not required for Illumina Single Cell 3' RNA Prep Kits.</td></tr><tr><td>Barcode Position</td><td>Required</td><td>Defaults to 0_7+11_16+20_25+31_38</td></tr><tr><td>UMI/BI Position</td><td>Required</td><td>Defaults to 39_41</td></tr></tbody></table>
{% endtab %}

{% tab title="NextSeq 1000/2000" %}

<table><thead><tr><th width="202">Parameter Name</th><th width="122">Required?</th><th>Description</th></tr></thead><tbody><tr><td>Description</td><td>Optional</td><td>Optional Text Field</td></tr><tr><td>Library Prep Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Index Adapter Kit</td><td>Required</td><td>Auto-populated from previous step</td></tr><tr><td>Reference Genome</td><td>Required</td><td><p>Select the appropriate genome reference for the sample type.</p><p>When selecting human, it is recommended to use linear references for RNA analysis. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support">DRAGEN Reference Support</a> for more information.</p></td></tr><tr><td>RNA Annotation File</td><td>Optional</td><td><p>For custom references, use this field to select the corresponding GTF file to use for annotation.</p><p>For built in references, use this field to override default annotations. The following list shows the default GTFs being used for annotation.</p><ul><li><p>GENCODE v19</p><ul><li>Homo sapiens [UCSC] hg19 v5</li><li>Homo sapiens [UCSC] hg19 v5 Pangenome</li><li>Homo sapiens [NCBI] hs37d5 v5</li><li>Homo sapiens [NCBI] hs37d5 v5 Pangenome</li></ul></li><li><p>GENCODE v44</p><ul><li>Homo sapiens [1000 Genomes] hg38 v5</li><li>Homo sapiens [1000 Genomes] hg38 v5 Pangenome</li></ul></li><li><p>GENCODE vM23</p><ul><li>Mus musculus [UCSC] mm10</li></ul></li><li><p>ENSEMBL 98</p><ul><li>Rattus norvegicus [UCSC] rn6</li></ul></li></ul></td></tr><tr><td>Feature Barcode Reference</td><td>Required for Illumina Single Cell CRISPR Library Prep</td><td>Specify a CSV feature reference file that contains feature barcode information as specified in the <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#inputs">DRAGEN documentation</a>.</td></tr><tr><td>Custom Adapters</td><td>Optional</td><td>Select if custom adapter reads will be specified instead of those in the index adapter kit.</td></tr><tr><td>Adapter Read 1</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Adapter Read 2</td><td>Optional</td><td>Use this field to specify custom adapter reads.</td></tr><tr><td>Index Reads</td><td>Required</td><td>Defaults to 2 indexes</td></tr><tr><td>Read Type</td><td>Required</td><td>Defaults to Paired End</td></tr><tr><td>Read Lengths</td><td>Required</td><td><p>Defaults to 45:10:10:72. May be different if running multiple applications in a single run. The default is compatible with 150 cycle SBS kits. If using a larger kit, the Read 2 cycle information can be increased.</p><p>There are diminishing returns for increased read lengths as the insert will read through the cDNA sequence into the poly-A region with longer read lengths.</p><p>Read 1 should not be updated as it contains the cell barcode and binning index. Longer read lengths will need to be trimmed. Shorter read lengths will impact cell barcode identification</p></td></tr><tr><td>Sample Table</td><td>Required</td><td><p>The sample table should be filled out based on how the sample will be prepared based on the library preparation kit used. See <a href="https://support.illumina.com/sequencing/sequencing_kits/illumina-single-cell-prep.html">Illumina Single Cell 3' RNA Prep Documentation</a> for more information. The following fields are included in the table:</p><ul><li>Sample Name</li><li>Expression Index ID</li><li>Feature Index ID*</li><li>Thresholding Method - specify the method for determining the count threshold value. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other#cell-filtering">DRAGEN documentation</a> for more details.</li><li>Expected Number of Cells</li><li>Project - used to specify the associated BaseSpace Project to output data to. If left empty, Project will default to the Project name derived from the Experiment/Run name.</li></ul><p>*Feature fields are required for Illumina Single Cell CRISPR Library Prep</p></td></tr><tr><td>Barcode Mismatches Index 1</td><td>Required</td><td>The allowed number of index read 1 mismatches. The default is 1 and the maximum value is 2.</td></tr><tr><td>Barcode Mismatches Index 2</td><td>Required</td><td>The allowed number of index read 2 mismatches. The default is 1 and the maximum value is 2.</td></tr><tr><td>Override Cycles</td><td>Required</td><td>Defaults to U45;I10;I10;Y72. May be different if running multiple applications in a single run.</td></tr><tr><td>Configuration Type</td><td>Optional</td><td>Specify either Illumina Single Cell 3’ RNA or Illumina Single Cell CRISPR</td></tr><tr><td>Barcode Read</td><td>Required</td><td>Defaults to Read 1</td></tr><tr><td>RNA Library Type</td><td>Required</td><td>Defaults to Stranded Forward</td></tr><tr><td>Barcode Sequence File</td><td>Optional</td><td>Not required for Illumina Single Cell 3' RNA Prep Kits. Specify a file containing valid cell barcode sequences. Maps to --single-cell-barcode-sequence-whitelist in command line arguments.</td></tr><tr><td>Barcode Position</td><td>Required</td><td>Defaults to 0_7+11_16+20_25+31_38</td></tr><tr><td>UMI/BI Position</td><td>Required</td><td>Defaults to 39_41</td></tr></tbody></table>
{% endtab %}
{% endtabs %}

## Step 4: Run Review

Once all details are captured and pass validation, review the run information and choose the **Edit** option to correct any information.

For NovaSeq 6000/6000Dx, **Export** the sample sheet to be uploaded to the instrument.

For NovaSeq X Series and NextSeq 1000/2000, the run can be saved as a draft or as a planned run (via **Save as Draf**t and **Save as Planned** buttons respectively). Either selection will save the run to the Planned Runs screen on BaseSpace. Once a run is saved as Planned, it will appear on the instrument where it can be selected for sequencing.

{% hint style="info" %}
The sample sheet for Planned runs can be downloaded by selecting the planned run and **File** -> **Download** -> **SampleSheet**.
{% endhint %}

For more information about the auto-launch, refer to [Cloud Analysis Auto-launch](https://help.connected.illumina.com/analysis/analysis_autolaunch). For additional information on run planning, refer to [Plan Runs on Basespace Sequence Hub](https://help.connected.illumina.com/basespace-sequence-hub/sequence/plan-runs).


# Sample Sheet Requirements

The DRAGEN Single Cell RNA software has optional and required fields in addition to general sample sheet requirements. Below is a description of the fields in each section.

{% hint style="info" %}
The preferred method for creating sample sheets is to use [Run Planning in BaseSpace Sequence Hub](/dragen-single-cell-rna/run-set-up-in-bssh/run-planning)
{% endhint %}

## \[Sequencing\_Settings]

<table><thead><tr><th width="186">Parameter</th><th width="132">Required?</th><th>Details</th></tr></thead><tbody><tr><td>LibraryPrepKits</td><td>Required</td><td>Accepted values are: IlluminaSingleCell3RNAPrep</td></tr></tbody></table>

## \[BCLConvert\_Settings]

<table><thead><tr><th width="195">Parameter</th><th width="129">Required?</th><th>Details</th></tr></thead><tbody><tr><td>SoftwareVersion</td><td>Required</td><td>The DRAGEN component software version. DRAGEN Single Cell RNA software requires 4.4.0.</td></tr><tr><td>NoLaneSplitting</td><td>Required</td><td>TRUE for DRAGEN Single Cell RNA software</td></tr><tr><td>TrimUMI</td><td>Required</td><td>0 for DRAGEN Single Cell RNA software</td></tr><tr><td>OverrideCycles</td><td>Required</td><td>U45;I10;I10;Y72 for DRAGEN Single Cell RNA software. May be different if running multiple applications in a single run.</td></tr><tr><td>FastqCompressionFormat</td><td>Required</td><td>gzip</td></tr></tbody></table>

## \[BCLConvert\_Data]

<table><thead><tr><th width="220">Parameter</th><th width="151">Required?</th><th>Details</th></tr></thead><tbody><tr><td>Sample_ID</td><td>Required</td><td>Must match a Sample_ID listed in the [Cloud_DragenSingleCellRna_Data] and [Cloud_Data] section.</td></tr><tr><td>Index</td><td>Required</td><td>Index 1 sequence</td></tr><tr><td>Index2</td><td>Required</td><td>Index 2 sequence</td></tr><tr><td>Lane</td><td>Only for NovaSeq 6000/6000 Dx workflow</td><td>Indicates which lane corresponds to a given sample. Enter a single numeric value per row. Cannot be empty, i.e. the analysis fails if the Lane column is present without a value in each row.</td></tr></tbody></table>

## \[Cloud\_DragenSingleCellRna\_Settings]

<table><thead><tr><th width="207">Parameter</th><th width="137">Required?</th><th>Details</th></tr></thead><tbody><tr><td>SoftwareVersion</td><td>Required</td><td>The DRAGEN component software version. DRAGEN Single Cell RNA software requires 4.4.0.</td></tr><tr><td>EnablePipseqMode</td><td>Required</td><td>TRUE for Illumina Single Cell 3’ RNA kit. Maps to --scrna-enable-pipseq-mode in command line arguments.</td></tr><tr><td>ReferenceGenomeDir</td><td>Required</td><td>Location of reference genome TAR containing a DRAGEN hash table and optionally a GTF.</td></tr><tr><td>BarcodeRead</td><td>Required</td><td>Read1 for Illumina Single Cell 3’ RNA kit</td></tr><tr><td>RnaLibraryType</td><td>Required</td><td>SF for Illumina Single Cell 3’ RNA kit (stranded forward). Maps to --rna-library-type in command line arguments.</td></tr><tr><td>BarcodePosition</td><td>Required</td><td>0_7+11_16+20_25+31_38 for Illumina Single Cell 3’ RNA kit. Maps to --scrna-barcode-position in command line arguments.</td></tr><tr><td>UmiPosition</td><td>Required</td><td>39_41 for Illumina Single Cell 3’ RNA kit. Maps to --scrna-umi-position in command line arguments.</td></tr></tbody></table>

## \[Cloud\_DragenSingleCellRna\_Data]

<table><thead><tr><th width="213">Parameter</th><th width="135">Required?</th><th>Details</th></tr></thead><tbody><tr><td>Sample_ID</td><td>Required</td><td>Must match a Sample_ID listed in the [BCLConvert_Data] and [Cloud_Data] section.</td></tr></tbody></table>

## \[Cloud\_Settings] for Auto-launch

<table><thead><tr><th width="225">Parameter</th><th width="128">Required?</th><th>Details</th></tr></thead><tbody><tr><td>GeneratedVersion</td><td>Not Required</td><td>The cloud version used to create the sample sheet. Optional if manually updating a sample sheet. (ex: 1.17.0.202411192008).</td></tr><tr><td>Cloud_Workflow</td><td>Not Required</td><td>ica_workflow_1</td></tr><tr><td>BCLConvert_Pipeline</td><td>Required</td><td><p>The value is a universal record number (URN). The valid value is:</p><p>urn:ilmn:ica:pipeline:730df76f-715a-45bf-9500-e6e0ce1ab224#BclConvert_v4_3_13</p></td></tr><tr><td>Cloud_DragenSingleCellRna_Pipeline</td><td>Required</td><td><p>The value is a URN in the following format:</p><p>urn:ilmn:ica:pipeline:b3c5ab5f-2853-4873-93c4-61a807f844a7#DRAGEN_Single_Cell_RNA_4-4-2_-_Sequencer_Integration_Only</p></td></tr></tbody></table>

## \[Cloud\_Data] for Auto-Launch

<table><thead><tr><th width="229">Parameter</th><th width="128">Required?</th><th>Details</th></tr></thead><tbody><tr><td>Sample_ID</td><td>Required</td><td>Must match a Sample_ID listed in the [BCLConvert_Data] and [Cloud_DragenSingleCellRna_Data] section.</td></tr><tr><td>ProjectName</td><td>Not Required</td><td>The BaseSpace Sequence Hub project name</td></tr><tr><td>LibraryName</td><td>Not Required</td><td>Combination of sample ID and index values in the following format: sampleID_Index_Index2.</td></tr><tr><td>LibraryPrepKit</td><td>Required</td><td>The Library Prep Kit used</td></tr><tr><td>IndexAdapterKitName</td><td>Required</td><td>The Index Adapter Kit used</td></tr></tbody></table>


# Manual Launch on BaseSpace

The DRAGEN Single Cell RNA analysis can be manually launched to analyze previously generated FASTQ files by using a BaseSpace App.

Use the steps below to create a manually launch the DRAGEN Single Cell RNA app in BaseSpace. To get to the app, open BaseSpace Sequence Hub and navigate to the **Apps** page by using the navigation bar or by opening the menu on the left-hand side. Select or search for the **DRAGEN Single Cell RNA** app from the list of available apps. Select **Launch Application** to provide details for your analysis. Detailed steps are provided below.

### Select Input Data

The DRAGEN Single Cell RNA app only supports Biosample inputs. For more information on Biosamples refer to the [BaseSpace Data Model](https://help.connected.illumina.com/basespace-sequence-hub/overview/data-model).

### Configuration

| Parameter Name         | Required? | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ---------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Analysis Name          | Required  | Name of the analysis                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| Save Results To        | Required  | Select the project that will store the analysis results.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| Library Kit            | Required  | Select either Illumina Single Cell 3' RNA or Illumina Single Cell CRISPR                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| Input Type             | Required  | Select "Expression input only" for samples with only gene expression libraries. Select "Paired expression and feature inputs" for samples with gene expression and feature inputs. Depending on the option chosen, the biosamples section below will disable either the "Select Biosample(s)" option or the table with the option to pair expression and feature biosamples.                                                                                                                                                                                                                                                                                                                                                               |
| Biosample(s)           | Required  | <p>Depending on the Input Type chosen browse for and select either of the following:</p><ul><li>A set of gene expression biosamples</li><li>A set of paired gene expression and feature biosamples. Use the Sample Name field to designate the name of the paired sample containing both gene expression and feature information</li></ul>                                                                                                                                                                                                                                                                                                                                                                                                 |
| Map/Align Output       | Required  | Select whether to output the alignments in BAM or CRAM format. The default is to not output an alignments file, which decreases compute time.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| Barcode/MI Source      | Required  | <p>Select the appropriate setting that matches how FASTQ files were generated.</p><ul><li>FASTQ Header – the FASTQ files were generated with the OverrideCycles sample sheet setting writing the R1 sequence to the FASTQ header</li><li>Barcode/UMI Read - the Read 1 FASTQ files were created without setting OverrideCycles in the sample sheet so the Read 1 FASTQ file contains the full sequencing read.</li></ul>                                                                                                                                                                                                                                                                                                                   |
| Reference              | Required  | Select the reference genome to use in the analysis. The app provides support for common human, mouse, and rat genomes in addition to supporting custom references built by the DRAGEN Reference Builder app. When selecting human, it is recommended to use linear references for RNA analysis. See [DRAGEN Reference Support](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support) for more information.                                                                                                                                                                                                                                                                                     |
| Custom Reference Files | Optional  | <p>Custom references can be generated from a FASTA file and optionally a GTF file with the DRAGEN Reference Builder app. For more information, refer to <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support/prepare-a-reference-genome">Prepare a Reference Genome</a>.</p><ul><li>Ensure "Include RNA Data in Reference" is enabled</li></ul>                                                                                                                                                                                                                                                                                                                                       |
| Gene Annotation File   | Optional  | <p>For custom references, select the corresponding GTF file to use. For built in references, the following list shows the default GTFs being used. This can be overridden for custom annotations by using this field.</p><ul><li><p>GENCODE v19</p><ul><li>Homo sapiens \[UCSC] hg19 v5</li><li>Homo sapiens \[UCSC] hg19 v5 Pangenome</li><li>Homo sapiens \[NCBI] hs37d5 v5</li><li>Homo sapiens \[NCBI] hs37d5 v5 Pangenome</li></ul></li><li><p>GENCODE v44</p><ul><li>Homo sapiens \[1000 Genomes] hg38 v5</li><li>Homo sapiens \[1000 Genomes] hg38 v5 Pangenome</li></ul></li><li><p>GENCODE vM23</p><ul><li>Mus musculus \[UCSC] mm10</li></ul></li><li><p>ENSEMBL 98</p><ul><li>Rattus norvegicus \[UCSC] rn6</li></ul></li></ul> |

### Library Kit Configuration

| Parameter Name             | Required? | Description                                                                                                                                                                                 |
| -------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Barcode Position           | Required  | Defaults to 0\_7+11\_16+20\_25+31\_38 for Illumina Single Cell Prep Kits.                                                                                                                   |
| UMI/BI Position            | Required  | Defaults to 39\_41 for Illumina Single Cell 3' RNA Prep Kits.                                                                                                                               |
| Barcode/MI Read            | Required  | Defaults to Read 1 for Illumina Single Cell 3' RNA Prep Kits.                                                                                                                               |
| Barcode Sequence List File | Optional  | Specify a file containing valid cell barcode sequences. Maps to --single-cell-barcode-sequence-whitelist in command line arguments. Not required for Illumina Single Cell 3' RNA Prep Kits. |
| RNA Library Type           | Required  | Auto-populated with forward for Illumina Single Cell 3' RNA Prep Kits.                                                                                                                      |

### Cell Hashing and Feature Counting

| Parameter Name                    | Required? | Description                                                                                                                                                                                                                                                                                                                             |
| --------------------------------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Cell Hashing and Feature Counting | Optional  | Use the checkboxes to enable cell hashing and feature counting using feature barcode UMI.                                                                                                                                                                                                                                               |
| Feature Barcode UMI Position      | Optional  | Feature barcode UMI position is in the format of \<start index>\_\<end index>. ex: 11\_18 specifies an 8 bp sequence from positions 11 to 18 (inclusive). The first position is 0.                                                                                                                                                      |
| Cell Hashing Reference            | Optional  | Specify a CSV or FASTA cell-hashing reference file that contains sample-specific oligo-tags. Maps to --single-cell-cell-hashing-reference in command line arguments.                                                                                                                                                                    |
| Detect Doublets                   | Optional  | Select the checkbox to enable doublet detection in cell-hashing sample demultiplexing. Maps to --single-cell-demux-detect-doublets in command line arguments.                                                                                                                                                                           |
| Feature Barcode Reference         | Optional  | Specify a CSV feature reference file that contains feature barcode information as specified in the [DRAGEN documentation](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#inputs). Maps to --single-cell-feature-barcode-reference in command line arguments. |

### Demultiplexing

| Parameter Name        | Required? | Description                                                                                                                           |
| --------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| Demultiplexing Method | Optional  | Select genotype-based or genotype-free sample demultiplexing.                                                                         |
| Sample VCF            | Optional  | Specify a VCF file for genotype-based demultiplexing. Maps to --single-cell-demux-sample-vcf in command line arguments.               |
| Reference VCF         | Optional  | Specify a VCF file for genotype-free demultiplexing. Maps to --single-cell-demux-reference-vcf in command line arguments.             |
| Number of Samples     | Optional  | Specify the number of samples for genotype-free demultiplexing. Maps to --single-cell-demux-number-samples in command line arguments. |
| Detect Doublets       | Optional  | Enable doublet detection in sample demultiplexing. Maps to --single-cell-demux-detect-doublets in command line arguments.             |

### Advanced Settings

| Parameter Name           | Required? | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| ------------------------ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Poly-A Trimming          | Optional  | Disabled for Illumina Single Cell 3' RNA Prep Kits.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| Expected Number of Cells | Optional  | Specify the expected number of cells. The DRAGEN default of 400 is used if not set. Adjust only if the expected number of cells is so far from the default that DRAGEN does not call the correct cell filtering threshold automatically.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| Thresholding Method      | Optional  | <p>Specify the method for determining the count threshold value.</p><ul><li>Ratio: DRAGEN estimates the count threshold as max(Te, Tm). Tm is 10% of the count seen in the cell at the 10th percentile of the expected cells. Te is 50% of the count seen in the least abundant expected cell.</li><li>Inflection: DRAGEN estimates the count threshold by analyzing inflection points in the cumulative distribution of counts.</li><li>Fixed: The count threshold is set to force the expected number of cells.</li></ul><p>Maps to --single-cell-threshold in command line arguments. See <a href="https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other#cell-filtering">DRAGEN documentation</a> for more details.</p> |
| Scratch Size Per Node    | Optional  | Scratch size for DRAGEN process. Defaults to "8 TiB".                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |

### Additional Arguments

Use the Additional Arguments section to acknowledge disclaimer prior to defining any custom settings. Below are some commonly used additional arguments.

| Argument                               | Description                                                                                                                                                                                                                                        |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --annotation-file-ignore-biotypes=none | When selecting the Illumina Single Cell 3’ RNA Library Prep Kit, the pipeline will automatically ignore pseudogenes, shortRNA, and rRNA biotypes during mapping. This behavior can be disabled by adding "--annotation-file-ignore-biotypes=none". |

### Launch Application

Accept the BaseSpace Labs disclaimer and **Launch Application** to begin your analysis.


# Installation and Setup

DRAGEN Single Cell RNA can be run on an OnPrem HPC using the DRAGEN in SW-mode and an Illumina Connected API Key. When running in this mode, the STAR aligner will be used instead of the DRAGEN RNA aligner.

### Prerequisites

1. Access to high-performance compute with an EL8 or EL9 based operating system.
2. Your own domain which has an BioInsight Platform (previously ICA) Subscription, including the free version - Platform Core Basic. For more information on subscription setup, refer to the [BioInsight Platform Setup Guide](https://help.connected.illumina.com/account-management/rg-registration).

### Generating an Illumina Connected API Key

To generate an Illumina API key, log in through the [Illumina login](https://login.illumina.com/login) to access Illumina BioInsight Platform. The "API Keys" page in the left navigation menu allows you to view and manage you API keys.

To create a new API key, click the "Generate" button. Provide a name for the key, then choose to either include all workgroups or select specific workgroups that the key should have access to. Copy this generated key for inclusion in the license credential file. For more information on managing and generating API keys refer to the [Platform Home support documentation](https://help.connected.illumina.com/account-management/platform-home#api-keys).

A license credential file is used by DRAGEN in runtime to perform credential based authentication for users running DRAGEN SW-mode. DRAGEN must have access to the DRAGEN license server at runtime. To verify connectivity to the license server, you can query the health check endpoint which will return a 200 status code and a small JSON body if successful.

```
curl https://license.dragen.illumina.com/healthcheck/version --header 'Content-Type: application/json'
{"version":"<server version>"}
```

Use the generated API key to create the license credential file with the following format:

```
credentials-1=IlluminaPlatform
credentials-2=<user api key>
```

Pass the License Credential file to DRAGEN at runtime using the `--lic-credentials` parameter as specified.

###


# Pre-Trimming With Cutadapt when using the STAR Mapper

### Recommended pre-trimming for DRAGEN v4.5 (DRAGEN-STAR workflow)

The standard DRAGEN hardware mode (run with FPGA acceleration) trims reads automatically. However, when processing scRNA-seq data with DRAGEN v4.5 software mode, using the DRAGEN-STAR workflow, trimming is not performed. We recommend pre-trimming FASTQs to remove TSO, polyA, and polyG sequences that can negatively impact alignment (leading to a variable drop in mapping rates, depending on the sample type and read 2 length). Note: *Trimming should only be performed on gene expression libraries.*\
TSO sequences are typically observed in <1–5% of reads from Illumina Single-Cell Prep Libraries and are more common in shorter, unfragmented library molecules.\
PolyA sequences occur when reads extend into the polyA tail or reverse-complement polyT capture sequence.\
PolyG sequences arise from reads containing short inserts, which generate dark cycles after processing past the end of the library molecule. Dark cycles are as G bases on Illumina instruments.\
Manual pre-trimming is recommended until support for these artifacts is incorporated in a future DRAGEN release (v4.6).\
The workflow below produces trimmed FASTQ pairs while trimming only Read 2, allowing the output files to be used directly with the DRAGEN-STAR pipeline.The standard DRAGEN hardware mode (run with FPGA acceleration) trims reads automatically. However, when processing scRNA-seq data with DRAGEN v4.5 software mode, using the DRAGEN-STAR workflow, trimming is not performed. We recommend pre-trimming FASTQs to remove TSO, polyA, and polyG sequences that can negatively impact alignment and mapping rates depending on the sample type and read 2 length.

* **TSO** sequences are typically observed in <1–5% of reads from Illumina Single-Cell Prep Libraries and are more common in shorter, unfragmented library molecules.
* **PolyA** sequences occur when reads extend into the polyA tail or reverse-complement polyT capture sequence.
* **PolyG** sequences arise from reads containing short inserts, which generate dark cycles after processing past the end of the library molecule. Dark cycles are as G bases on Illumina instruments.

Manual pre-trimming is recommended until support for these artifacts is incorporated in a future DRAGEN release (v4.6).

The workflow below produces trimmed FASTQ pairs while trimming only Read 2, allowing the output files to be used directly with the DRAGEN-STAR pipeline.

For more information about using the STAR Mapper with DRAGEN, see the following: <https://developer.illumina.com/news-updates/your-pipseq-workflow-is-consolidating-into-dragen>

#### 1. Install Cutadapt

<pre><code><strong># installation
</strong><strong>conda create -n cutadapt -c conda-forge -c bioconda cutadapt
</strong><strong># Activate it
</strong><strong>conda activate cutadapt
</strong><strong># verify installation
</strong><strong>cutadapt --version
</strong></code></pre>

#### 2. Set up the input + output paths and run the cutadapt command

```
r1="/path/to/r1.fastq.gz"
r2="/path/to/r2.fastq.gz"
sample_id="sample_name"

cutadapt \
  --cores 8 \
  --no-indels \
  --pair-filter=any \
  -e 0 \
  -O 12 \
  -U 1 \
  -G AGAGTGAATGGG \
  -G TCAACGCAGAGT \
  -A AAAAAAAAAAAA \
  -A "GGGGGGGGGGGGX;o=6;e=0.15" \
  -m 1:20 \
  -o "${sample_id}_R1.pretrimmed.fastq.gz" \
  -p "${sample_id}_R2.pretrimmed.fastq.gz" \
  "$r1" "$r2"
```

#### Parameter Details

<table data-header-hidden data-search="false"><thead><tr><th></th><th></th></tr></thead><tbody><tr><td><strong>Cutadapt option</strong></td><td><strong>Purpose</strong></td></tr><tr><td>-U 1</td><td>Trims 1 base from the 5′ end of Read 2, the transcript read.</td></tr><tr><td>-G AGAGTGAATGGG</td><td>Removes the TSO sequence from the 5′ side of the transcript read.</td></tr><tr><td>-G TCAACGCAGAGT</td><td>Removes the SMART PCR sequence from the 5′ side of the transcript read.</td></tr><tr><td>-A AAAAAAAAAAAA</td><td>Removes poly-A sequence from the 3′ side of the transcript read.</td></tr><tr><td>-A GGGGGGGGGGGG o=6;e=0.15"</td><td>Removes poly-G sequence from the 3′ side of the transcript read, requiring a 6 bp minimum overlap (o) and allowing 15% errors (e) in the matched region</td></tr><tr><td>-O 12</td><td>Requires a 12-base adapter overlap, matching the intended 12-base trimming stringency.</td></tr><tr><td>-e 0.10</td><td>Allows up to 10% mismatch, consistent with DRAGEN adapter trimming behavior.</td></tr><tr><td>--no-indels</td><td>Disallows insertions/deletions during adapter matching, making the behavior closer to DRAGEN’s mismatch-based adapter trimming.</td></tr><tr><td>-m 1:20</td><td>Keeps Read 1 effectively unfiltered by length, while requiring the transcript read to be at least 20 nt after trimming.</td></tr><tr><td>-p &#x3C;R2 path></td><td>Path to output trimmed R2 output</td></tr><tr><td>-o &#x3C;R1 path></td><td>Path to output R1 (untrimmed, but matching reads retained from R2)</td></tr></tbody></table>


# FASTQ Processing

For data processed with the Illumina Single Cell 3’ RNA Prep Kit, each read in R1 includes a cellular barcode sequence followed by a 3-base binning index (BI) sequence. R2 includes the sequences cDNA constructs created from the captured mRNA, which contain random cut sites that serve as intrinsic molecular identifiers (IMIs) and are used for molecular counting.

<figure><img src="/files/gSN9kpapiKhrTX1KB5oK" alt=""><figcaption></figcaption></figure>

For more information about FASTQ processing refer to the [DRAGEN documentation](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#fastq-processing) detailing the DRAGEN PIPseq scRNA Pipeline.


# Transcript Counting

Within each barcode and gene combination, IMIs are grouped in one of 64 bins, based on the 3-base binning index. For each bin, all identical IMIs are collapsed into a single count, since they are likely PCR duplicates of the same fragment generated during library prep.

Any barcode and gene combination that has ten or fewer unique binning indexes is assigned the number of unique binning indexes as its final count estimate. The pipeline then totals the number of IMIs associated with each remaining barcode and gene combination, and divides that number by the IPM correction factor, which accounts for the additional copies generated from a single captured molecule during five amplification cycles. The final count is the maximum between the floor of this value and the number of unique binning indexes for this barcode and gene.

Because all IMIs from the same parent molecule share a binning index, the number of unique binning indexes observed within a specific barcode and gene is determined by the number of molecules and is not impacted by the number of IMIs that were produced by the molecules. This means that the probabilistic relationship between the number of unique bins and the true number of molecules in a barcode and gene combination is constant and is the result of random sampling from the 64 possible bin indexes when each molecule is captured. For the subset of barcode and gene combinations with between 5 and 32 unique bin indexes, dividing the total number of IMIs by the average number of molecules expected based on the number of unique bin indexes gives you the estimated average IMIs per molecule (IPM).

The estimated molecular count for a barcode and gene is the total number of IMIs divided by the IPM, rounded down. The more true molecules a barcode and gene combination has, the true average IMIs per molecule should approach the average IPM of the sample. For barcode and gene combinations with very few molecules, the number of unique bins is expected to be a better predictor of the molecular count than the number of IMIs because the variance in the true IMIs per molecule among this group is high since the number of molecules in each individual barcode and gene combination is low. For this reason, IPM correction is applied for barcode and gene combinations with more than 10 unique bin indexes, and otherwise the corrected count is equal to the number of unique bin indexes.

<figure><img src="/files/MkHCpxB7lOa9nFcDJoGX" alt=""><figcaption></figcaption></figure>

For more information about transcript counting refer to the [DRAGEN documentation](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#estimating-the-correction-factor) detailing the DRAGEN PIPseq scRNA Pipeline.


# PIPseq CRISPR Mode

DRAGEN also supports processing samples from Illumina's CRISPR Single Cell kits using PIPseq technology. Setting `--scrna-enable-pipseq-crispr-mode` to true activates this mode.

Activating PIPseq CRISPR mode automatically configures DRAGEN for processing feature reads containing the CRISPR guide RNA (gRNA) sequences. This includes handling offsets in the cell-barcode position for the gRNA reads, transforming the gRNA cell-barcodes to match the gene expression ones, utilizing the "hook and grab" approach for identifying the gRNA reads, and counting the gRNA reads (disregarding BIs and IMIs). Both gene expression and gRNA reads are processed in the same single cell workflow, so extra steps are added to identify the hook sequence of gRNA reads. Note: unmapped reads do not contribute to gene expression read counts but are still included in gRNA counts if they match the hook sequence.

The “hook and grab” method is a targeted approach for identifying CRISPR perturbation reads. It leverages a conserved sequence within the guide RNA structural region as a “hook” to locate the guide RNA and then “grabs” the specific guide by mapping it to a database of known sequences based on their displacement from the hook.

<figure><img src="/files/ot2FgurtxcVk9asSxmLx" alt=""><figcaption></figcaption></figure>

For more information about PIPseq CRISPR mode refer to the [DRAGEN documentation](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#crispr-mode) detailing the DRAGEN PIPseq scRNA Pipeline.


# Guide RNA Calling

DRAGEN now also supports guide RNA calling - identifying cells that express guide RNA sequences. Although the counts matrix reports how many feature counts were found for each cell, there is a chance that the feature reads happened to match to the feature barcode reference sequence or the feature count was assigned to the cell by chance rather than due to real signal. Guide RNA calling applies an additional filtering step to distinguish between cells truly expressing guide RNA sequences and noise (similar to how cell filtering is performed to distinguish true passing cells from noise)

For more information on Guide RNA calling refer to the [DRAGEN documentation](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#pipseq-crispr-mode) detailing the Gaussian Mixture Model approach.


# Accessing Results

For information on tracking and viewing run and analysis results in BaseSpace Sequence Hub, refer to [View Data on Basespace](https://help.connected.illumina.com/basespace-sequence-hub/data/view-data).

To view results on BioInsight Platform Core, you may either click on "View Files in Platform Core" in the top right corner of your BSSH Analysis page, or directly access the analysis in Platform Core. It will be in a BSSH managed project with the same name as your BSSH workgroup. For information on viewing analysis results on your BioInsight Platform Core account, refer to [Viewing Data on Platform Core](https://help.connected.illumina.com/illumina-connected-analytics/project/p-data#viewing-data).


# DRAGEN Report

Running the Illumina DRAGEN Single Cell app produces a DRAGEN report in HTML format which includes QC metrics for trimming, fastQC, mapping, and single cell analysis as well as a barcode rank plot and UMAP. Below is a description of metrics and plots on each tab of the DRAGEN report.

#### [Single-Cell RNA](/dragen-single-cell-rna/analysis-results/dragen-report/single-cell-rna)

#### [Single-Cell Clustering](/dragen-single-cell-rna/analysis-results/dragen-report/single-cell-clustering)

#### [Trimmer](#dragen-fastqc)

#### [DRAGEN-FastQC](/dragen-single-cell-rna/analysis-results/dragen-report/dragen-fastqc)

#### [Mapping](/dragen-single-cell-rna/analysis-results/dragen-report/mapping)


# Single-Cell RNA

### Barcode Rank Plot

The barcode rank plot (often referred to as the “knee plot”), orders barcodes based on the number of molecules associated with them. Typically, the cell barcodes are concentrated at the top of the rank plot, whereas the background barcodes are concentrated in the lower portion of the plot. The purpose of the cell calling is to find a point in the first “knee” area that separates the cells from the background.&#x20;

A more negative slope and a poorly defined knee can be indicative of lower quality samples. It is not recommended to call cells too far down into the transition zone (area between first knee and second knee) as this risks adding compromised cells and background noise.

* Homogenous sample types and cells should be called towards the upper part of the transition zone
* More heterogenous sample types should be called ⅓ or ½ way down the transition zone

<figure><img src="/files/0tqn7cVE0vlpF42ji0s4" alt="" width="375"><figcaption></figcaption></figure>

### Single Cell RNA Metrics

| Metric                    | Description                                                                                                                                                                                                                                                                                                                                                             |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Input Reads               | Total number of reads that were successfully sequenced for each library and demultiplexed during BCL-convert. This includes reads missing barcodes, reads with exactly matching barcodes, reads with corrected barcdoes, and reads with non-matching barcodes.                                                                                                          |
| % Barcoded Reads          | Percentage of total input reads with a cell barcode that match the whitelist. This includes reads with exactly matching barcodes and reads with corrected barcodes (either an exact match or within one edit distance away for each of the four barcode tiers). Typically 85-95% for cell lines and RNA-rich cells or 70-90% for nuclei and cells with low RNA content. |
| % Mapped to Transcriptome | <p>Percentage of total barcoded reads that are mapped to the transcriptome and are included in downstream analysis (barcode matching, counting, etc.). This includes:</p><ul><li>unique exon matching reads</li><li>unique intron matching reads</li><li>mitochondrial reads</li></ul>                                                                                  |
| % Sequencing Saturation   | Represents the probability that adding additional sequenced read will have already been observed. The higher the sequencing saturation, the less benefit of uncovering new information with additional sequencing reads.                                                                                                                                                |
| Passing Cells             | Total unique barcodes that pass the count threshold as set by the ratio, inflection, or forced cell calling approaches.                                                                                                                                                                                                                                                 |
| % Reads in Passing Cells  | Number of reads in passing barcodes divided by the total number of barcoded reads in all cells                                                                                                                                                                                                                                                                          |
| Mean Reads per Cell       | Total input gene expression reads divided by the number of passing cells. Generally, should fall within range of 20,000-40,000 reads per (called) cell at minimum to ensure sufficient read depth. If mean reads per cell is lower than 20,000, it might be necessary to sequence at a higher read depth.                                                               |
| Median Molecules per Cell | The median number of captured RNAs per passing barcode. A molecule represents an individual transcript that was captured onto a PIP, was converted into a cDNA library molecule (with one or more daughter molecules during fragmentation), sequenced, then properly barcoded, mapped, and IMI-corrected.                                                               |
| Median Genes per Cell     | Median number of unique gene assignments observed across all molecules in the passing cell fraction                                                                                                                                                                                                                                                                     |
| Genes Detected            | Total unique genes detected in passing cells                                                                                                                                                                                                                                                                                                                            |

### Mapping Metrics

| Metric                    | Description                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| % Mapped to Transcriptome | <p>Percentage of total barcoded reads that are mapped to the transcriptome and are included in downstream analysis (barcode matching, counting, etc.). This includes:</p><ul><li>unique exon matching reads</li><li>unique intron matching reads</li><li>mitochondrial reads</li></ul>                                                                                     |
| % Mapped to Genome        | <p>Percentage of total barcoded reads that are:</p><ul><li>unique exon matching</li><li>unique intron matching</li><li>mitochondrial reads</li><li>filtered antisense reads</li><li>filtered ambiguously matching reads</li><li>filtered low MAPQ reads</li></ul><p>Filtered read categories are discarded from downstream analysis (barcode matching, counting, etc.)</p> |
| % Exon Matching           | Percentage of total barcoded reads that map to exons in genes                                                                                                                                                                                                                                                                                                              |
| % Intron Matching         | Percentage of total barcoded reads that map to introns in genes                                                                                                                                                                                                                                                                                                            |
| % Mitochondrial           | Percentage of total barcoded reads that map to mitochondria                                                                                                                                                                                                                                                                                                                |
| % Antisense               | Percentage of total barcoded reads that map to the opposite strand of a gene                                                                                                                                                                                                                                                                                               |
| % Ambiguous               | Percentage of total barcoded reads that map to multiple genes equally well                                                                                                                                                                                                                                                                                                 |
| % Low MAPQ                | Percentage of total barcoded reads that below the MAPQ threshold (default: 4). These reads map to multiple positions in the genome. This metric is not included if feature counting is enabled.                                                                                                                                                                            |
| % Non-Matching            | Percentage of total barcoded reads that do not map to any genes                                                                                                                                                                                                                                                                                                            |

### Feature Metrics

This section is only available for analyses with feature counting enabled.

| Metric                            | Description                                                                                                                                                                                                                                                              |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Input Feature Reads               | Total number of input feature reads                                                                                                                                                                                                                                      |
| % Barcoded Feature Reads          | Percentage of input feature reads with barcode matching whitelist after error-correction. This includes reads with exactly matching barcodes and reads with corrected barcodes (either an exact match or within 1 edit distance away for each of the four barcode tiers) |
| % Feature Matching Reads          | Percentage of total barcoded reads that match to a feature reference sequence                                                                                                                                                                                            |
| % Feature Reads in Cells          | Percentage of total barcoded feature reads that are reads in passing cells                                                                                                                                                                                               |
| % Feature Matching Reads in Cells | Percentage of feature matching CRISPR reads that are found in passing cells                                                                                                                                                                                              |
| % Cells with Features             | Percentage of passing cells with at least 1 CRISPR read count detected                                                                                                                                                                                                   |
| Mean Feature Reads per Cell       | Mean number of feature reads per passing cell                                                                                                                                                                                                                            |
| Median Feature Molecules per Cell | Median number of feature reads observed per passing cell. Note: CRISPR and feature read applications use read-based counting in DRAGEN because they do not contain IMIs                                                                                                  |
| Median Features per Cell          | Median number of unique guide RNA types (or feature types) present per passing cell                                                                                                                                                                                      |
| Features Detected                 | Total number of features observed across all cells in the dataset                                                                                                                                                                                                        |

### Guide Assignment Metrics

| Metric                      | Description                                                         |
| --------------------------- | ------------------------------------------------------------------- |
| Passing Cells with Features | Total number of passing cells containing at least one feature count |
| Cells with ≥ 1 Guide        | Number of feature-positive cells positive for at least 1 guide      |
| Cells with ≥ 2 Guides       | Number of feature-positive cells positive for 2 or more guides      |


# Single-Cell Clustering

<figure><img src="/files/f1ohch9mQhcw7xGrnDpj" alt=""><figcaption></figcaption></figure>

### UMAP Plot

The UMAP plot allows visualization of individual cells in 2D space to capture the similarities between cells. Louvain clustering is performed to quantify and visualize heterogeneity within the cell population and begin to identify different cell types.&#x20;

### Top Marker Genes

The Top Marker Genes table includes the top 10 marker genes for each cluster of the UMAP. These genes can be used to identify subpopulations that correspond to actual biology. For each gene, the gene name, ENSEMBL gene ID, log2 fold change, and pValue are shown. The contents of the table are also available in CSV format by selecting **Download CSV.**


# Trimmer

### Trimmed Reads

| Metric                    | Description                                                                    |
| ------------------------- | ------------------------------------------------------------------------------ |
| Input Reads               | Total number of input reads to DRAGEN                                          |
| Max Read Length           | Maximum detected input read length                                             |
| Average Input Read Length | Average input read length to DRAGEN, after any adapter trimming by BCL Convert |
| Masked                    | Total number 3’ Poly-G bases masked from mapping                               |
| Trimmed                   | Total number of reads trimmed by DRAGEN                                        |
| Filtered                  | Total number of reads removed from the input by DRAGEN                         |

### Trimmer Metrics

| Metric       | Description                                |
| ------------ | ------------------------------------------ |
| Fixed-Length | Total number of fixed-length trimmed reads |
| Adapter      | Total number of adapter trimmed reads      |


# DRAGEN FastQC

### DRAGEN FASTQC Plots

| Plot                               | Description                                                                                                                                                                                                                                                                                        |
| ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Base Quality by Position           | Phred-scale quality value for bases at a given location                                                                                                                                                                                                                                            |
| Mean Base Quality by Position      | Average Phred-scale quality value of bases with a specific nucleotide and at a given location in the read                                                                                                                                                                                          |
| Read Length Distribution           | Total number of reads with each observed length                                                                                                                                                                                                                                                    |
| Read Quality Distribution          | Total number of reads with each observed average Phred-scale quality score                                                                                                                                                                                                                         |
| %GC Content                        | Percentage of sequences with each GC content across the whole length of each sequence compared to a modelled normal distribution of GC content                                                                                                                                                     |
| Read Quality by %GC Content        | Average Phred-scale read mean quality for reads with each GC content percentile between 0% and 100%                                                                                                                                                                                                |
| Ambiguous Base Content by Position | Percent ambiguous bases at a given location                                                                                                                                                                                                                                                        |
| Base Content by Position           | Percent of bases of each specific nucleotide at given locations in the read                                                                                                                                                                                                                        |
| Adapter Content by Position        | Percentage of the proportion of your library which has seen each of the adapter sequences at each position. Once a sequence has been seen in a read it is counted as being present right through to the end of the read so the percentages you see will only increase as the read length increases |


# Mapping

### Read-Level Metrics

| Metric            | Description                                                                     |
| ----------------- | ------------------------------------------------------------------------------- |
| Total input reads | Total number of input reads                                                     |
| QC-failed reads   | Total number of reads failing one or more quality checks                        |
| % QC-failed       | Percentage of reads failing one or more quality checks                          |
| Unique reads      | Total number of unique reads                                                    |
| % Unique          | Percentage of reads that are unique                                             |
| Mapped reads      | Total number of mapped reads, adjusted for filtered and excluded targets        |
| % Mapped          | Percentage of reads that are mapped, adjusted for filtered and excluded targets |

### Base-Level Metrics

| Metric      | Description                                          |
| ----------- | ---------------------------------------------------- |
| Total Bases | Total number of input bases                          |
| Mapped R1   | Total number of mapped bases on R1                   |
| % Mapped R1 | Percentage of R1 bases mapped                        |
| % Q30 R1    | Percentage of R1 bases with phred quality score >=30 |

### Deduplication

The deduplication bar chart reflects the ratio of unique and deduplicated reads.

### Read MAPQs

The read MAPQs bar chart reflects the percent of reads in various categories of phred quality score (Q0-Q10, Q10-Q20, Q20-Q30, Q30-Q40, Q40+).


# Secondary Analysis Results

The following table describes the files created by DRAGEN Single Cell RNA:

| File                                                       | Description                                                                                                                               |
| ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| \<Sample\_ID>.scRNA.bam                                    | Binary Alignment Map (BAM) files containing information about all reads in the input FASTQ files that were mapped to the reference genome |
| \<Sample\_ID>.scRNA.bam.bai                                | Index file for the BAM for use by downstream applications                                                                                 |
| \<Sample\_ID>.scRNA.barcodeCounts.txt                      | Text file containing the counts per barcode                                                                                               |
| \<Sample\_ID>.scRNA.barcodeSummary.tsv                     | Summary of barcode statistics including number of reads, genes, and molecules in each barcode                                             |
| \<Sample\_ID>.scRNA\_metrics.csv                           | Single cell metrics summary with assay sensitivity and quality metrics                                                                    |
| \<Sample\_ID>.scRNA.moleculeInfo.h5                        | Summary of the read counts for each barcode, gene, BI, and IMI combination in HDF5 format.                                                |
| \<Sample\_ID>.scRNA.matrix.mtx.gz                          | Sparse matrix with rows that represent genes and features detected, and columns that consist of all barcodes that were detected           |
| \<Sample\_ID>.scRNA.features.tsv.gz                        | Information about the features corresponding to the rows of the sparse matrix                                                             |
| \<Sample\_ID>.scRNA.barcodes.tsv.gz                        | List of barcodes corresponding to the columns of the sparse matrix                                                                        |
| \<Sample\_ID>.scRNA.h5ad                                   | Count matrix in AnnData format. Barcodes are stored in the OBS group and genes are stored in the VAR group.                               |
| \<Sample\_ID>.scRNA.filtered.matrix.mtx.gz                 | Filtered sparse matrix with rows that represent genes and features detected, and columns that consist of all barcodes that were detected  |
| \<Sample\_ID>.scRNA.filtered.features.tsv.gz               | Information about the features corresponding to the rows of the filtered sparse matrix                                                    |
| \<Sample\_ID>.scRNA.filtered.barcodes.tsv.gz               | List of barcodes corresponding to the columns of the filtered sparse matrix                                                               |
| \<Sample\_ID>.scRNA.filtered.h5ad                          | Filtered count matrix in AnnData format. Barcodes are stored in the OBS group and genes are stored in the VAR group.                      |
| \<Sample\_ID>.scRNA.positive\_cell\_guide\_assignments.csv | CSV containing all passing cells with information on guide RNA sequences called for each barcode                                          |
| \<Sample\_ID>.scRNA.singlet\_cell\_guide\_assignments.csv  | CSV containing only cells for which exactly one guide RNA feature was called                                                              |
| \<Sample\_ID>.scRNA.guide\_metrics.csv                     | CSV metrics file containing guide calling metrics                                                                                         |


# Third-party CRISPR Guide Assignment Toolkit

Peer-Reviewed Option for performing CRISPR gRNA thresholding

[Crispat](https://academic.oup.com/bioinformatics/article/40/9/btae535/7750392) is a peer-reviewed CRISPR guide calling package that conveniently bundles many popular gRNA thresholding approaches, enabling users to compare the performance of multiple approaches side-by-side.&#x20;

Here, we implement one of these approaches, the Gaussian-Gaussian Mixture Model, which we've found performs best with the Illumina's Single-Cell CRISPR Prep kit.&#x20;

[Instructions for how to download and install](https://github.com/aamayzhang/crispat_iscp/tree/main) our optimized fork of crispat can be found[ here](https://github.com/aamayzhang/crispat_iscp/tree/main).

In brief, the model does the following:

* Read in DRAGEN filtered matrix files from an DRAGEN output directory.
* Creates and reads in a .h5ad file (via scanpy).
* Performs 2-component Gaussian-Gaussian mixture modeling on the gRNA counts.
* Outputs gRNA assignments for each cell barcode along with plots showing the fit of the model to the raw counts.

A [step-by-step tutorial](https://github.com/aamayzhang/crispat_iscp/blob/main/tutorial/iscp_demo.py) and a [small test dataset](https://github.com/aamayzhang/crispat_iscp/tree/main/example_data) can be found in the repo.

You should expect the plots of the models fit to nonzero count data to look similar to this:&#x20;

<figure><img src="/files/G5QOryG0K4du0jpWdl4N" alt=""><figcaption></figcaption></figure>

The output file *assignments.csv* provides a full list of all cell barcodes containing a guide that exceeded the threshold determined by the gaussian mixture model.  The outputs should look as follows:

<div align="left"><figure><img src="/files/dy17ZzePjecJ4XUwXk01" alt=""><figcaption></figcaption></figure></div>

`cell`:  the cell barcode sequence.

`gRNA` : the gRNA ID that was assigned to the cell.

`read_counts`:  the number of read counts observed for each cell-gRNA assignment.


# Illumina Connected Multiomics

Illumina Connected Multiomics (ICM) is available for further tertiary analysis of single-cell and other multiomic data.

## Getting Started

Refer to the following links to the ICM user guide to get started with ICM:

* [Registration and Login](https://help.connected.illumina.com/icm/introduction/icm)
* [Data Inputs](https://help.connected.illumina.com/icm/introduction/data-inputs)
* [Creating a Study from a Platform Core Project](https://help.connected.illumina.com/icm/studies/create-study)
* [Viewing Results and Navigating in ICM](https://help.connected.illumina.com/icm/studies/enter-study)

## Default Sample Settings

Each Illumina DRAGEN Single Cell sample contains a feature-barcode matrix file. By default, if the feature ID is not unique, the feature will be summarized by 'Mean' as the Deduplication method and the feature name is used as the primary feature identifier. The count value format is 'Raw counts' (this is the same whether the 'filtered' sparse matrix files are used or the un-filtered files). All features and cells with a total read count of at least 400 are reported. This information is visually depicted below.

<figure><img src="/files/UFkT7TzsPcVjnBqmio3l" alt=""><figcaption></figcaption></figure>

If any changes to the default Illumina DRAGEN Single Cell sample settings are desired, add data to your study using the scRNA feature-barcode-matrix option and then create a Custom: Third Party Assays analysis. After initiating the analysis, file format options can be adjusted through the required user input on the analysis. Once finished, the import will proceed with the options selected.

## Default Single Cell Analysis

Below are explanations of the steps that are run in the default single-cell analysis that is automatically launched on import of single-cell data in ICM. Also included below are the instructions to launch each step manually if input parameters need to be adjusted from the default settings.

### Normalize Counts

Because different cells will have a different number of total counts, it is important to normalize the data prior to downstream analysis. For droplet-based single cell isolation and library preparation methods that use a 3' counting strategy, where only the 3' end of each transcript is captured and sequenced, we recommend the following normalization:

1. Divide by the sum
2. Multiply by 10000
3. Add 1
4. Log e

This accounts for differences in total IMI counts per cell and log transforms the data, which makes the data easier to visualize. While the above normalization is already run in a default analysis, additional normalization can be done by following the steps below.

* Click the counts node you wish to normalize
* Click **Normalization and scaling** in the context-sensitive task menu on the right
* Click **Normalization**
* Click **Use recommended** to add the recommended normalization scheme

This divides by the Sum, Multiplies by 10000, adds one, then takes log e to the *Normalization order* panel. Normalization steps are performed in descending order.

* Click **Finish** to apply the normalization

<figure><img src="/files/MtJ79lNNNuAciLfnPdun" alt=""><figcaption></figcaption></figure>

### PCA

Principal components (PC) analysis (PCA) is an exploratory technique that is used to describe the structure of high dimensional data by reducing its dimensionality. Because PCA is used to reduce the dimensionality of the data prior to clustering as part of a standard single cell analysis workflow, it is useful to examine the results of PCA for your data set prior to clustering.

* Select a data node containing the normalized and filtered count matrix
* Click **Exploratory analysis** in the task menu
* Click **PCA** from the drop-down list
* Select the number of features to include
* Select the number of PCs to calculate

You can choose *Features contribute* **equally** to standardize the genes prior to PCA or allow more variable genes to have a larger effect on the PCA by choosing **by variance**. By default, we take variance into account and focus on the most variable genes.

If you have multiple samples, you can choose to run PCA for each sample individually or for all samples together by selecting or not selecting the *Split by sample* option.

* Click **Finish** to run

<figure><img src="/files/ltyc3xPWADQcaQmBq3u0" alt=""><figcaption></figcaption></figure>

A new *PCA* task node will be produced on the task graph for the analysis. When complete, double-click the **PCA** task node to open the 3D PCA scatter plot in data viewer.

Beside PCA coordinates of the cells, PCA task report also includes, the Scree plot, the component loadings table, and the PC projections table.

The Scree plot lists PCs on the x-axis and the amount of variance explained by each PC on the y-axis, measured in Eigenvalue. The higher the Eigenvalue, the more variance is explained by the PC. Typically, after an initial set of highly informative PCs, the amount of variance explained by analyzing additional PCs is minimal. By identifying the point where the Scree plot levels off, you can choose an optimal number of PCs to use in downstream analysis steps like graph-based clustering, UMAP and t-SNE.

### Graph-based Clustering

Graph-based clustering identifies groups of similar cells using PC values as the input. By including only the most informative PCs, noise in the data set is excluded, improving the results of clustering.

* Click the *PCA* data node
* Click **Exploratory analysis** in the task menu
* Click **Graph-based clustering**

Clustering can be performed on each sample individually or on all samples together.

* Select the **Clustering algorithm** to use. The default Single-Cell analysis uses the Louvain algorithm.
* Check **Compute biomarkers** to compute features that are highly expressed when comparing each cluster
* Select the number of **PCs to use**
* Click **Configure** to access the *Advanced options*

The *Number of principal components* can be set based on the your examination of the Scree plot and component loadings table. The default value is likely exhaustive for most data sets; altering this value may introduce noise that influences the number of clusters that are distinguished.

* Click **Finish** to run the task

<figure><img src="/files/5uao3lFm2FRDTgdtAzJr" alt=""><figcaption></figcaption></figure>

A new *Graph-based clusters* data and *Biomarkers* data node will be generated along with the task nodes

* Double-click the **Graph-based clusters** node to see the cluster results and statistics. The *Graph-based clustering result* lists the *Total number of clusters* and what proportion of cells fall into each cluster.
* Double-click the **Biomarkers** node to see the computed biomarkers if you have selected this option. The *Biomarkers* node includes the top features for each graph-based cluster. It displays the top-10 genes that distinguish each cluster from the others. **Download** at the top left of the table can be used to view and save more features. These are calculated using an ANOVA test comparing the cells in each group to all the other cells, filtering to genes that are 1.5 fold upregulated, and sorting by ascending p-value. This ensures that the top-10 genes of each cluster are highly and disproportionately expressed in that cluster.

### UMAP

Uniform Manifold Approximation and Projection (UMAP) is a dimensional reduction technique. UMAP aims to preserve the essential high-dimensional structure and present it in a low-dimensional representation. UMAP is particularly useful for visually identifying groups of similar samples or cells in large high-dimensional data sets such as single cell RNA-Seq.

* Click the **Graph-based clusters** or **PCA** node
* Click **Exploratory analysis** in the task menu
* Click **UMAP**
* Select the number of **PCs to use**
* Click **Configure** to access the *Advanced options*
* Click **Finish** to run

If you have multiple samples, you can choose to run UMAP for each sample individually or for all samples together using the *Split cells by sample* option.

Like Graph-based clustering, UMAP takes PC values as its input and further reduces the data down to two or three dimensions. For consistency, you should use the same number of PCs as the input for UMAP that you used for Graph-based clustering.

A new *UMAP* task node will be produced. When complete, double-click the **UMAP** node to open the UMAP task report. Use the panel on the left to modify the plot or add more plots to this Data viewer session.

The UMAP scatter plot is interactive and can be viewed in 2D or 3D. The UMAP plot is 3D by default. You can rotate the 3D plot by left-clicking and dragging your mouse or using **Control** under *Configure*. You can zoom in and out using your mouse wheel. You can pan by right-clicking and dragging your mouse. You can use **Style** to modify *color*, *shape*, *size*, and *labeling* (e.g. add a fog effect to improve depth perception on the plot). Add a 2D plot clicking **New plot,** selecting **2D Scatter plot** and selecting UMAP as the source of the data.

## Other Single-Cell Analysis Tasks

### QA/QC

The Single-cell QA/QC task in ICM enables you to visualize several useful metrics that will help you include only high-quality cells. To invoke the Single-cell QA/QC task:

* Click a **Single cell counts** data node
* Click the **QA/QC** section of the task menu
* Click **Single cell QA/QC**

By default, all samples are used to perform QA/QC. You can choose to split the sample and perform QA/QC separately for each sample.

You will be prompted to choose the genome assembly and annotation file by the Single cell QA/QC configuration dialog and ideally this closely matches the references used in the DRAGEN secondary analysis. Note, it is still possible to run the task without specifying an annotation file. If you choose not to specify an annotation file, the detection of mitochondrial counts will not be possible.

<figure><img src="/files/rcCJeYWgWww6K7aCyAnH" alt=""><figcaption></figcaption></figure>

**For Additional information regarding QA/QC see the following page in the ICM user guide**: [Single-cell QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/single-cell-qa-qc)

### Filter Features

A common task in single-cell RNA-Seq analysis is to filter the data to include only informative genes (features). Because there is no gold standard for what makes a gene informative or not and ideal gene filtering criteria depends on your experimental design and research question, ICM has a wide variety of flexible filtering options. The Filter features step can also be performed before normalization or after normalization.

* Select a data node containing the count matrix
* Click **Filtering** in the task menu
* Click **Filter features**
* Select the **Filter type** and **Filter criteria** desired

There are four categories of filter available - noise reduction, statistics-based, feature metadata, and feature list.

The noise reduction filter allows you to exclude genes considered background noise based on a variety of criteria. The statistics-based filter is useful for focusing on a certain number or percentile of genes based on a variety of metrics, such as variance. The metadata, saved list, and manual list filters allow you to filter your data set to include or exclude particular genes.

For example, you can use a noise reduction filter to exclude genes that are not expressed by any cell in the data set, but were included in the matrix file. To do so:

* Click the **Noise reduction filter** check box
* Set the *Noise reduction filter to Exclude features where* **value <= 0 in at least 99.9% of cells** using the drop-down menus and text boxes
* Click **Finish** to apply the filter

<figure><img src="/files/KNoceELMYixSL1tHpjWm" alt=""><figcaption></figcaption></figure>

The default single cell pipeline uses the statistics-based filter to filter for the top 10% of features with the highest variance.

### t-SNE

t-Distributed Stochastic Neighbor Embedding (t-SNE) is a dimensional reduction technique that prioritizes local relationships to build a low-dimensional representation of the high-dimensional data that places objects that are similar in high-dimensional space close together in the low-dimensional representation. This makes t-SNE well suited for analyzing high-dimensional data when the goal is to identify groups of similar objects, such as cell types in single cell RNA-Seq data.

* Click the **Graph-based clusters** or **PCA** node
* Click **Exploratory analysis** in the task menu
* Click **t-SNE**
* Select the number of **PCs to use**
* Click **Configure** to access the *Advanced options*
* Click **Finish** to run

The t-SNE scatter plot visualization has the same functionality and style elements as the UMAP plot described above.

### Differential Analysis

A common goal in single cell analysis is to identify genes that distinguish a cell type. To do this, you can use the differential analysis tools in ICM.

* Click the **Normalized counts** results node
* Click **Statistics** in the toolbox
* Click **Differential Analysis**
* Select **ANOVA** as the Method to use for differential analysis and click **Next** (note that other single cell suggested models include the **Hurdle model** or **Wilcoxon** but you are not limited)
* Select and add the categorical and/or numeric factors for analysis
* Click **Next**

<figure><img src="/files/qPGXGqYG3zcTzJ1e5oIE" alt=""><figcaption></figcaption></figure>

The differential analysis tool can be used to compare one group of cells to another group of cells to identify genes or features that distinguish cells. Common examples include determining distinguishing genes between one cell type and all others, two cell types, or the same cell type between two experimental conditions.

The comparison builder can be used to create any of these tests. The top panel is the numerator for fold-change calculations so usually the experimental or test groups are selected in the top panel. The bottom panel is the denominator for fold-change calculations so the control group is often selected in the bottom panel.

* Add attributes/classifications to the **numerator**
* Add attributes/classifications to the **denominator**
* Select **Combine** for a single comparison or **Pairwise** for a factorial set of comparisons
* Select **Add comparison**
* Optionally select the checkbox to **Apply lowest average coverage filter** to exclude a feature if the geometric average of its values over all samples is less than the specified value. This can be useful if no noise reduction filter has already been applied in the pipeline.
* Click **Configure** to access the *Advanced options* which includes other Multiple test correction options.
* Click **Finish** to run

<figure><img src="/files/CI2bPk5a4MoXd9bsHzf0" alt=""><figcaption></figcaption></figure>

When completed, double click the newly generated data node to open the **ANOVA** task report. The ANOVA task report lists genes on rows and the results of the statistical test (p-value, fold change, etc.) on columns. Genes are listed in ascending order by the p-value of the first comparison so the most significant gene is listed first.

#### Filter for Significant Genes

Using the filter control panel on the left, we can filter to just the genes that are significantly different for the comparison using the p-value and/or multiple test correction value (FDR step-up by default). The number of genes at the top of the filter control panel updates to indicate how many genes are left after the filters are applied.

Click **Generate filtered node** to generate a filtered version of the table for downstream analysis. This new data node containing the filtered genes will run in the Analyses pipeline to generate a filtered **Feature feature list** data node which will be available in the task graph by closing the ANOVA report and navigating to the Analyses pipeline; this filtered list can now be used for downstream tasks.

### Gene set enrichment

While a long list of significantly different genes is important information about a cell type, it can be difficult to identify what the biological consequences of these changes might be just by looking at the genes one at a time. Using enrichment analysis, you can identify gene sets and pathways that are over-represented in a list of significant genes, providing clues to the biological meaning of your results.

* Click the **Feature list** data node produced by the Differential analysis filter
* Click **Biological interpretation** in the task menu
* Click **Gene set enrichment**
* Select the **Database** to use. ICM distributes the gene sets from the Gene Ontology Consortium, but Gene set enrichment can work with any custom or public gene set database.
* Choose the latest assembly available from the Gene set drop-down
* Click **Finish**

When completed, double-click the *Gene set enrichment* task node to open the task report.

The **Gene set enrichment** task report lists gene sets on rows with an enrichment score and p-value for each. It also lists how many genes in the gene set were in the input gene list and how many were not. Clicking the Gene set ID links to the geneontology.org or KEGG page for the gene set.

### Hierarchical clustering / heatmap

Since we have filtered to a list of significantly different genes, we can visualize these genes by generating a heatmap or bubble map.

* Click the **Filtered feature list** data node produced by the Differential analysis filter
* Click **Exploratory analysis** in the toolbox
* Click **Hierarchical clustering** **/ heatmap**

This task is used to generate the heatmap or bubble map; choose **Heatmap** as the plot type. You can choose to **Cluster** features (genes) and cells (samples) under *Feature order* and *Cell order* in the *Ordering* section which will perform hierarchical clustering producing a dendrogram which is useful for determining relationships. For single cell data sets, you may choose to forgo clustering the cells in favor of ordering them by the attribute of interest (e.g. drag and drop to order the attribute in a way that makes sense). Both ordering methods help to make the heatmap more comprehensible. [Please click here for more information on this task.](https://help.multiomics.illumina.com/partek/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering#invoking-hierarchical-clustering)

* Select **Feature order**
* Select **Cell order**
* Optionally add any additional **Filtering**
* Click **Configure** to access the *Advanced options*
* Click **Finish** to run

### Cell Typing with ScType

ScType allows automated cell-type identification based on scRNA-seq data along with a comprehensive cell marker database as background information.

* Click the data node containing the non-normalized count matrix
* Click on **Classification** > **Single cell type** in the toolbox
* Select the marker database from the drop-down menu, the original full ScType database is provided by default
* Select categorical attributes to **Categorize by** (e.g. graph-based clusters)
* Optionally **Filter tissue types**
* Select the **SC Type algorithm** to use
* Click **Configure** to access and change any *Advanced options*
* Click **Finish** to run

A new *scType classification* task node will be produced. When complete, double-click the **Single cell type** node to open the results of the cell-type identification. For each cell, the tissue, sctype result, and typescore are reported.


# Troubleshooting


# FAQ


# Introduction

The DRAGEN miRNA software is integral to the Secondary Analysis step in Illumina's miRNA prep End-to-End (E2E) solution, accessible via the BaseSpace Sequence Hub (BSSH) and Illumina Connected Analytics (ICA) cloud platforms. This section aims to illustrate how counts align within the workflow and to outline the other components of the E2E solution (see Figure 1).

<figure><img src="/files/N9cYmeethOFuIZgzNk9F" alt=""><figcaption><p>Figure 1. Illumina miRNA prep E2E worfklow.</p></figcaption></figure>

#### Illumina miRNA E2E workflow

After preparing samples with the Illumina miRNA Prep kit, customers can create a Sample Sheet during the **BSSH Run Planning** step and upload or select it on the **sequencer**.

Once sequencing is complete, they can demultiplex the data using the **BCLConvert** app in BSSH or a local installation to generate FASTQ files.

These FASTQs are then used for secondary analysis (DRAGEN miRNA app), which produces small RNA count matrices, mapping statistics, and a full Quality Control **DRAGEN Report**.

Final outputs, including count matrices and DRAGEN Report, can be downloaded from BaseSpace or **ICA**, depending on the customer’s subscription.

Customers can perform tertiary analyses—such as differential expression analysis—using Illumina Connected Multiomics. Refer to the[ Illumina Connected Multiomics](/dragen-mirna/tertiary-analysis/illumina-connected-multiomics) section for additional details.


# Counting Algorithm

The Illumina miRNA sequencing fastqs are uploaded to either BSSH or ICA, where reads are processed through the following steps in the DRAGEN miRNA app:

1. **Calibration of miRBase entries**\
   For miRNA entries with identical or nearly identical sequences in the miRBase mature database, manual calibration is performed. A combined entry is generated for each overlapping miRNA set. For instance, the sequence of *hsa-miR-151b* is entirely contained within *hsa-miR-151a-5p*; therefore, the resulting entry is reported as *hsa-miR-151b/151a-5p*.
2. **Adapter and quality trimming**\
   Reads are processed with *cutadapt* ([documentation](https://cutadapt.readthedocs.io/en/stable/guide.html)) to remove 3′ adapters (AACTGTAGGCACCATCAAT) and low-quality bases. Reads lacking adapter sequences are separately tallied as *no\_adapter\_reads*.
3. **Insert and UMI identification**\
   After trimming, insert and UMI sequences are extracted. Reads with inserts shorter than 16 bp (*too\_short\_reads*) or UMIs shorter than 10 bp (*UMI\_defective\_reads*) are discarded.
4. **Insert sequence alignment**\
   A unique sequence set is generated across all samples within a submitted job. Insert sequences are annotated using a sequential alignment strategy with *bowtie* (bowtie-bio.sourceforge.net). Alignments proceed in the following order:

   * Perfect match to miRBase mature
   * miRBase hairpin
   * Noncoding RNA, mRNA, otherRNA
   * Secondary alignment to miRBase mature (allowing up to two mismatches)

   At each step, only unmapped sequences are passed forward. Read counts are reported per RNA category (e.g., *miRNA\_Reads, hairpin\_Reads, piRNA\_Reads, tRNA\_Reads, rRNA\_Reads, mRNA\_Reads*). miRBase is used for miRNAs (v21 or v22), while piRNABank is referenced for piRNAs.

   For human, mouse, and rat, a species-specific miRBase mature database is used, followed by genome alignment of remaining sequences to identify potential novel miRNAs (human: GRCh38, mouse: GRCm38, rat: Rnor\_6.0). For all other species, a comprehensive miRBase mature database is applied.
5. **Counting reads and unique molecules**\
   For each sample, all reads assigned to a given miRNA or piRNA ID are tallied, and UMIs are aggregated to calculate unique molecule counts. Results are reported as follows:
   * *miRNA\_piRNA* sheet: read counts and UMI counts for miRNAs and piRNAs
   * *tRNA* and *otherRNA* sheets: results for tRNAs and other RNAs
   * *notCharacterized\_mappable* sheet: reads and clustered UMIs aligned to the genome in the final step (human, mouse, rat only)
   * *notCharacterized\_notMappable*: tally of all remaining unmapped reads


# Reference Database

DRAGEN miRNA app offers users the option to select from two reference databases. This section outlines the process of constructing these databases.

### Reference v21

Mature miRNA and miRNA hairpin libraries were obtained from [miRBase.org](https://www.mirbase.org/) using version v21. Human (*Homo sapiens*), mouse (*Mus musculus*), rat (*Rattus norvegicus*), fruitfly (*Drosophila melanogaster*), nematode (*Caenorhabditis elegans*) and zebrafish (*Danio rerio*) mRNA and other noncoding RNA libraries were obtained from Ensembl ([www.ensembl.org/](http://www.ensembl.org/)), unless otherwise denoted. Human tRNAs were obtained from the Genomic [tRNA Database](https://pubmed.ncbi.nlm.nih.gov/18984615/). Human snoRNA was obtained from the [snoRNABase](https://www-snorna.biotoul.fr/).

### Reference v22

The libraries used when selecting reference v22 are available here [SourceForge](https://sourceforge.net/projects/mirge3/files/miRge3_Lib/). Mature miRNA and miRNA hairpin libraries were obtained from [miRBase.org](https://www.mirbase.org/) using version v22.

### Available Species

List of species supported in the software (compatible with both v21 and v22):

* Human (*Homo sapiens*)
* Mouse (*Mus musculus*)
* Rat (*Rattus norvegicus*)
* Fruitfly (*Drosophila melanogaster*)
* Nematode (*Caenorhabditis elegans*)
* Zebrafish (*Danio rerio*)


# Read Structure

The Illumina miRNA Prep Kit is a single-end read solution that includes adapters and Unique Molecular Indices (UMIs). These UMIs tag each miRNA at an early stage, reducing PCR and sequencing bias.

The 3′ adapter sequence (AACTGTAGGCACCATCAAT) is followed by the random UMI sequence, which is typically a 12 bp long random sequence. This is immediately followed by the TruSeq adapter sequence (AGATCGGAAGAGCACACGTCTGAACTCCAGTCA)

<figure><img src="/files/X30tc0YfXQeaLWVT1iTV" alt=""><figcaption><p>Figure 2. Read structure of Illumina miRNA prep reads.</p></figcaption></figure>

The Truseq adapter sequence is removed at the demultiplexing stage (via BCLConvert) prior to secondary analysis.

The 3’ adapter sequence is removed at the secondary analysis step, after trimming, reads are classified accordingly:

* Reads are processed with cutadapt to remove 3′ UMI adapter and low-quality bases.
* Reads lacking 3’ UMI adapter sequence are separately tallied as no\_adapter\_reads.
* Insert and UMI identification After trimming, insert and UMI sequences are extracted. Reads with inserts shorter than 16 bp (too\_short\_reads) or UMIs shorter than 10 bp (UMI\_defective\_reads) are discarded


# Prerequisites

* NovaSeq 6000/6000Dx, NextSeq 1000/2000, NextSeq 500/550, MiniSeq, MiSeq or iSeq 100
* Illumina miRNA Prep Kit
  * For more information about the library preparation kit, refer to the [Illumina miRNA Prep Support Site](https://support.illumina.com/sequencing/sequencing_kits/).
* A cloud account with a valid subscription. For information on registering your BaseSpace Sequence Hub or Illumina Connected Analytics account, refer to [Software Registration page](https://help.connected.illumina.com/account-management/rg-registration).


# Sample Sheet Creation

Before launching the DRAGEN miRNA app on Illumina's cloud platforms (BSSH or ICA), ensure that samples are demultiplexed and FASTQ files are created. Use BCLConvert locally or on the cloud for this purpose.

To run BCLConvert a sample sheet is required. A sample sheet is a comma-separated value (\*.csv) file format used by Illumina instruments, platforms, and analysis pipelines to store settings and data for sequencing and analysis. For general information on the v2 sample sheet, refer to [Illumina Connected Software - Sample Sheet](https://help.connected.illumina.com/run-set-up/overview).

To create a sample sheet compatible with BCLConvert using the [BSSH Run Planner](https://ilmn.basespace.illumina.com/dashboard), follow these steps:<br>

1. Select the **Runs** tab, and then select the **New Run** drop-down.
2. Select **Run Planning**.
3. Run Settings wizard will be loaded.
4. In the Run Name field, enter a unique name of your preference to identify the current run. The run name can contain a maximum of 255 alphanumeric characters, spaces, dashes, and underscores.
5. **\[Optional]** In the Run Description field, enter a description of the current run. The run description can contain a maximum of 255 alphanumeric characters.
6. Select the Instrument Platform (one of the following: **NovaSeq 6000/6000Dx, NextSeq 1000/2000, NextSeq 500/550, MiniSeq, MiSeq or iSeq 100**).
7. Select the analysis location. Depending on the selected instrument type, not all options may be available.
   * **BaseSpace** - Analyze sequencing data in the cloud.
   * **Local** - Analyze sequencing data on-instrument or generate a Sample Sheet v2 for Local or Hybrid mode.
8. **\[Optional]** In the Library Tube ID field, optionally enter the library tube ID of the current run. The library tube id can contain a maximum of 255 alphanumeric characters.
9. Select Next
10. Configuration wizard will be loaded
11. Select an analysis type and version. Choose the latest version of **BCL Convert** (e.g., DRAGEN BCL Convert 4.3.13)
12. Select the following instrument settings. Depending on the library prep kit, recommended options are automatically selected. Some library prep kits have hard-coded number of indexes reads and read types, which cannot be changed.
    * Library prep kit: **Library miRNA prep**
    * Index adapter kit: one of the following: **Illumina miRNA UDI Indexes Set A, B, C, or D**
    * Number of index reads: **2 indexes**
    * Read type: **Single Read**
    * Read Lengths:
      * Read 1: 72
      * Index 1: 10
      * Index 2: 10
      * Read 2: 0
    * Override Cycles:
      * Read 1: Y72
      * Index 1: I10
      * Index 2: I10

\
For more information on different instruments, refer to [Illumina support documentation](https://help.basespace.illumina.com/sequence/plan-runs).


# Analysis in BSSH

### Overview

This tutorial guides you through running the Illumina DRAGEN miRNA pipeline using [BSSH](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub.html). You’ll begin by setting up a project with demultiplexed sequencing data (FASTQ files), selecting the DRAGEN miRNA app and launching an analysis run.

***

### Set-up

This tutorial assumes you already have an existing project in BaseSpace.

If you do not yet have a project:

* Create a new project from the **Projects** page in BaseSpace.
* Upload or import demultiplexed FASTQ files into the project.

You will also need access to the DRAGEN miRNA App. This app is typically available through your BaseSpace subscription or via your organization’s licensed apps.

### Preparing Your Project

#### Uploading or Importing Data

1. Go to **Projects** in BaseSpace.
2. Open your target project.
3. Upload FASTQ files or import them from an instrument run or external storage.
4. Confirm that all required FASTQ files appear in your project’s **Samples/Data** section.

### Apps

After setting up your project and data, you can run analysis apps.

#### DRAGEN miRNA App

This example demonstrates how to run the DRAGEN miRNA app in your BaseSpace project using your uploaded FASTQ data.

Required inputs:

* FASTQ files stored in your BaseSpace project
* A reference database
* Species selection

### Launching the DRAGEN miRNA App

1. From your project page, click on the rocket icon (or go to the **Apps** tab).
2. Search for and select **DRAGEN miRNA**.
3. Click **Launch App**.

#### **Initial Setup**

You will be prompted to:

* Enter a meaningful **analysis name**
* Select a **project** for output storage
* Confirm your **compute/subscription** options (if applicable)

**Input Configuration**

In the app configuration screen, provide the following inputs:

**FASTQ/Biosample Files**

* Select the FASTQ files from your project.
* Add them to the input section.

**Reference**

* Choose the reference database for mapping and annotation.
* The reference corresponds to a miRBase version.
* Default reference is typically **miRBase\_v22**.

**Species**

* Select the organism of interest.
* Supported model organisms are available in a dropdown list.
* Default species is **Human**.

**Analysis Output**

Choose the project where your analysis results should be saved.

### Running the Analysis

1. Review all configuration settings.
2. Click **Launch Application**.
3. Monitor progress from the **Runs** or **Analysis** dashboard.

#### **Launching the Illumina DRAGEN miRNA Published Pipeline with Demo Data**

To run DRAGEN miRNA using Illumina demo data, follow these steps:

1. **Add Demo Data:**
   * Navigate to the **Demo Data** tab in your BaseSpace Sequence Hub (BSSH) dashboard.
   * Search for "MiSeq i100: Illumina miRNA Prep."
   * Select and import the "Project", choose the context to accept the invite, and click **Accept**.
2. **Data Processing:**
   * The demo project contains data already processed through BCLConvert (FASTQ).
   * Proceed to analyze the FASTQs with DRAGEN miRNA by following the above steps.


# Analysis in ICA

### Overview

This tutorial guides you through running the Illumina DRAGEN miRNA pipeline using [Illumina Connected Analytics](https://help.ica.illumina.com/). You'll begin by setting up a project with demultiplexed instrument run data (FASTQ), adding pipelines and running a pipeline.

***

### Set-up <a href="#set-up" id="set-up"></a>

This tutorial assumes you already have an existing project in ICA. To create a new project, please see instructions in the [Projects](https://help.ica.illumina.com/home/h-projects#create-new-project) page.

Additionally, you will need the Illumina DRAGEN miRNA Bundle linked to your existing ICA project. The DRAGEN miRNA Demo Bundle is an an [entitled bundle](https://help.ica.illumina.com/home/h-bundles#access-an-entitled-bundle) provided by Illumina with all standard ICA subscriptions and includes DRAGEN pipelines, references, and demo data.

#### Linking an Bundle <a href="#linking-an-entitled-bundle" id="linking-an-entitled-bundle"></a>

For general steps on creating and linking bundles to your project, see the [Bundles](https://help.ica.illumina.com/home/h-bundles) page. This tutorial explores the DRAGEN miRNA Published Pipeline, so we will need to link the DRAGEN miRNA Bundle to our existing project.

Steps:

* Go to the **Projects > your\_project > Project settings > Details** page
* Click **Edit**
* Click the **+** button, under Linked bundles
* Select the *DRAGEN miRNA 1.0.0 Bundle*, followed by **Link**.
* Finally, **save** the change.

DRAGEN miRNA 1.0.0 Bundle assets will now be available on the *Data* and *Pipelines* pages.

### Pipelines <a href="#pipelines" id="pipelines"></a>

After setting up the project in ICA and linking a bundle, we can run various pipelines.

#### DRAGEN miRNA Published Pipeline <a href="#dragen-germline-published-pipeline" id="dragen-germline-published-pipeline"></a>

This example demonstrates how to run the DRAGEN\_miRNA\_1-0-0 Published Pipeline in your ICA project using the demo data from the linked DRAGEN miRNA 1.0.0 Bundle.

The required pipeline input assets for this tutorial include:

* Under **Projects > your\_project > Data**
  * Illumina DRAGEN miRNA Demo Data folder
* Under **Projects > your\_project > Flow > Pipelines**
  * Illumina DRAGEN miRNA

#### **Launching the Illumina DRAGEN miRNA Published Pipeline with Demo Data**

From the **Projects > your\_project > Flow > Pipelines** page, select Illumina DRAGEN\_miRNA\_1-0-&#x30;*,* and then click **Start analysis.** Initial set-up details require pipeline run name meaningful to the user and a *Subscription* from the drop-down menu under Pricing.

Running the Illumina DRAGEN miRNA pipeline uses the following inputs which are to be added in the Input Files section:

**FASTQ files**

Select the FASTQ files in the *Illumina DRAGEN Enrichment Demo Data* folder and select **Add**.

**Reference**

Select the reference database to be used for mapping and annotation. The database name indicates the collection of reference files used for small RNA classification and is based on the corresponding miRBase version. Default Reference is set to "miRBase\_v22".

**Species**

Select the species of interest. Libraries from six model organisms are available for miRNA classification. Default species is set to "Human".

**Resources**

Choose the expected storage space required for storing analysis results.


# Accessing Results

For information on tracking and viewing run and analysis results in BaseSpace Sequence Hub, refer to [View Data on Basespace](https://help.connected.illumina.com/basespace-sequence-hub/data/view-data).

To view results on ICA, you may either click on "View Files in ICA" in the top right corner of your BSSH Analysis page, or directly access the analysis in ICA. It will be in a BSSH managed project with the same name as your BSSH workgroup. For information on viewing analysis results on your Illumina Connected Analytics account, refer to [Viewing Data on ICA](https://help.connected.illumina.com/illumina-connected-analytics/project/p-data#viewing-data).


# Secondary Analysis Results

### Folder structure

Outputs are split in three main folders:

* counts
* fastqc
* reports

#### File and Folder Structure (`counts/`)

The `counts/` directory contains the following subfolders and files:

| `annotation_logs/`                          | Annotation log files                                                                                                                                                                                                                                                                                                                                                            |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `mapping_logs/`                             | Mapping log files                                                                                                                                                                                                                                                                                                                                                               |
| `<Sample_ID>/`                              | Run parameters for each sample (`<Sample_ID>`)                                                                                                                                                                                                                                                                                                                                  |
| `READs/`                                    | <p>Read count files by RNA type, e.g.:</p><ul><li><code>all\_samples.concatenated.allRNAs.READs.txt</code></li><li><code>all\_samples.miRNA.READs.txt</code></li><li><code>all\_samples.mRNA.READs.txt</code></li><li><code>all\_samples.otherRNA.READs.txt</code></li><li><code>all\_samples.piRNA.READs.txt</code></li><li><code>all\_samples.tRNA.READs.txt</code></li></ul> |
| `run_logs/`                                 | Logs from the counting step                                                                                                                                                                                                                                                                                                                                                     |
| `UMIs/`                                     | <p>UMI count files by RNA type, e.g.:</p><ul><li><code>all\_samples.concatenated.allRNAs.UMIs.txt</code></li><li><code>all\_samples.miRNA.UMIs.txt</code></li><li><code>all\_samples.mRNA.UMIs.txt</code></li><li><code>all\_samples.otherRNA.UMIs.txt</code></li><li><code>all\_samples.piRNA.UMIs.txt</code></li><li><code>all\_samples.tRNA.UMIs.txt</code></li></ul>        |
| `all_samples.notCharacterized_mappable.txt` | Metrics for uncharacterized mappable reads                                                                                                                                                                                                                                                                                                                                      |
| `all_samples.summary.txt`                   | Summary metrics for all samples (text format)                                                                                                                                                                                                                                                                                                                                   |
| `all_samples.V1.summary.xlsx`               | Summary metrics for all samples and counts files as tabs for all RNA types (Excel format)                                                                                                                                                                                                                                                                                       |
|                                             |                                                                                                                                                                                                                                                                                                                                                                                 |

The `fastqc/` directory contains subfolders and files generated by FastQC quality control analysis. Each sequencing sample has its own `<Sample_ID>`-specific outputs, including both an interactive HTML report and a tab-delimited metrics file.

| `<Sample_ID>_fastqc/`         | <p>Subfolder containing FastQC output files for a given sample.<br><strong>Example:</strong> <code>QK2\_S2\_L001\_R1\_001\_fastqc/</code></p>   |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `<Sample_ID>_fastqc_data.txt` | <p>Tab-delimited summary of FastQC metrics for a sample.<br><strong>Example:</strong> <code>QK2\_S2\_L001\_R1\_001\_fastqc\_data.txt</code></p> |
| `<Sample_ID>_fastqc.html`     | <p>Interactive FastQC report (HTML format) for a sample.<br><strong>Example:</strong> <code>QK2\_S2\_L001\_R1\_001\_fastqc.html</code></p>      |

The `report/` directory contains per-sample subfolders along with an aggregated HTML report generated by DRAGEN.

| `<Sample_ID>/`                               | <p>Subfolder containing DRAGEN output files for a given sample.<br><strong>Example:</strong> <code>QK2\_S2/</code>, <code>QL4\_S4/</code></p> |
| -------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `<Sample_ID>/<Sample_ID>.bowtie_metrics.csv` | sample level mapping metrics                                                                                                                  |
| `<Sample_ID>/<Sample_ID>.fastqc_metrics.csv` | sample level fastq quality metrics                                                                                                            |
| `DRAGEN_Reports.html`                        | Consolidated interactive report summarizing metrics across all samples.                                                                       |


# DRAGEN Report

The **DRAGEN Report** is generated as part of the **DRAGEN miRNA Software** and provides key quality-control information for miRNA sequencing runs.\
The report is organized into the following sections:

* **Experimental Summary**
* **FastQC**
* **Mapping**

### Experimental Summary

The *Experimental Summary* tab contains two tables that describe the run configuration and a general overview of the data.

#### General Run Settings

| **Library Prep Kit**   | Illumina library preparation kit used to generate miRNA and other small RNA sequencing libraries (e.g., *miRNA Illumina Prep Kit*). |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| **Reference Database** | miRNA reference database version used for alignment and quantification (e.g., *miRBase\_v21*).                                      |
| **Software Version**   | Version of the DRAGEN miRNA analysis pipeline used to process the FASTQ files (e.g., *1.0.0*).                                      |
| **Species Used**       | Organism reference used for mapping input RNA (e.g., *Human*).                                                                      |

***

#### General Summary

<table data-header-hidden><thead><tr><th valign="top"></th><th valign="top"></th></tr></thead><tbody><tr><td valign="top">Sample</td><td valign="top"><br></td></tr><tr><td valign="top">Input Reads</td><td valign="top">Total input reads</td></tr><tr><td valign="top">Mapped Reads</td><td valign="top">Total mapped reads</td></tr><tr><td valign="top">% Mapped</td><td valign="top">Percentage of mapped reads</td></tr><tr><td valign="top">Average Read Length</td><td valign="top">Estimated average read length per sample</td></tr><tr><td valign="top">Average @ Score</td><td valign="top">Estimated average Q score per sample</td></tr><tr><td valign="top">Average % GC Content</td><td valign="top">Estimated Average %GC content per sample</td></tr></tbody></table>

### FastQC

This tab displays the following graphs:

* ﻿﻿Read Quality Distribution
* ﻿﻿%GC Content
* ﻿﻿Ambiguous Base Content by Position

### Mapping

This tab displays the following graph and tables:

• Characterized Reads

This graph contains the percentage of total reads classified to a given small RNA type (e.g., miRNA, hairpin, RNA, tRNA, piRNA, mRNA, other RNA reads). The missing percentage is attributed to unmapped reads.

#### Mapping Summary

<table data-header-hidden><thead><tr><th valign="top"></th><th valign="top"></th></tr></thead><tbody><tr><td valign="top">Index</td><td valign="top">Sample index</td></tr><tr><td valign="top">Sample</td><td valign="top">Sample ID defined by fasta names</td></tr><tr><td valign="top">Input Reads</td><td valign="top">Total reads input both mapped and unmapped</td></tr><tr><td valign="top">Mapped Reads</td><td valign="top">Number of mapped reads</td></tr><tr><td valign="top">% Mapped</td><td valign="top">Percentage of mapped reads from total reads</td></tr><tr><td valign="top">Unmapped Reads</td><td valign="top">Number of unmapped reads</td></tr><tr><td valign="top">% Unmapped</td><td valign="top">Percentage of unmapped reads from total reads</td></tr></tbody></table>

#### Mapped Reads

<table data-header-hidden><thead><tr><th valign="top"></th><th valign="top"></th></tr></thead><tbody><tr><td valign="top">Index</td><td valign="top">Sample Index</td></tr><tr><td valign="top">Sample</td><td valign="top">Sample ID defined by fasta names</td></tr><tr><td valign="top">miRNA</td><td valign="top">Number of reads mapped to micro RNAS</td></tr><tr><td valign="top">rRNA</td><td valign="top">Number of reads mapped to ribosomal RNA</td></tr><tr><td valign="top">tRNA</td><td valign="top">Number of reads mapped to transfer RNA</td></tr><tr><td valign="top">mRNA</td><td valign="top"><p>Number of reads mapped to messenger</p><p>RNA</p></td></tr><tr><td valign="top">piRNA</td><td valign="top">Number of reads mapped to piwi-interacting RNA</td></tr><tr><td valign="top">Hairpin</td><td valign="top">Number of reads mapped to hairpin</td></tr><tr><td valign="top">Other RNA</td><td valign="top">Number of reads mapped to other RNA types than the ones described above</td></tr></tbody></table>

#### Unmapped Reads

<table data-header-hidden><thead><tr><th valign="top"></th><th valign="top"></th></tr></thead><tbody><tr><td valign="top">Index</td><td valign="top">Sample Index</td></tr><tr><td valign="top">Sample</td><td valign="top">Sample ID</td></tr><tr><td valign="top">Too Short</td><td valign="top"><p>Reads with a length shorter than 16 nucleotides after adapter trimming with</p><p>Cutadapt</p></td></tr><tr><td valign="top">No Adapter</td><td valign="top"><p>Reads containing no adapter sequence</p><p>(AACTGTAGGCACCATCAAT)</p></td></tr><tr><td valign="top">UMI defective</td><td valign="top">Reads lacking a unique molecular identifier are defective or missing</td></tr><tr><td valign="top">Uncharacterized Mappable</td><td valign="top">During multiple alignment processes, reads are assigned their corresponding RNA identities. Uncharacterized mappable reads refer to those not classified as a specific RNA type but can be aligned to the reference genome</td></tr><tr><td valign="top">Uncharacterized Unmappable</td><td valign="top">Uncharacterized and unmappable RNA refers to RNA sequences that cannot be assigned to a specific RNA category or aligned with a reference genome</td></tr></tbody></table>


# Demo Data

For more information about this product, check the Demo Data tab on the [BSSH website](https://ilmn.basespace.illumina.com/datacentral). You can import a sequencing run, containing raw sequencing data, or a project with secondary analysis results.

The [MiSeq i100 Illumina miRNA prep](https://ilmn.basespace.illumina.com/datacentral) Demo Data includes samples prepared using the Illumina miRNA sequencing library preparation protocol and sequenced on [MiSeq i100](https://www.illumina.com/systems/sequencing-platforms/miseq-i100.html).

The [MiSeq i100 Illumina miRNA prep](https://ilmn.basespace.illumina.com/datacentral) Demo data is also available within Illumina Connected Multiomics via a bundle, please refer to[ Analysis in ICA](broken://spaces/S2VRwfbiFY1ysT1i6Ron/pages/IWxg1BkOIwy7JkmeLdYb) to learn how to link this demo data to your ICA project.

### Experimental Details

The [MiSeq i100 Illumina miRNA prep](https://ilmn.basespace.illumina.com/datacentral) Demo Data is composed of replicates of human brain total RNA (HBTR), human liver total RNA (HLTR), and total RNA isolated from human whole blood were prepared using the Illumina miRNA prep kit at 100ng input. The libraries were sequenced on MiSeq i100 with 100M flowcell using a single read at 72bp read length with dual indexing. The samples were analyzed with the Basespace/ICA Dragen miRNA app.

### Tertiary Analysis

After successfully generating data with Illumina Sequencers and completing secondary analysis with the DRAGEN miRNA app, you can gain further insights through the [Illumina Connected Multiomics](https://www.illumina.com/products/by-type/informatics-products/connected-multiomics.html).

In the following session, there will be instructions on how to analyze the miRNA data further, going from the miRNA counts matrix to differential expression analysis. To do so, users shall input information about their data through a metadata file.

### Creating a Metadata file

The metadata file, is a .tsv or .csv file that shall contain the following mandatory column:<br>

* "SampleID": containing the sample names as written in the count matrix file

Users can then add extra columns that describes samples such as "Phenotype", "Sex", or "Status".

The following table illustrates the metadata format:

| SampleID | Status  | Sex    |
| -------- | ------- | ------ |
| SampleA  | Disease | Female |
| SampleB  | Disease | Male   |
| SampleC  | Control | Female |


# Illumina Connected Multiomics

Illumina Connected Multiomics (ICM) is available for tertiary analysis of miRNA and other multiomic data.

### Getting Started

### [Logging into Connected Multiomics](https://help.multiomics.illumina.com/icm/introduction/readme-1#log-in-to-connected-multiomics)

### [Creating a study and adding projects from Illumina Connected Analytics](https://help.multiomics.illumina.com/icm/studies/create-study)

### [Viewing results and navigating in Connected Multiomics](https://help.multiomics.illumina.com/icm/analyses/enter-analysis)

### Data / Task nodes and Performing tasks in Connected Multiomics

* Within a study, the *Analyses tab* contains two elements: task nodes (rectangles) and data nodes (circles) connected by lines and arrows. Collectively, they represent a data analysis pipeline.
* Clicking a data node brings up a context sensitive menu on the right. This menu changes depending on the type of data node. It will only present tasks which can be performed on that specific data type. Hover over the task to obtain additional information regarding each option.
* Select the task you wish to perform from the menu. When configuring task options, additional information regarding each option is available. Click **Finish** to perform the task.
* Depending on the task, a new data node may automatically be created and connected to the original data node. This contains the data resulting from the task. Tasks that do not produce new data types will not produce an additional data node.
* To view the results of a task, click the data node and choose the **Task report** option on the menu.

### Viewing and saving data

* All data contained in data nodes can be downloaded to the local machine by selecting the node and navigating to the bottom of the toolbox then choose **Download data**.
* The [Data Viewer](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/BEpNAu7pFMMuYvu7XKHz) can be used to plot, modify, and save data. In this walkthrough the PCA data node and Hierarchical clustering / heatmap node can be automatically opened in the Data viewer by double-clicking the data node or opening the Task report from the toolbox.
* To save an individual image within the Data Viewer to your machine, click **Plot** then **Export image** & select the format, size, and resolution then click **Save**. Use the plot-specific tools for this.
* All visualizations within a sheet in the Data Viewer can be exported as one image (e.g. use one image with all plots for a poster). Use the **Export** drop-down at the top of the data-viewer for this and select **Export image**.

## Import Data

microRNA data generated from secondary analysis contains sample by miRNA raw count matrix in a .txt file format. The demo data contains 6 samples from 3 different tissue types, 2 samples in each tissue type: brain, blood and liver. The libraries were sequenced on MiSeq i100 with 100M flowcell using a single read at 72bp read length with dual indexing. The samples were analyzed with the Basespace DRAGEN miRNA app and using hg38 assembly and miRBase reference v22.

First, create a study to upload data.

* Click **+ New Study**

  <figure><img src="/files/dN7E9LhZWbzVlvCfV4RM" alt=""><figcaption><p>Create a new study</p></figcaption></figure>
* Add **Study Name** and **Description**
* Click **Create**
* Click **+ Add Data**

<figure><img src="/files/HL4YD1ZESoRO8jB6vkbk" alt=""><figcaption><p>Add data to the study</p></figcaption></figure>

* Choose **Select from ICA project**
* Choose **Bulk > miRNA**
* Click the **+ Add Demo Data** button
* Select the '*Multiomics-Demo-Data'* folder
* Select the '*Transcriptomics'* folder
* Select the *'ILMN miRNA Demo data'* folder
* Check the '*all\_samples.miRNA.UMIs.metadata.csv'* and '*all\_samples.miRNA.UMIs.txt'* files
* Click **Add selected data to your study**
* After creating a study and adding data to the study, click **+ New Analysis**
* Give the Analysis a name, select the **Analysis Type > Custom: Illumina miRNA**, use all Illumina miRNA Prep samples, which is the default, and click **Run Analysis**

<figure><img src="/files/uy01FYyPYjmy1eaNCFm1" alt=""><figcaption></figcaption></figure>

* The Status will show as Complete once it is finished.

{% hint style="warning" %}
Click the Refresh button to see change the Status in real-time.
{% endhint %}

<figure><img src="/files/L9dwJzhtxPH5ef6IID6r" alt=""><figcaption></figcaption></figure>

* Click on the name of the analysis to open the task graph, a data node circle labled as miRNA is displayed

<figure><img src="/files/Ijy5TzRuTx2uFrXqxpZo" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Hover over nodes to see details about the data node. Below, the number of samples, features, and data size is shown in the node.
{% endhint %}

Single-click on miRNA node, choose **Imported count matrix report** in QA/QC section to check the distribution of the data. After the task is finished, double click on output report to view it.

<figure><img src="/files/9loXEY7EUmiGk8jAGCsf" alt=""><figcaption></figcaption></figure>

In the table, each row is a sample, columns are descriptive statistics metrics of the miRNA count information. It is recommended to filter out low expression miRNAs to reduce false positive in the downstream analysis, the criteria of low expression depends on the data, this table is helpful to make a decision, e.g. the median value of the matrix is 13.5, you might want to use this value as noise background value to perform low expression filter for the first iteration of downstream analysis.

## Filter Low Expression miRNA

Click on the miRNA data node, choose **Filtering > Filter features,** select **Noise reduction** filter type which is default, exclude features where max is <=13.5, click **Finish**.

<figure><img src="/files/HbotZY7dRX2WPA51pLIs" alt="" width="563"><figcaption></figcaption></figure>

Mouse over the output Filtered counts data node to check how many features in the filtered data node. If the number is too low, you might want to redo the filter and use a more lenient filter criteria.

<figure><img src="/files/vJEA7U2vJe1Px6PG7IWT" alt=""><figcaption></figcaption></figure>

## Annotate Features (Optional)

If you want to get genomic location information on the miRNA, you can link the data to the miRBase annotation.

* Single-click the *miRNA* node.
* Select **Annotate features** under the *Pre-analysis tools* section in the toolbox on the right.
* Choose the **genome** and **annotation** files that match those used in DRAGEN then click **Finish**. This data is using hg38 miRBase mature microRNAs v22

<figure><img src="/files/gUMi2KEZwsGZCm8oMpCN" alt=""><figcaption></figcaption></figure>

In this example, we skip this step since genomic location of miRNA is not used for the downstream analysis.

## Normalization and Scaling

Normalize the data to prepare for downstream analysis.

* Single-click the **Filtered counts** node, then select the **Normalization** task from the *Normalization and Scaling* section.
* Click the **"Use Recommended"** button or select an alternative method. We recommend the widely used *Median ratio (DESeq2 only)* method.

<figure><img src="/files/2LL8JqvBdovfUyOLbHfd" alt=""><figcaption></figcaption></figure>

The task output Normalized counts node.

## Principal Component Analysis (PCA)

Visualize sample clustering and variance.

* From the *Normalized Counts* node, select **PCA** under *Exploratory Analysis*.
* Use default settings in the dialog, click **Finish**.

<div align="left"><figure><img src="/files/CN4O9XGGpdqkSPFplIme" alt="" width="473"><figcaption></figcaption></figure></div>

Double click on the PCA node to view the scatterplot in Data Viewer

<figure><img src="/files/7YIb4GtkRfCK4Pbl2MNa" alt=""><figcaption></figcaption></figure>

When color the dots with TissueType, we see separate clusters based on different tissue types, and blood samples are more different from the other two tissues.

## Differential Analysis

Compare miRNA expression across experimental groups.

* From the *Normalized counts* node, select **Differential Analysis** from the *Statistics* section.
* Choose your preferred model and set up the comparison. Note that we have chosen the [DESeq2 method](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/71qP0wdPqQaJ9aO0DnP2) and used the corresponding normalization prior.
* Select **TissueType** to **Add factors** and click **Next**

<div align="left"><figure><img src="/files/2RpLB8oMZmmW0jZJcvFA" alt="" width="248"><figcaption></figcaption></figure></div>

Set up the three comparisons and leave other settings as default, click **Finish**

<div align="left"><figure><img src="/files/9eRgbmlh3dWMqUEJv210" alt="" width="563"><figcaption></figcaption></figure></div>

DESeq2 report is generated, double click on it to open the report:

<div align="left"><figure><img src="/files/PtCl229rOh649j4VnHcy" alt="" width="563"><figcaption></figcaption></figure></div>

## Filter Feature List

Refine the list of genes/features based on criteria.

* On filter panel, specify filter criteria to generate significant miRNA lists
* Here is an example of generate a list of miRNA that are significantly different between liver vs blood based on FDR <=0.05, fold change up and down 2 fold.
* Check **FDR step up** section, choose **Per contrast** option
* Specify live vs blood value FDR <=0.05, you can uncheck other contrasts, or leave them unchanged, since default FDR <=1 is not filtering anything
* Check **Fold change** section, choose **Per contrast** option
* Specify exclude range from -2 to 2 for only liver vs blood
* Click **Generate filtered node**

<div align="left"><figure><img src="/files/urM6mOHDdaEo4biMKrfw" alt=""><figcaption></figcaption></figure></div>

*Option: specify similar criteria to generate significant feature list of liver vs brain and blood vs brain.*

## Hierarchical clustering / Heatmap

Visualize the significant miRNA list using heatmap. Select the filtered feature list based on the criteria for liver vs blood,

* Select **Hierarchical clustering / Heatmap** from the *Exploratory analysis* section.
* Filter only include TissueType in **blood** and **liver**, leave everything else as default, click **Finish**

<div align="left"><figure><img src="/files/pJYAwwAiHeSVCzIcmt0s" alt=""><figcaption></figcaption></figure></div>

* Double-click on the output node to visualize the results in the Data viewer.

<div align="left"><figure><img src="/files/8x4tNVybqwuEET7G6doj" alt=""><figcaption></figcaption></figure></div>

## Generate Targeted mRNA List

Click on a filtered feature list, choose miRNA integration > **Get targeted mRNA** on the pop-up menu, use TargetScan database, click **Finish.**

<div align="left"><figure><img src="/files/5K7zIoOqx4k4uJG1GA0V" alt=""><figcaption></figcaption></figure></div>

The task generates a Targeted mRNA node, double click on it to open the report:

<div align="left"><figure><img src="/files/6dWrYaGdNqatFK2LH1xZ" alt=""><figcaption></figcaption></figure></div>

* Context++ score: estimating the strength of repression, more negative values suggest stronger predicted repression, this is the key metric for ranking predicted targets
* Context++ score percentile: percentile ranking of the score compared to all predictions, higher percentile means stronger predicted effect
* Weighted Context++ score: adjusted score considering multiple sites in the same transcript
* Predicted relative KD: predicted relative dissociation constant is to measure the affinity between miRNA and targeted mRNA, lower KD means stronger binding affinity, higher KD means weaker binding

## Gene Set Enrichment

Identify enriched biological pathways or gene sets. Click on Targeted mRNA node

* Select **Gene Set Enrichment** from the *Biological Interpretation* section.
* Choose between **KEGG Pathway Enrichment** or **Gene Set Ontology,** click **Finish**
* Double click on the output *Pathway enrichment* node to open the report
* Filter the report based e.g. Enrichment score > 5, when the table has less than 100 rows, Data viewer plots are enabled

<figure><img src="/files/oLeJrSiHIZciYWdrrqag" alt=""><figcaption></figcaption></figure>

* Click on **View plots in Data Viewer** to display the pathways in barchart and scatterplot

<figure><img src="/files/OFWGgvJkNGNBslBHXzK9" alt=""><figcaption></figcaption></figure>


# Common Questions

* **Q1. My run failed; how can I troubleshoot it**?

  You can access the "Failed steps details" link to review the ICA Analysis Failed Step Logs:

<figure><img src="/files/To0wGoUXh3nm310jAj5L" alt=""><figcaption></figcaption></figure>

* **Q2. Should I trim adapters beforehand?**

  In the E2E workflow, Truseq Adapter trimming is managed by BCL Convert when using a Run Planner-configured sample sheet (when Illumina miRNA prep is selected). Meanwhile, the DRAGEN miRNA software automatically trims the 3' UMI adapter during its process. If the Truseq Adapter is not trimmed before analysis, it does not affect the actual analysis but alters the final read statistics.
* **Q3. What are the most common errors?**
  1. Sample Names not following Illumina’s name convention
  2. Memory issues at the counting step (fixed during SW verification)
  3. Sample sheet configuration mistakes (best to use BSSH Run Planning)
* **Q4. Can DRAGEN miRNA be used to analyze TruSeq Small RNA data?**
  * The software was not designed to analyze TruSeq Small RNA and analysis will fail.
* **Q5. How were reference sequences for small RNAs obtained?**
  * Check the [Reference Database](/dragen-mirna/readme/reference-database) page for further information
* **Q6. What are the software component versions used in DRAGEN miRNA 1.0.0?**
  * Reference database: miRBase v21 and v22 available
  * Analysis Pipeline Version: 1.0.0
  * DRAGEN Report Version: 4.5.0-alpha.7
  * Cutadapt Version: 1.10
  * Bowtie Version: 1.1.2
* **Q7. What species will the analysis support?**

  Supports the following species:

  * Human
  * Mouse
  * Rat
  * Zebrafish
  * Nematode
  * Fruitfly

  MiRNA references are based upon the version v21 and v22 of the miRBase, customers can choose between the two versions when running an analysis
* **Q8. How do I demultiplex my samples?**
  * One can use the BCLConvert app available in BSSH. Check the following page for further information: <https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/bcl-convert.html>


# Known Issues and Limitations

## DRAGEN miRNA 1.0.0

### Known Issues​<br>

* Sample-level metrics: DRAGEN Reports currently do not display sequencing quality graphs at the sample level.
* UMI support: DRAGEN Reports currently display metrics for Reads, but not for Unique Molecular Identifiers (UMIs).​ Please refer to the UMI output files (e.g., all\_samples.miRNA.UMIs.txt) for UMI-specific information.​
* Summary file counts: In the files all\_samples.V1.summary.xlsx and all\_samples.V1.summary.txt, the reported total reads​ for each small RNA type (e.g., miRNA\_Reads, hairpin\_Reads, piRNA\_Reads, rRNA\_Reads, tRNA\_Reads, mRNA\_Reads) represent​ the sum of all reads mapped to that RNA category. However, only small RNAs with at least 10 read matches are listed in​ the individual small RNA output files.•

### Known Limitations​ ​

* RNA detection is restricted to the reference sequences provided by the software for the supported model organisms.​ Custom or user-supplied references are not supported.​
* Only RNAs included in the reference database are detected; the software does not support discovery or prediction of​ novel or unannotated RNAs.


# About Connected Multiomics

**Illumina Connected Multiomics** is a cloud-based software platform designed for biologists and bioinformaticians to perform tertiary analysis of multiomics data for research purposes. It enables users to organize and manage biological data into studies, apply statistical methods, reference knowledge sources for biological interpretation, correlate results across various omics data types and modalities, and use interactive visualizations to interpret results. This functionality streamlines the process from samples to multiomics insights, accelerating data interpretation and publication efforts.

{% embed url="<https://youtu.be/OHQfIJLikD8?si=HrJiR_MIVEWqqeN9>" %}

Explore our [Interactive Demos](/icm/about/interactive-demos) for an introductory experience of available Multiomics workflows.

{% hint style="warning" %}
You can search our help documentation or ask questions with AI-generated answers using the search-box in the top of the page.

Navigate and explore using the left-panel.
{% endhint %}

### Plans and Pricing

Illumina Connected Multiomics offers two subscription tiers:

{% stepper %}
{% step %}

#### Illumina BioInsight Platform (formerly Connected Software) Basic

Includes preconfigured analysis workflows supporting select Illumina assays and is free to use. Requires sufficient [BioInsight Credits](/icm/reference/icredits) to meet data storage and analysis needs based on throughput.

<figure><img src="/files/sadrF5I1RUc4TANZYgu0" alt=""><figcaption><p>Default analysis for single cell analysis</p></figcaption></figure>
{% endstep %}

{% step %}

#### Illumina Connected Multiomics Professional

Includes preconfigured analysis workflows supporting select Illumina assays and the ability to create new custom workflows for key Illumina multiomic assays and select third-party assays. Pathway analysis is enabled with the integration of Illumina Correlation Engine or KEGG Pathway. Requires purchase of an annual or monthly license and sufficient [BioInsight Credits](/icm/reference/icredits) to meet data storage and analysis needs based on throughput.

<figure><img src="/files/Q0NKldmiyHgKwjmVCtIs" alt=""><figcaption><p>Example custom workflow for single cell analysis</p></figcaption></figure>
{% endstep %}
{% endstepper %}

For more information about supported analysis, see [Analysis Options](/icm/studies/view-studies/view-studies-2#analysis-options) and [Supported Data Types](/icm/reference/supported-data-types).

### Supported Regions

Connected Multiomics cloud infrastructure currently is deployed in the following [regions](https://aws.amazon.com/about-aws/global-infrastructure/regions_az/):

| Regional Code | Name                     | Geography      |
| ------------- | ------------------------ | -------------- |
| apn2          | Asia Pacific (Seoul)     | South Korea    |
| aps1          | Asia Pacific (Singapore) | Singapore      |
| aps2          | Asia Pacific (Sydney)    | Australia      |
| cac1          | Canada (Central)         | Canada         |
| euc1          | Europe (Frankfurt)       | Germany        |
| euw2          | Europe (London)          | United Kingdom |
| use1          | US East (N. Virginia)    | United States  |

### Additional Resources

Additional resources for Connected Multiomics:

| Link                                                                                                         | Description                                                                                                |
| ------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- |
| [Product Page](https://www.illumina.com/products/by-type/informatics-products/connected-multiomics.html)     | Learn more about Connected Multiomics including key features and product resources.                        |
| [Order Page](https://www.illumina.com/products/by-type/informatics-products/connected-multiomics/order.html) | Order Connected Multiomics and see data sheet, overview information, FAQs and available plans and pricing. |
| [Video playlist](https://youtube.com/playlist?list=PLKRu7cmBQlahGuzvRfL41txLRPAV3e2uZ\&si=8FMGNzdAAYdTv0Lu)  | Access all available how to and demo videos in a playlist.                                                 |
| [News and Updates](https://developer.illumina.com/illumina-connected-multiomics)                             | Latest blog posts, news, and other updates about Connected Multiomics.                                     |
| [How-to Videos and Webinars](/icm/reference/how-to-videos)                                                   | How-to videos from the video playlist and recorded webinars are linked on page.                            |

{% hint style="warning" %}
Illumina offers paid training to making getting started with Connected Multiomics easy.

Training includes domain registration/setup, data transfer and ingestion, analysis launch, overview of statistical tools/ workflows, and workflow demos. Includes (4) hours of product training delivered virtually.
{% endhint %}

### Additional Multiomics Software

**Illumina Multiomics Software** offers a comprehensive range of solutions to streamline your journey from sample collection to multiomic insights. Discover assays designed for protein analysis, spatial data, single-cell RNA data, and more.

Our powerful study management tool helps organize your data from ingestion to analysis, while our advanced analysis platform generates actionable insights to drive your research forward.

Select an application below to learn more.

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>Provides normalized protein counts for NGS-based proteomics</td><td><a href="/files/kqiVtgvBYQHM7uKZW0s0">/files/kqiVtgvBYQHM7uKZW0s0</a></td><td><a href="/spaces/nyQb4WG1K4VQZKovLv5v">/spaces/nyQb4WG1K4VQZKovLv5v</a></td></tr><tr><td>Processes single-cell RNA-Seq data into gene expression matrices</td><td><a href="/files/foXfh9rR5bMe0csnNBX2">/files/foXfh9rR5bMe0csnNBX2</a></td><td><a href="/spaces/qVEYIKB8JFfdScsTocFN">/spaces/qVEYIKB8JFfdScsTocFN</a></td></tr><tr><td>Generates methylation and genomic variant calls on 5-base DNA assay data</td><td><a href="/files/LwOk9dy1r3JnOSngIJk3">/files/LwOk9dy1r3JnOSngIJk3</a></td><td><a href="/spaces/L2tTN7buOERM9NKPYYlg">/spaces/L2tTN7buOERM9NKPYYlg</a></td></tr><tr><td>Transforms miRNA data into small RNA count matrices</td><td><a href="/files/6Yy4Pt4mtVOaI0gNOzbE">/files/6Yy4Pt4mtVOaI0gNOzbE</a></td><td><a href="/spaces/RvUv32d8wOzBh0okiOjK">/spaces/RvUv32d8wOzBh0okiOjK</a></td></tr><tr><td>Links experimental results to curated studies, pathways, and clinical data for enhanced biological interpretation</td><td><a href="/files/2eKFKkHaMzlwtbACUk0C">/files/2eKFKkHaMzlwtbACUk0C</a></td><td><a href="/spaces/WolK2fDgPCCcGurM8A73">/spaces/WolK2fDgPCCcGurM8A73</a></td></tr></tbody></table>

{% embed url="<https://www.youtube.com/watch?v=Xsyp0qqKzkY>" %}


# Interactive Demos

Explore multiomics data with an interactive click-through demo of Connected Multiomics. Click <img src="/files/iJeMI2rEqn1wO78WkOzG" alt="" data-size="line"> to expand the selection for easier viewing.

## Illumina Proteomics

{% @storylane/embed subdomain="app" url="<https://app.storylane.io/demo/ar9lbsmvluks>" linkValue="ar9lbsmvluks" %}

## Illumina 5-base DNA

{% @storylane/embed subdomain="app" url="<https://app.storylane.io/share/gwnf17hfwc6v>" linkValue="gwnf17hfwc6v" %}

## Illumina Single Cell

{% @storylane/embed subdomain="app" url="<https://app.storylane.io/share/0tuoiwjfsmxj>" linkValue="0tuoiwjfsmxj" %}


# Getting Started

During the software registration process, you will create a domain and workgroup. [**Domains** ](https://help.connected.illumina.com/account-management/admin-console/domain)and [**Workgroups**](https://help.connected.illumina.com/account-management/admin-console/workgroups) are used by Illumina Software to control access to different customer’s data and assets. Make sure all users are added to the workgroup and that the workgroup has the necessary permissions to access the Connected Multiomics application, the users need to have the subscription of the software.

Here are the details of the steps to register in video format and as a step-by-step walkthrough:

{% embed url="<https://youtu.be/CEVA0wkbB0U?si=dFRyLwLd73HRCJWe>" %}

<details>

<summary>Step 1: Domain Registration</summary>

To register your Illumina Connected Software, click the registration link in the e-mail provided by Illumina.

<div align="center"><figure><img src="/files/2V1FtuZtcN6nAmHsWoD1" alt="" width="535"><figcaption></figcaption></figure></div>

This will bring you to the Illumina Software Registration Portal. If you do not have an Illumina Connected Software account, click **Sign Up** to create a new account.

<div align="center"><figure><img src="/files/VsyxbxPJel2rc42BHs1r" alt="" width="504"><figcaption></figcaption></figure></div>

If you already have a user account for Illumina Connected Software, enter those credentials into the login screen.

<figure><img src="/files/X1LIdYKpb8ytuB08YF7F" alt="" width="503"><figcaption></figcaption></figure>

Once logged in, you can review the order details for each order. Select your order for Illumina Connected Multiomics and click **Setup**. Multiple Connected Multiomics subscriptions can be selected and registered at one time.

<figure><img src="/files/KW700tuavxi9cIGgZP2r" alt="" width="563"><figcaption></figcaption></figure>

You will need to select a region and create a domain to which the software will be registered. Choose an existing domain or create a new domain for your account.

<figure><img src="/files/8NUacaOlkGzCTwFYHsri" alt="" width="563"><figcaption></figcaption></figure>

When creating a new domain, you will be considered the administrator for that domain. Type a display name for the domain, as well as the domain URL. Also enter specific e-mail addresses or e-mail extensions that are permitted in this domain. The list of allowed emails and e-mail extensions can be updated later in the Admin Console.

Click **Setup** to complete order registration.

<figure><img src="/files/Vp8J7T7Gzw2rcIMattMf" alt="" width="545"><figcaption></figcaption></figure>

A confirmation e-mail will be sent once the domain URL is activated. In the confirmation e-mail, access your domain by clicking on the domain URL link provided.

<figure><img src="/files/cXzwh9QjNWCwRxAUy32g" alt="" width="540"><figcaption></figcaption></figure>

Adding the link to your browser's bookmark is recommended for easier access in the future.

Sign in to your domain by entering your credentials into the login screen. Once signed in, the product dashboard will display all Illumina Connected software platforms you have access to. The Admin Console is a platform designed for administrators to manage domain access and control workgroup permissions.

</details>

<details>

<summary>Step 2: Workgroup Creation (Optional)</summary>

Workgroups are groups of users that can share projects and data. If you do not create a workgroup, you can get started with analysis in your personal context; however, collaborators will not be able to view and contribute to your analysis.

In product dashboard, click on **Admin** on the left panel to go to Admin Console

<figure><img src="/files/ilVygYXswodC1eNDybZ0" alt="" width="540"><figcaption></figcaption></figure>

Click on **Create Workgroup** in the top right corner. Workgroups can only be created by domain administrators.

<figure><img src="/files/VuxgQ4h4g4IYpqcKJz2M" alt=""><figcaption></figcaption></figure>

Then specify a workgroup name, as well as a description. Type in *Administrator Email* and click **Create workgroup**.

Note: You can allow users outside this domain to be added in a workgroup if you select *Allow collaborators outside of this domain* option. This allows you to invite Illumina's support scientists to your workgroup for assistance. The setting cannot be updated at a later time.

<figure><img src="/files/d46JwKfXq9dReEHxWArT" alt=""><figcaption></figcaption></figure>

Any domain user can be assigned to be the workgroup administrator. The workgroup administrator can manage workgroup membership by adding or removing domain users.

If you wish to add a collaborator to a workgroup, that collaborator must first be added as a user in that domain.

</details>

<details>

<summary>Step 3: Add Users (Optional)</summary>

To invite new users to a domain, click on the **Domain** tab on the left-hand side of the *Admin Console*. Then click **User management** on the left panel, followed by **Invite**.

<figure><img src="/files/7E7CWiN4Tn8YyuhOqoa9" alt=""><figcaption></figcaption></figure>

Type in the e-mail address of the user you wish to invite to the domain, then click **Invite to domain**. An invitation e-mail will be sent to the invited user.

The user should follow the instructions in the email to register an account with the domain.

<figure><img src="/files/yNA3faKVwDJ1BB5IMFO6" alt="" width="508"><figcaption></figcaption></figure>

The domain administrator can verify the successful registration of the invited user in the Admin Console. Click on the **Domain** tab on left-hand side, select **User management** on the left panel, click **Users** tab in the *User Management* page. Click the <img src="/files/zA0mVHNFjU658NLYHs0j" alt="" data-size="line"> info icon. The user status information, such as Active, will be displayed at the top right corner of the page.

<figure><img src="/files/TPzECXToTlr15Oe0xmNg" alt=""><figcaption></figcaption></figure>

The newly registered user can now be added as a collaborator in a workgroup.

To add the registered user to a workgroup, navigate to the **Workgroups** tab on the left panel of the Admin Console.

<figure><img src="/files/FhwzfJieNL4enxxA8XcM" alt=""><figcaption></figcaption></figure>

Select the workgroup you want to add users. Then in the **Users** tab, click **Invite**.

<figure><img src="/files/UPafnMneN8nJg3rrhskV" alt=""><figcaption></figcaption></figure>

Type in the user's email and specify product access. Click **Invite to workgroup**.

Note: The options to invite via current domain, public (accounts) domain or collaborative domain only appear if the Workgroup was setup as collaborative. These can be used to invite Illumina support or other outside email addresses to your domain. This setting cannot be changed afterwards.

<div data-with-frame="true"><figure><img src="/files/OYLglXHGjxVxl25Xnz9U" alt="" width="375"><figcaption></figcaption></figure></div>

The users added to this workgroup will be listed on the *Users* page.

</details>

<details>

<summary>Step 4: Subscription Assignment</summary>

For Illumina Connected Multiomics Professional licenses, users need to be assigned to its subscription in order to access the associated application.

To give a user access to Illumina Connected Multiomics for example, go to the product dashboard, click on **Subscriptions** on the left side panel.

At the Subscriptions panel, look for Illumina Connected Multiomics and click **Assign**.

<figure><img src="/files/KairWHcEVO8klvbvp6Tz" alt="" width="563"><figcaption></figcaption></figure>

Only domain administrators can assign subscription. Enter the e-mail address of a registered domain user that you wish to assign the subscription to and click **Assign.**

<figure><img src="/files/cLH5Lj8Z5vJrJdYZ4yGG" alt="" width="563"><figcaption></figcaption></figure>

Once a user is assigned to a domain and workgroup with necessary permission, the user will be able to perform data analysis in Connected Multiomics.

If the current user wanted to reassign the seat, the subscription page can be found directly from the Connected Multiomics application by clicking the 'Subscriptions' at the person icon in the top right-hand corner.

<figure><img src="/files/pNh9FfbndyQ0nIBr1lJd" alt=""><figcaption></figcaption></figure>

</details>

## Log In To Connected Multiomics

Once you have successfully set up your account and created your domain and workgroup, you can log in to Connected Multiomics. When logging into Connected Multiomics, you will be required to select a domain and workgroup to establish the context for your work. Follow the steps below.

{% stepper %}
{% step %}
Access the software by clicking this link: login.illumina.com/login

Adding the URL link to your browser bookmark is recommended for easy access in the future.
{% endstep %}

{% step %}
Log into your domain by entering your credentials into the login screen, then click Sign In.

<p align="center"><img src="/files/7SyYyM6C621bnMK4mIIn" alt=""><br></p>
{% endstep %}

{% step %}
Once signed in, you will be directed to Product Dashboard page. Select the Illumina Connected Multiomics application tile from your Product Dashboard.

<figure><img src="/files/oH1sjF7Xz1RhWVMSRUX6" alt="" width="563"><figcaption></figcaption></figure>
{% endstep %}

{% step %}
Upon first login, if one workgroup is available, the default will be the workgroup context. If more than one or no workgroups are available, the default will be personal context.

Once you select a workgroup, it will automatically be chosen during future logins. This selected workgroup will not be applied to other Illumina Connected Software applications. If you have only one workgroup, it will be selected for you by default.
{% endstep %}
{% endstepper %}

You will now be directed to the Studies overview screen. In the top right, you can click on your profile button to see your current workgroup or personal context, domain, and subscription information. To change workgroups, you can click on the workgroup name and select **Switch workgroup** at the bottom of the dropdown.

<figure><img src="/files/QvypoYy97zleWrdoyyfq" alt="" width="304"><figcaption></figcaption></figure>

## Explore Onboarding Resources

When you log into Connected Multiomics for the first time, you will be greeted with a **Welcome Screen**.

<div align="center"><img src="https://help.multiomics.illumina.com/~gitbook/image?url=https%3A%2F%2F580316046-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FWMxqQAMFOJtu98OBk9KN%252Fuploads%252Fgit-blob-b51fc0db536f4659636a9f6deb6c8dfbfcb12a74%252Ficm-welcome.png%3Falt%3Dmedia&#x26;width=768&#x26;dpr=4&#x26;quality=100&#x26;sign=b2634c15&#x26;sv=2" alt=""></div>

Click **Get started** to take you directly to the home page of the application and begin navigating on your own.

To explore a read-only interactive environment with pre-assembled studies, analyses, results, and visualizations, choose **Try with demo data** to bring you to the Connected Multiomics **Tutorial Study**. The source demo files used in the Tutorial Study are available to add to and use in [your own Studies](https://help.multiomics.illumina.com/icm/studies/enter-study/overview) later, too.

<div align="center"><img src="https://help.multiomics.illumina.com/~gitbook/image?url=https%3A%2F%2F580316046-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FWMxqQAMFOJtu98OBk9KN%252Fuploads%252Fgit-blob-916c1df11b28eb5e265ed390f523340a05c6c258%252FICM-tutorial-study.png%3Falt%3Dmedia&#x26;width=768&#x26;dpr=4&#x26;quality=100&#x26;sign=f2a294ee&#x26;sv=2" alt=""></div>

Select **Help resources** for easy links to user guide content, including a helpful query box, detailed walkthrough tutorials, instructional videos, and software release notes.

<div align="center"><img src="https://help.multiomics.illumina.com/~gitbook/image?url=https%3A%2F%2F580316046-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FWMxqQAMFOJtu98OBk9KN%252Fuploads%252Fgit-blob-2dad6439f8156d1e8d36912461ec13c6be086acb%252FICM-help-resources.png%3Falt%3Dmedia&#x26;width=300&#x26;dpr=4&#x26;quality=100&#x26;sign=ba467b1d&#x26;sv=2" alt=""></div>

You may return to the Welcome Screen and these selections anytime you'd like by selecting the ![](https://help.multiomics.illumina.com/~gitbook/image?url=https%3A%2F%2F580316046-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FWMxqQAMFOJtu98OBk9KN%252Fuploads%252Fgit-blob-d17d5ab19b6a3f0a3454daa5bcd2dceeb42d1e40%252FICM-WalkMe.png%3Falt%3Dmedia\&width=300\&dpr=4\&quality=100\&sign=d8f07516\&sv=2) icon on the top right of the application window next to your profile.

## Navigate Connected Multiomics

In the Connected Multiomics navigation on the left panel, you'll see the following structure:

* **Studies** are your primary work locations which contain your data and tools to execute your analyses. Studies can be considered as a binder for your work and information.
* **Analyses** are a list of the analyses you have launched. Here, you can access graphs, view visualizations, and gather insights related to your data.
* **Sample Groups** are collections of samples that are organized based on one or more metadata fields, allowing you to group and manage data effectively for analysis.

Click on each section below to learn more.

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>Organize and manage data</td><td><a href="/files/06EMSw6e0WXXb25qGvLU">/files/06EMSw6e0WXXb25qGvLU</a></td><td><a href="/pages/QzgRvv4RpaI5Toqtoxrd">/pages/QzgRvv4RpaI5Toqtoxrd</a></td></tr><tr><td>Analyze and interpret data</td><td><a href="/files/R130AuhTSJ2VmQwtPQZe">/files/R130AuhTSJ2VmQwtPQZe</a></td><td><a href="/pages/vkjPki0B01TpHeS6geI3">/pages/vkjPki0B01TpHeS6geI3</a></td></tr><tr><td>Set and categorize samples</td><td><a href="/files/eA4B9CcwBTPZb0LJFugt">/files/eA4B9CcwBTPZb0LJFugt</a></td><td><a href="/pages/JGz6RHWxwKotYXp6zmFA">/pages/JGz6RHWxwKotYXp6zmFA</a></td></tr></tbody></table>


# Data Inputs

To have data to be used as input for Connected Multiomics, data can be:

* [Generated from a DRAGEN pipeline](#use-data-from-dragen-pipeline)
* [Added from ICA project](#add-data-from-ica-project)
* [Add demo data](#add-demo-data)
* [Uploaded from local drive](#add-data-from-local-drive)
* [Add Generic Count Matrix](#add-generic-count-matrix)
* [Transferred from a public BaseSpace account](#transfer-data-from-public-basespace)

### Use data from DRAGEN pipeline

Connected Multiomics uses data from DRAGEN pipelines within the Illumina Connected ecosystem.

When kicking off DRAGEN pipelines in BaseSpace, ensure the same workgroup is used in BaseSpace as Connected Multiomics so that the DRAGEN outputs are visible to Connected Multiomics. To change the workgroup on BaseSpace, select the desired workgroup in the top right-hand corner.

<figure><img src="/files/cMu7gLNpsFTvga8sLcmZ" alt=""><figcaption></figcaption></figure>

If DRAGEN pipeline are kicked off in BioInsight Platform Core or in BaseSpace using Personal, these results can be shared with the workgroup by adjusting the permissions set as indicated in the [Add data from Core project](#add-data-from-ica-project) section.

### Add data from BioInsight Platform Core project

Connected Multiomics inputs are the results of secondary analysis pipelines that are stored on projects in [BioInsight Platform Core](https://help.ica.illumina.com/) (Core). Secondary analysis results may be generated from [auto-launched pipelines](https://help.ica.illumina.com/sequencer-integration/analysis_autolaunch), run manually or come from [other sources](https://help.multiomics.illumina.com/icm/reference/supported-data-types) such as legacy pipelines or commercial applications.

To access Core in your software, select the BioInsight Platform Core application tile from your Product Dashboard. For guidance on uploading and managing data in Core, please refer to the instructions [here](https://help.ica.illumina.com/project/p-data). For data types and files format supported for analysis in the Connected Multiomics, please refer [Supported Data Types](/icm/reference/supported-data-types).

To ensure users can proceed smoothly with their data exploration in Connected Multiomics, the workgroup that is in use in Connected Multiomics needs to be added to the BioInsight Platform Core **Project(s)** at [**Team**](https://help.ica.illumina.com/project/p-team) settings, with the following permissions set:

* Contributor or higher [role](https://help.ica.illumina.com/project/p-team#project-access)
* [Download](https://help.ica.illumina.com/project/p-team#upload-and-download-rights) permissions
* [Upload](https://help.ica.illumina.com/project/p-team#upload-and-download-rights) permissions

After data files are uploaded and workgroup is added, you may click on the Product Dashboard icon ![](/files/xAUZi7UwYbfycn4WtgmD) to navigate to Illumina Connected Multiomics to start working on your data.

{% embed url="<https://www.youtube.com/watch?v=6ZeaGXXo5DI>" %}

### Add demo data

For new Connected Multiomics users, a tutorial study is linked to all Connected Multiomics domains. This is a read-only study that offers users the ability to view example pipelines and provide visibility into each task and result.

Connected Multiomics also offers the ability to import demo data into any study. This allows the ability to create new analysis tasks and pipelines.

In order to import demo data into a specific study, first clock on the "+ Add Data" > "Select from Core Project".

<div align="left"><figure><img src="/files/xZtUlioU3sylZt4M7gdv" alt=""><figcaption></figcaption></figure></div>

From the new window, select any data type to import. For the purposes of this instruction, DRAGEN Single Cell RNA will be selected. Afterwards, click on "Select format" on the bottom right corner of the screen.

<div align="left"><figure><img src="/files/TEIFZ68gv2U43s3ZqhIK" alt=""><figcaption></figcaption></figure></div>

In the next window, click on "Add Demo Data" button in the upper right section of the screen.

<div align="left"><figure><img src="/files/u11Ktze44wcNXHfgC3XB" alt=""><figcaption></figcaption></figure></div>

After the screen loads momentarily, a notification will be displayed that the "Resource Bundle Linked" and a new folder named "Multiomics-Demo-Data" will appear at the bottom of the page.

<div align="left"><figure><img src="/files/U8H8YP4TX2ra5eip9RAw" alt=""><figcaption></figcaption></figure></div>

Navigating further into the directory will allow the specific files to be imported into Connected Multiomics as samples.

### Add data from local drive

In a Connected Multiomics study, click on Add Data and choose Upload data.

<div align="left"><figure><img src="/files/xZtUlioU3sylZt4M7gdv" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

There are three options:

* **Upload files and folders and add to study**: the wizard will guide you create samples in a study after the files are uploaded
* **Upload files and folders only**: files will be stored in a BioInsight Platform Core project. If you want to create samples from those files later, you need to choose "Select from BioInsight Platform Core project" when you create samples
* **Upload metadata files**: upload [sample metadata](/icm/studies/view-studies/sample-metadata) file and add the metadata into a study

When choosing the first option, data type needs to be selected from the data types panel on the left

<div align="left"><figure><img src="/files/cRQFXQAdu8FSQJNJmNoU" alt=""><figcaption></figcaption></figure></div>

Select a data type and click **Continue**

The files will be uploaded to BioInsight Platform Core, specify a BioInsight Platform Core project folder and drag & drop files from your local device or click on the Browse Files button to select files, click **Upload to BioInsight Platform Core**

<figure><img src="/files/ABrpLGfIWBCnpuQspdIS" alt=""><figcaption></figcaption></figure>

A progress bar will be displayed during uploading, once it is done, the status is displayed as Uploaded

Files will be uploaded to the folder of the BioInsight Platform Core project name, subfolder "**uploads**" and Connected Multiomics project name subfolder

<div align="left"><figure><img src="/files/z53iYfxsfCMZTnJV9nuA" alt=""><figcaption></figcaption></figure></div>

Click on **Ingest to Study** to generate samples in the study.

### Add Generic Count Matrix

For bulk RNA-seq, bulk Proteomics, bulk miRNA, microarray data, sample-by-feature matrix text files are supported. Refer to the [Support Data Types](/icm/reference/supported-data-types) to see which file extensions are accepted. After selecting the matrix file, the input file format needs to be specified, as well as sample identifier and feature identifier location:

<div align="left"><figure><img src="/files/yDyQ56xzTJZaMspa1OLm" alt=""><figcaption></figcaption></figure></div>

User should carefully review the file layout; sample identifiers are in blue color, while feature identifiers are in green color. If multiple files are selected, the format of the files should be the same. The order of the features can be different. All the samples and features from different files will be combined after import. If features are missing in one file, the value of the samples in this file of the missing features will be imputed as 0s.

### Transfer data from public BaseSpace

If the data is in BaseSpace in public domain, see [instructions](https://knowledge.illumina.com/software/basespace-sequence-hub/software-basespace-sequence-hub-reference_material-list/000003654) on how to transfer data to an enterprise domain.

{% embed url="<https://youtu.be/loTlLFdike0?si=elS0hP7kjBb8P_k0>" %}


# Transfer Partek Flow Project

Illumina provides a managed transfer service to help you migrate projects from Partek Flow (Illumina-Hosted) to Connected Multiomics.

Illumina Support performs the transfer on your behalf after receiving the required information.

### Request a Project Transfer

To transfer Partek Flow projects to Connected Multiomics, users must have access to both the source (Partek Flow) and destination (Connected Multiomics) environments.

To initiate a transfer, email Illumina Support at **<techsupport@illumina.com>** and provide the following information:

* **Source:** domain, project ID(s) or indicate to transfer all projects, user token
* **Destination:** domain, workgroup

Instructions for locating this information are provided in the sections below.

#### Source domain

The source domain is part of your Partek Flow URL and appears as the first portion of the address. For example, in `test.partek.com`, the source domain is **test**.

#### Source project

You can choose to transfer all projects or select specific projects to transfer. To select specific projects, a Partek Flow project ID is needed for the transfer. There are two ways to find the project ID:

1. On Partek Flow homepage, mouse over the project name, the information of the project is displayed at the lower-left corner of the page, the number in the url is the project ID

<figure><img src="/files/4cRi46LwN7F4gUvPV7NL" alt=""><figcaption></figcaption></figure>

2. Open the project in Partek Flow, the project ID number can be found in the url on the top of the page

<figure><img src="/files/8FYtbfQ7XiAGftC0vquu" alt=""><figcaption></figcaption></figure>

#### Source user token

When transfering a project, we need a token of the user who has access to export the project — in other words, the user has to be the owner, collaborator or viewer of the project. To get the token, after logging in to Partek Flow, click on user name on the top-right corner of the page, choose **Settings;** at the bottom of the System information page, click on **Generate token**.

<figure><img src="/files/jCR0gpLNWDXwnQkGiruB" alt=""><figcaption></figcaption></figure>

#### Destination domain and workgroup ID

The destination is Illumina Connected Multiomics. On Connected Multiomics homepage, click on the user name on the upper-right corner to open the profile to get the workgroup ID and domain name.

<figure><img src="/files/ohMVTqa95wk468d6tAqo" alt=""><figcaption></figcaption></figure>

### Transfer completion

After the transfer is complete, the user will be notified. When the user logs in to Connected Multiomics, a new study named "*Partek Flow - Migrated projects*" is created. All transferred Partek Flow projects can be found on the *Analyses* page within this study.

<div align="left"><figure><img src="/files/3vwq5C07TSOokLbZe7AT" alt="" width="563"><figcaption></figcaption></figure></div>

{% hint style="warning" %}
Secondary analysis steps, such as alignment, cannot be performed in Connected Multiomics. While the migrated analysis will include all the secondary analysis steps, those steps cannot be re-run.
{% endhint %}

Example analysis where lighter blue tasks and data nodes are secondary analysis steps that cannot be re-run.

<div align="left"><figure><img src="/files/ffhy7KiFQGdMpJEADYsa" alt=""><figcaption></figcaption></figure></div>


# Create Study

To Create a new study, click <img src="/files/rRCLgrIkYorieJ42ddKb" alt="" data-size="line">in the top right and follow the steps below.

{% stepper %}
{% step %}

### Enter a Study Name

Enter a name for your study between 1 and 255 characters in length (including special characters). The study name can be updated later.
{% endstep %}

{% step %}

### Enter a Study Description

Enter an optional description between 1 and 500 characters in length (including special characters). The study description can be updated later.
{% endstep %}

{% step %}

### Select an Existing Project

If the data files are stored in a BioInsight Platform Core project, check the box of "I want to choose my preferred BioInsight Platform Core project" and select a BioInsight Platform Core project to ingest data from. The project list will only display the projects accessible to your workgroup. When adding data to your study, the available data will come from the BioInsight Platform Core project you select.

If you don't have a BioInsight Platform Core project before creating study, you can type in new project name to create one, and then upload files to the project in BioInsight Platform Core after new project creation.
{% endstep %}
{% endstepper %}

<div align="left"><figure><img src="/files/kmjDtshynqiKIBL4Avek" alt="" width="563"><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

Click **Create** to create your study. You will be redirected to the new study. For information on how to navigate within your study, see [View Studies](/icm/studies/view-studies).

{% embed url="<https://youtu.be/X-NO54Hjg2Y?si=-PuD_-ssuUiEKaa3>" %}
How to get started with Illumina Connected Multiomics
{% endembed %}


# Add Data

To add data into a study, click the <img src="/files/WVOXx0olgnt7hvkO9l0E" alt="" data-size="line"> button on the top right.

<div align="left"><figure><img src="/files/xZtUlioU3sylZt4M7gdv" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

You can select data from an [BioInsight Platform Core project](/icm/introduction/upload-data-files) or[ upload local files](/icm/introduction/upload-data-files). Selecting data from a BioInsight Platform Core project will open a new screen where you can select your data. At the top of the screen, you'll see the name of your study. On the left panel, is a list of [supported data types](/icm/reference/supported-data-types) in Connected Multiomics, with *Illumina data types* on the top section, third-party data types at the *Other data types* section. You can search data type by typing into the *Search data types* search box.

<div align="left"><figure><img src="/files/9UODOolUMpBsNj5Fa1sc" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

For each of the Other data types, click on the arrow (>) to expand the data type to view details:

<div align="left"><figure><img src="/files/dMYVdqB4VFhWUBzfFzdX" alt=""><figcaption></figcaption></figure></div>

Click to select a data type from the left panel, you will see more details about the selected data type on the right, such as accepted file types, example filename, and input data options. For example, if you choose DRAGEN Single Cell RNA, you will see that you have the options to add either a .h5ad file or a set of feature, barcode, matrix file, per sample.

<figure><img src="/files/TEIFZ68gv2U43s3ZqhIK" alt=""><figcaption></figcaption></figure>

After understanding the required files format, click on **Select format**, and you will be directed to files selection page. Navigate to the folder containing your data files. Select an entire folder or click into the folder to select files individually. You can search files by typing into the *Search* text box on the top right to filter the file list.

<figure><img src="/files/PQA0FpF6xIug6FPhzj7m" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure>

For data types with multiple files per sample, you can tick **Group by sample** checkbox on the top right to automatically group and select files by samples.

<figure><img src="/files/p4ou4wl6iYM6crkkt5IH" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure>

Demo data that is used to generate the [Tutorial Study](/icm/introduction/icm#explore-onboarding-resources) is available to you to add to your Study by clicking <img src="/files/Bq3B3aNg0jOjLjDlXoUM" alt="" data-size="line">.

Once you've selected your data, click <img src="/files/9sNbOWxptkNQJbv1FLpK" alt="" data-size="line"> to add the data to your study. Upon submitting, you'll be redirected to the **Samples** tab, and a message will indicate your data ingestion was successful. Here you can [upload metadata](/icm/studies/view-studies/sample-metadata) from local storage, [add additional data](#add-data) from a BioInsight Platform Core project, create [sample groups](/icm/studies/view-studies/view-studies-1), or create [new analysis](/icm/studies/view-studies/view-studies-2#create-analysis).

<figure><img src="/files/3O19N0taByg46Ju59lPY" alt=""><figcaption></figcaption></figure>

Now, if you click back to the **Overview** tab, you'll see that the **Date Modified** was updated and the **Data Metrics** reflect the newly added files and samples.

Repeat the process to add data for additional data types.


# View Studies

A **Study** is used to manage your samples, metadata, comparison groups, and analyses for a specific set of data.

### Overview

The Study overview page will display the following **Study Information**:

| Information                      | Description                                                                                                                                                                                                                                                        |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Date Created                     | The date and time the study was created.                                                                                                                                                                                                                           |
| Created By                       | The user who created the study.                                                                                                                                                                                                                                    |
| Date Modified                    | The last date and time the study was modified. This date updates whenever a user performs an action within the study, such as adding data, creating a sample group, starting a new analysis, etc. Users may need to manually refresh their browser to see updates. |
| Modified By                      | The user who last modified the study.                                                                                                                                                                                                                              |
| BioInsight Platform Core Project | The BioInsight Platform Core project the study imports data from.                                                                                                                                                                                                  |
| Description                      | The description of the study, provided by the user during study creation.                                                                                                                                                                                          |

You will also see the following **Data Metrics**:

| Information             | Description                                                                                                                                                                                                                                                                                            |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Sample Files            | The number of sample files within the study, including both manually added and auto-ingested sample files. Please see the [Supported Data Types](/icm/reference/supported-data-types) for a list of possible sample file extensions.                                                                   |
| Metadata Files          | The number of metadata files within the study, identified by the `.tsv` file extension.                                                                                                                                                                                                                |
| Samples                 | The number of samples within the study. Note that multiple protein samples can come from a single ADAT file, whereas a single spatial or scRNA sample is created for each set of spatial/scRNA files. Please see the [Supported Data Types](/icm/reference/supported-data-types) for more information. |
| Analyses                | The number of analyses within your study, including both in-progress and completed analyses.                                                                                                                                                                                                           |
| Attributed Sample Count | The number of samples in the study with linked metadata.                                                                                                                                                                                                                                               |
| Metadata Entry Count    | The number of metadata attributes displayed in the study. This corresponds with the number of metadata columns shown next to your samples.                                                                                                                                                             |

If there are analyses in your study, you will see the **Recent Analyses** section at the bottom of the **Overview** screen.

<div align="left"><figure><img src="/files/gkpuTOzQq5GybZL8jYup" alt=""><figcaption></figcaption></figure></div>


# Samples

The Samples page lists all the samples added into the study with meta data information.

<figure><img src="/files/PadeSiRvgu1VKegvOmWz" alt=""><figcaption></figcaption></figure>

Top left is a search bar, type in Sample ID to search. It is doing partial matching, it is case insensitive.

Filter (<img src="/files/YDxtfjFqcJ10bxAhhbFb" alt="" data-size="line">): click on filter icon, meta data field list is displayed on the right side. The search bar on the top is to search for the field name. Click on > to expand a field, search bar within each field is to search for subgroup name. Check the box in front of each subgroup name to select. Deselected groups will be filtered out on the table.

<div align="left"><figure><img src="/files/F37fIWjVuZGzMZcCf9tA" alt="" width="288"><figcaption></figcaption></figure></div>

Columns (<img src="/files/s70MGVfuCbOVzDemtoP2" alt="" data-size="original">): allows you to customize which columns you want to view in the table, click on the columns icon, meta data field list is displayed on the right.

<div align="left"><figure><img src="/files/xPU9WvmMu8V3cfHZXboE" alt="" width="280"><figcaption></figcaption></figure></div>

Check the box in front the field name to display the field in the table. Unchecked filed names are hidden from the table.

On the table, check the box in front of the sample ID to select samples. When at least one sample is selected, the **Create sample group** button (<img src="/files/IeCFZm9VxdHGhRSRk7OO" alt="" data-size="line"> ) will be enabled (upper-right of the table), click on this button, the created sample group will be displayed in Sample Groups page.

When at least one sample is selected, the **Create analysis** button (<img src="/files/FgwZshxBbRfzy6qG7jAA" alt="" data-size="line"> ) will be enabled (upper-right of the table), click on this button to create a new analysis. In the **New Analysis** dialog, give a name to the analysis, select an [analysis type](/icm/studies/view-studies/view-studies-2#analysis-types), then click **Run Analysis** to create the analysis. The created analysis will be displayed in Analyses page.

<figure><img src="/files/lBFuxQMEeWamibjw53o4" alt=""><figcaption></figcaption></figure>

Selected sample meta data can be downloaded by clicking <img src="/files/gu6ao34hK4OuMp8lGiHO" alt="" data-size="line">; selected samples will be removed from the study when click <img src="/files/qtsjD1MYy9K8Jon7PZYu" alt="" data-size="line">. A warning message will be displayed:

<div align="left"><figure><img src="/files/AAXoqmSODzDIMa7kX04K" alt="" width="563"><figcaption></figcaption></figure></div>

Number of the samples displayed in the table, total number and number of selected samples are displayed at the low-right corner of the table.

<div align="left"><figure><img src="/files/oCJWeEjdJg7MnKbaJc8x" alt="" width="563"><figcaption></figcaption></figure></div>

Click the refresh icon <img src="/files/bQhthCGB6NtR1Yp6yjth" alt="" data-size="line"> to refresh the sample view after adding more data or to clear your filters and reset to the default view. You can also drag and drop columns to reorder columns.

### Sample Naming

#### Proteomic Data

Sample IDs for proteomic data will be extracted from the sample names provided in the ADAT files. Each file can contain multiple samples. For example, the Sample ID in the ADAT file (left image) corresponds to the sample name in a row of the Connected Multiomics table (right image). A new sample will be created in Connected Multiomics for each Sample ID listed in the ADAT file.

<figure><img src="/files/eiELgK4nRpukEU3u2Rx3" alt=""><figcaption></figcaption></figure>

#### Single Cell RNA Data

Sample IDs for scRNA data are obtained by removing the file extension from the file name. As a result, scRNA samples will be paired with their corresponding files based on the sample name. For example, the name of the scRNA files (left image) matches the sample name in a row of the ICM table (right image). Refer to [Supported Data Types](/icm/reference/supported-data-types) for information on the required files that make up a single scRNA sample.

<figure><img src="/files/W7QUffvHKW5yfNSZAe4x" alt=""><figcaption></figcaption></figure>

#### Spatial Data

Sample IDs for spatial data are derived by removing the file extension from the file name. Thus, spatial samples are linked to their corresponding files through the sample name. For instance, the name of the spatial data files (left image) corresponds to the sample name in a row of the Connected Multiomics table (right image). Refer to [Supported Data Types](/icm/reference/supported-data-types) for details on the required files that make up a single spatial sample.

<figure><img src="/files/vJJhodk0HsmYiMEpNCen" alt=""><figcaption></figcaption></figure>

### Sample Statuses

The following table describes the different sample statuses.

{% hint style="warning" %}
Note that protein samples only display the "Ready" status.
{% endhint %}

| Status    | Description                                                                                                          |
| --------- | -------------------------------------------------------------------------------------------------------------------- |
| Ingested  | The sample has been ingested to the Study but may not have been uploaded to Analysis yet.                            |
| Uploading | The sample is in the process of uploading to Analysis. Samples cannot be used for analyses while they are uploading. |
| Ready     | The sample is ready to be used in analysis.                                                                          |
| Failed    | The sample ingestion failed.                                                                                         |


# Sample Metadata

## Sample Metadata

Sample metadata allows you to add important additional information to your study that can be used to create sample groups or make comparisons as part of your study design. Sample metadata can be uploaded from local storage or added to a BioInsight Platform Core project.

### Metadata mapping

Metadata columns are populated based on the information in your metadata TSV file. Metadata files are required per study. Upon importing a metadata file, the samples will automatically be updated to reflect the metadata attributes. The sample names in the metadata file must match those in the data files. For example, in the figure below, you can see that the Sample ID in the ADAT file matches the SampleID in the metadata file. As a result, the attribute columns in the metadata file will correspond with the metadata columns available for selection in the **Samples** table.

{% hint style="warning" %}
The column order does not matter but you must have a column headed with "**SampleID**" (no space) and the column values need to match the Sample ID in the software.
{% endhint %}

<figure><img src="/files/TcOnijsYavNCacEySDGV" alt=""><figcaption></figcaption></figure>

You can add multiple metadata files. In this case, any samples that do not overlap with existing ones will be added as new entries in the **Samples** table.

{% hint style="warning" %}
Metadata cannot be overwritten. If an erroneous metadata file is added, a new study and analysis should be created.
{% endhint %}

### Download Sample Metadata file

Download a formatted Metadata file example for your study to review the metadata information. Make sure to correct the SampleID column header to match the exact name of the samples.

* Navigate to **Sample groups > Info> Download Metadata.** Click the Sample groups tab, under Action click the Info button (<img src="/files/qiUwxlCrN0w3gA9OtQZ5" alt="" data-size="line">), then use the **Download Metadata** button to download an example metadata file to your machine.

{% hint style="warning" %}
Metadata files should be saved as a .tsv or .csv extension.
{% endhint %}

<figure><img src="/files/cte2lgzl6A1kDOXECGCk" alt=""><figcaption><p>Download Metadata example file for modification with attribute information</p></figcaption></figure>

### Upload sample metadata file to Connected Multiomics

Upload the sample metadata file to Connected Multiomics to update the sample records of the study.

* Within a study, click **+ Add Data > Upload data > Upload metadata files**
* Upload data from your local or server based drives. Click Browse to select the files or drag and drop the files to the purple box.

{% hint style="warning" %}
This will not add the metadata file to the BioInsight Platform Core project as a stored file. The metadata uploaded a Connected Multiomics study will be used only in this study. It will need to be added again to other studies later if you want to use it again for other samples. To add metadata to the BioInsight Platform Core project you will need to do this within the BioInsight Platform Core interface.
{% endhint %}

<figure><img src="/files/PWJXYAGzKV7l9qJ7E4Fs" alt=""><figcaption><p>Upload metadata file to an existing study</p></figcaption></figure>

### Using a sample metadata file contained in the BioInsight Platform Core project

To add data, click the <img src="/files/WVOXx0olgnt7hvkO9l0E" alt="" data-size="line"> button in the top right. You can [select data from a BioInsight Platform Core project](/icm/studies/create-study/add-data#add-data) and choose your data type and format. TSV metadata files can always be selected. For example, if you choose Bulk > Proteomics as the data type, you will only see ADAT files and TSV files as options for upload.

### Sync metadata from Study

After the new metadata file has been added to the ***Study***, users will be able to update their metadata within an existing analysis by clicking the "***Sync metadata from study***" button in the ***Metadata*** tab.

A new confirmation dialog will appear to make sure users understand what will happen.

<div align="left"><figure><img src="/files/u5kZy4uH4648tEFnhSRu" alt="" width="545"><figcaption></figcaption></figure></div>

Once the above ***Sync metadata*** button is pushed, the metadata will be updated accordingly.

<figure><img src="/files/QoL7Z3hgFWHOz76jjhWc" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
No metadata will be applied to the duplicate samples if the metadata file does not contain the sample\_uuid.
{% endhint %}


# Cell metadata

### [Publish Cell attributes to a project](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/publish-cell-attributes-to-project)

Cell level attributes can be used for filtering, cell selection, classification and differential analysis using the [Publish Cell attributes to project task](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/publish-cell-attributes-to-project). Data nodes in the Analyses pipeline that contain this cell level information (e.g. Graph-based clusters data node) can be published to the project and will then be available in **Study > Analysis > Metadata > Manage** under Cell attributes where they can be modified and reordered.

### [Annotate cells](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells)

If you have attribute information about your cells, you can use the [Annotate cells task](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells) in Connected Multiomics to apply this information to the data. Once applied, these can be used like any other attributes, and thus can be used for cell selection, classification and differential analysis.

To run Annotate cells:

* Click a Single cell counts data node
* Click the **Annotation/Metadata** section in the toolbox
* Click **Annotate cells**

You will be prompted to specify annotation input options:

* Single file (all samples): requires one .txt file for all cells in all samples. Each row in the file represents a barcode and at least one barcode column which will match the barcodes in your data. It also requires a column containing Sample ID which must match the Sample name in the Metadata tab of your project.
* File per sample: requires the format of all of the annotation files to be the same. Each file has barcodes on rows, it requires one barcode column that will match the barcodes in your data in that sample. All files should have the same set of columns, column headers are case sensitive.


# Sample Groups

A sample group refers to a collection or subset of samples that are grouped together based on shared characteristics or attributes. You can view all your sample groups by clicking the **Sample Groups** tab in the top.

Created sample groups in the study are listed on this page.

### Create Sample Group

At **Samples** tab, check the box in front of the sample ID to select samples, when at least one sample is selected, the **Create sample group** button (<img src="/files/IeCFZm9VxdHGhRSRk7OO" alt="" data-size="line"> ) will be enabled (upper-right of the table).

You can also filter samples down into only those you want to be part of your sample group by clicking the filter icon at the top of each visible column or clicking **Filters** icon above the table <img src="/files/1FP1TvrejxYEWGWbP2D9" alt="" data-size="line"> to filter all columns based on metadata attributes.

Select all by clicking the check marks on the left-hand side. Now, click <img src="/files/VSexTC8nfQp15N1ECgJR" alt="" data-size="line"> to create a sample group. Choose to create a new sample group or add to an existing sample group. Give the sample group a name and click <img src="/files/EDXK3fKiFhttpRY7jOWD" alt="" data-size="line">.

<figure><img src="/files/q2ekgDBCpP0PI27W4Ujj" alt=""><figcaption></figcaption></figure>

### Manage Sample Groups

In the sample groups page, you can search for sample groups, and customize or filter the columns in your view.

You can also click on the action icons to view more details and update your sample groups.

<table><thead><tr><th width="112">Icon</th><th>Action</th></tr></thead><tbody><tr><td><img src="/files/6WMlTptf2zPYWcLohHAM" alt="" data-size="original"></td><td>Opens a pop-up with details of your sample group.<br><img src="/files/GXrA4UDcZKHuFclybHuK" alt=""></td></tr><tr><td><img src="/files/c5zz8t9EAQRzfnFQIRKZ" alt="" data-size="original"></td><td>Allows you to edit the sample group's name.<br><img src="/files/FqHU5K1iVReDE766vgaf" alt=""></td></tr><tr><td><img src="/files/IynLr6tKTEpmOBVDraYl" alt="" data-size="original"></td><td>Allows you to delete the sample group.<br></td></tr><tr><td></td><td></td></tr></tbody></table>


# Analysis and Data Management

In the **Analyses** tab, you'll see all the analyses in your study. Multiple analyses can be performed on each study, user can use sample group to create a subset of the data to perform analysis, or choose different analysis type: default or custom to analyze the data in a study.

### Card View vs List View

There are two types of display on *Analyses* page. On the upper-left corner of the page, you can switch between the two view styles, card view vs list view.

<div align="left"><figure><img src="/files/g2BimSnNeiysQyV5zNWG" alt=""><figcaption></figcaption></figure></div>

In **Card View**, each card will display the following information.

| Information   | Description                                                                                               |
| ------------- | --------------------------------------------------------------------------------------------------------- |
| Analysis Name | Name of analysis.                                                                                         |
| Data Type     | Data type for data imported into the analysis.                                                            |
| Status        | Status of analysis. See [Analysis Statuses](#analysis-statuses) for a table of statuses and descriptions. |
| Modified By   | Last user who modified the analysis.                                                                      |
| Date Modified | Date and time when the last update was made.                                                              |

The pin button ( ![](/files/jODdoDuEG6fNIRWTVDrp) ) pins an analysis to the top. The three dots ( ![](/files/NqL9iTeFtOuhicdvRvCP) ) opens to more options:

* **Explore in data viewer**: Open a saved data viewer session or a new data viewer session to create plots.
* **View details**: Show more details of the analysis, including error message when the analysis failed.
* **Share a copy**: Create a link to share the analysis.
* **Delete analysis**: Delete the analysis.

<div align="left"><figure><img src="/files/ycGGNeduEKikjbYj7EiT" alt="" width="470"><figcaption></figcaption></figure></div>

There are more controls in **List view**:

<div align="left"><figure><img src="/files/wbZibON25i6f521prBpa" alt=""><figcaption></figcaption></figure></div>

On column header, there is filter icon (![](/files/TseOqWNneIbzcwockGTL)), it allows to filter the rows in the table based on the selected value:

<div align="left"><figure><img src="/files/TvqiT4NQAWh5fSkZFhjo" alt=""><figcaption></figcaption></figure></div>

Click on the 3 dots next to the filter icon (![](/files/S9Vz5TeEJLKWBW3werlh)) to sort the table, choose columns to display

<div align="left"><figure><img src="/files/0Ynmlt7iDoh8yAAyVQCm" alt="" width="366"><figcaption></figcaption></figure></div>

### Create Analysis

To create a new analysis, go to Samples tab or Sample Groups tab, select at least one sample or one sample group, the Create analysis button ( <img src="/files/FgwZshxBbRfzy6qG7jAA" alt="" data-size="line"> ) will be enabled (upper-right of the table), click on this button to create a new analysis. In the **New Analysis** dialog, give your analysis a name, choose an analysis type from the dropdown menu.

<figure><img src="/files/lBFuxQMEeWamibjw53o4" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
If all samples are chosen for the analysis, a sample group will be automatically added to the Sample Groups page containing all of the samples in the analysis.
{% endhint %}

{% hint style="warning" %}
Current spatial multisample analysis does not yet support the import of a mix of samples that use pipeline-manifest.json and samples that do not. When running an analysis on a sample group that has both types of samples, only those samples with a pipeline-manifest.json file will import.
{% endhint %}

### Analysis Types

| Analysis Option                                            | Description                                                                                                                                                                                                                                                                                                                                                                                |
| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Default: Illumina Proteomics                               | General analysis pipeline for Illumina protein prep samples, with PCA and hierarchical clustering heatmap results. The maximum samples for analysis is 9,000.                                                                                                                                                                                                                              |
| Default: Illumina Single Cell Transcriptomics              | <p>General analysis pipeline for Illumina single cell prep samples, including the processing steps to output PCA and UMAP plots, plus graph-based clustering visualizations.</p><p>For each sample, all features are reported and raw counts are used as the count value. If the feature IDs are not unique mean is used for deduplication. The maximum samples for analysis is 1,000.</p> |
| Default: Illumina Single Cell Transcriptomics -Perturb-seq | General analysis pipeline for Illumina single cell prep samples, including the processing steps to output PCA and UMAP plots, plus graph-based clustering visualizations. gRNAs are available but not further analyzed since user input is required for clustering and differential expression to obtain desired results. The maximum samples for analysis is 1,000.                       |
| Default: Illumina Spatial Transcriptomics                  | General analysis pipeline for Illumina spatial samples, including a Spatial map with transcripts overlayed on the tissue image where each point is a grid. Graph-based clusters are also plotted on a UMAP and pie chart. The maximum samples for analysis is 100.                                                                                                                         |
| Default: Illumina Bulk Transcriptomics                     | General analysis pipeline for Illumina mRNA and Total RNA prep samples including PCA and hierarchical clustering heatmap results. The maximum samples for analysis is 500.                                                                                                                                                                                                                 |
| Default: Illumina miRNA                                    | General analysis pipeline for Illumina miRNA prep samples including PCA and hierarchical clustering heatmap results. The maximum samples for analysis is 500.                                                                                                                                                                                                                              |
| Default: Illumina 5-base DNA Methylation                   | General analysis pipeline for 5-base DNA prep samples including regional methylation PCA plots, k-means clustering, and QC. The maximum samples for analysis is 200.                                                                                                                                                                                                                       |
| Custom: Multiomics                                         | Available inputs include Illumina proteomics, single cell, spatial, and bulk transcriptomics, miRNA, 5-base DNA and Infinium methylation as well as third-party data types including Seurat (RNA), Seurat (ATAC), Somalogic ADAT, Gene counts in sf Format, and VCF. Two or more data types can be analyzed together in any combination. The maximum samples for analysis is 200.          |
| Custom: Illumina Proteomics                                | Starts with the quantified samples which have undergone prior normalization and offers flexible analysis options. The maximum samples for analysis is 9,000.                                                                                                                                                                                                                               |
| Custom: Illumina Single Cell Transcriptomics               | Starts with single cell counts and offers flexibility with the analyses pipeline step. For each sample, all features are reported and raw counts are used as the count value. If the feature IDs are not unique mean is used for deduplication. The maximum samples for analysis is 1,000.                                                                                                 |
| Custom: Illumina Single Cell Transcriptomics -Perturb-seq  | Starts with single cell counts and offers flexibility with the analyses pipeline step. For each sample, all features, both GEX and gRNAs, are reported and if feature ID's are not unique, Mean is used for Deduplication; raw counts are used as the count value format and cells with a total read count at least 400 are reported. The maximum samples for analysis is 1,000.           |
| Custom: Illumina Spatial Transcriptomics                   | Two starting nodes as cell-binned or grid-binned data, including the spatial image outputs, and offers flexibility with the analyses parameters. The maximum samples for analysis is 100.                                                                                                                                                                                                  |
| Custom: Illumina Bulk Transcriptomics                      | Starts with salmon format sample counts that have not been normalized and offers flexible analyses options. The assembly and annotation model used in secondary analysis is required to annotate the features. The maximum samples for analysis is 500.                                                                                                                                    |
| Custom: Illumina miRNA                                     | Starts with imported count matrix that have not been normalized and offers flexible analysis options. The maximum samples for analysis is 500.                                                                                                                                                                                                                                             |
| Custom: Illumina 5-base DNA Methylation                    | Starts with importing selected 5-base methylation data and offers flexible analysis options. The maximum samples for analysis is 200.                                                                                                                                                                                                                                                      |
| Custom: Illumina Infinium Methylation                      | Starts with importing selected idat files and offers flexible analysis options. No limit.                                                                                                                                                                                                                                                                                                  |
| Custom: Third-party analysis                               | Starts with data import. This option appears only when third-party assay data has been uploaded.                                                                                                                                                                                                                                                                                           |

### Run Analysis

Click <img src="/files/PWSxNI8XFatqswIa1Def" alt="" data-size="line"> to run the analysis. You will receive a notification that the analysis was successfully created.

<figure><img src="/files/9HNszxBJUS9yQbjAqdxT" alt=""><figcaption></figcaption></figure>

Wait for the analysis status to change from "Pending" to "Complete".

<figure><img src="/files/vsFckjMj7T0cez8Nk4iQ" alt=""><figcaption></figcaption></figure>

### Analysis Statuses

The following table describes the different analysis statuses.

<table><thead><tr><th width="193">Status</th><th>Description</th></tr></thead><tbody><tr><td>Pending</td><td>The analysis has been initiated but has not yet started. It is waiting in the queue to be processed.</td></tr><tr><td>In Progress</td><td>The analysis is currently being executed. The system is processing the data and generating results.</td></tr><tr><td>Complete</td><td>The analysis has finished, and the results are ready to be viewed. Click into the analysis to view it.</td></tr><tr><td>Importing</td><td>The analysis is importing from a shared copy. The status will change to Complete upon successful import.</td></tr><tr><td>Offline</td><td>During import, the source server might be temporarily unavailable. Click the refresh button. Contact tech support if the issue persists.</td></tr><tr><td>Error</td><td>The analysis has ended in error. Using the Analysis tab list view, click the info icon <img src="/files/CTob3vWY9LiKlE6TeIEo" alt="" data-size="line">to get more information about the error type. The error can be further diagnosed using the <strong>Inspect analysis</strong> link.</td></tr></tbody></table>

Once the analysis is complete, you can click into it to view the results. Refer to the "[**Enter Analysis**](/icm/analyses/enter-analysis)" section for more information on how to view your analysis.

## Data

In the **Data** tab, you'll see a list of all the data and metadata files within your study from which your samples are derived.

{% hint style="warning" %}
Note that deleting data or metadata files will not remove any samples that were added to Connected Multiomics from those files. The samples will remain in Connected Multiomics even if the original files are deleted, ensuring that your sample data stays intact.
{% endhint %}

<div align="left"><figure><img src="/files/wzAcwHK08KQOgDifX7cE" alt=""><figcaption></figcaption></figure></div>

###


# Study Settings

Within a study, there is a settings gear button allow more advanced editing for the study.

<div align="left"><figure><img src="/files/2n6vcAdXHm77ptvKQ6ih" alt=""><figcaption></figcaption></figure></div>

Study name, description and preferred project can be edited.

<div align="left"><figure><img src="/files/jvT0LQ7FXPc5Uc0d6LdR" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

When select *Automatically import data from pipelines*, new generated samples in preferred BioInsight Platform Core project in the selected assay will be automatically created.

<div align="left"><figure><img src="/files/WgO2GxnWapzwGx3m5wtr" alt=""><figcaption><p>ICA is now BioInsight Platform Core</p></figcaption></figure></div>

Click **Update** to save the change.


# View Analyses Across Studies

Click on the **Analyses** tab in the left panel to see a list of all analyses within your workgroup, including those across various studies. Use the **Columns** and **Filters** tabs on the right to adjust and refine your view. You can also search for specific analyses using the search bar in the top right. Additionally, you can perform actions on your analyses, such as renaming or deleting an analysis.

<figure><img src="/files/dzSlqk3fQTIsTbQbwSXV" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Analyses can only be created from within a study. See the [**Studies**](/icm/studies/view-studies) section for more information on how to create an analysis.
{% endhint %}


# Enter Analysis

Click into your analysis to see details about your analysis steps.

## Analyses

In the [**Analyses**](/icm/analyses/analysis-functionality) tab of your analysis within a study, you'll be able to see a depiction of your analysis pipeline, represented by data nodes and task nodes connected by arrows. Each rectangle represents a task, and each circle represents an output. You can click on each node to view more details on the right-side toolbox.

<figure><img src="/files/SAvEI5mkoyZTpYmZpQF3" alt=""><figcaption></figcaption></figure>

Double click on each node to see the task details appear in a separate window.

<figure><img src="/files/hj07URJtYQ6K94uud93B" alt=""><figcaption></figcaption></figure>

## Metadata

You can view the sample attributes of all the samples in your study in the [Samples](/icm/studies/view-studies/view-studies) tab of your Study. These sample attributes are then used in your Analysis for the samples you selected. Your analysis will display a table of the samples used only in the analysis, along with the sample groups they originated from. You can download the sample data here in the Metadata tab of the analysis and manage the sample attributes.

<figure><img src="/files/AluhCC5weW1xlyPoGDnE" alt=""><figcaption></figcaption></figure>

## Log

The **Log** tab displays a record of tasks from the task graph, including details such as the user who performed the task, as well as the start and end dates.

<figure><img src="/files/M1uzM3wngiaH8UWZaKvg" alt=""><figcaption></figcaption></figure>

## Project Settings

The **Project Settings** tab displays the details of the analysis, such as its name and description. From here, you can edit the analysis details, but note that the name cannot be changed.

<figure><img src="/files/oWAxGpPgP72Oqn7dXrZj" alt=""><figcaption></figcaption></figure>

## Data viewer

In the [**Data Viewer**](/icm/analyses/analysis-functionality/data-viewer) tab, you can return to your saved sessions or start new sessions. Saved sessions will retain all the graphs and settings from your last use. You can click on a session to continue where you left off, or create a new Data Viewer session to set up a different data view with new graphs.

<figure><img src="/files/1mjjVPXXG2DgtN0rkVKJ" alt=""><figcaption></figcaption></figure>

Click into the Data Viewer to access your session. For guidance on how to analyze your data, refer to the [Walkthroughs](/icm/analyses/walkthroughs).

<figure><img src="/files/bTxzokssNKACodPMZ3RX" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
You can navigate between studies and analyses using the breadcrumbs. The top breadcrumb allows you to move from studies to analyses, while the bottom breadcrumb shows either the analysis or the data viewer you have selected.

<img src="/files/Rr9O98CSWkg1cZHTnMEZ" alt="" data-size="original">
{% endhint %}


# Analysis Functionality


# Task Menu

The Task Menu lists all the tasks that can be performed on a specific node. It can be invoked from either a **Data** or **Task node** and appears on the right hand side of the *Analyses* tab. It is *context-sensitive*, meaning that it will only present tasks that the user can perform on the selected node. For example, selecting an *Differential analysis* report data node will not present normalization as options.

Clicking a **Data node** presents a variety of tasks:

* [AI analysis methods summary](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/LAv0DgjBxHMz2IXfYJ9O)
* [AI analysis suggestions](https://app.gitbook.com/o/-MWUoaZPOpY9hR_8vqOU/s/5WPPw051cYE3Zthy5U7m/~/edit/~/changes/101/analyses/analysis-functionality/ai-analysis-suggestions)
* [Task actions](/icm/analyses/analysis-functionality/task-menu/task-actions)
* [Data summary report](/icm/analyses/analysis-functionality/task-menu/data-summary-report)
* [QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc)
  * [Feature distribution](/icm/analyses/analysis-functionality/task-menu/qa-qc/feature-distribution)
  * [Imported count matrix report](/icm/analyses/analysis-functionality/task-menu/qa-qc/feature-distribution-1)
  * [Single-cell QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/single-cell-qa-qc)
  * [Cell barcode QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/cell-barcode-qa-qc)
  * [5-base Methylation QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/5-base-methylation-qc)
* [Annotation/Metadata](/icm/analyses/analysis-functionality/task-menu/annotation-metadata)
  * [Annotate cells](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells)
  * [Annotate features](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells-1)
  * [Publish cell attributes to project](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/publish-cell-attributes-to-project)
  * [Attribute report](broken://spaces/5WPPw051cYE3Zthy5U7m/pages/fOdrsK868SrQaRk5Ly1F)
* [Pre-analysis tools](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools)
  * [Pseudobulk](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/pool-cells)
  * [Split by feature type](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/split-matrix)
  * [Generate group cell counts](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/generate-group-cell-counts)
  * [Merge matrices](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/merge-matrices)
  * [Generate beta value](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/generate-beta-value)
  * [Generate methylation mvalues](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/generate-methylation-mvalues)
* [Filtering](/icm/analyses/analysis-functionality/task-menu/filtering)
  * [Filter features](/icm/analyses/analysis-functionality/task-menu/filtering/filter-features)
  * [Filter samples/cells](/icm/analyses/analysis-functionality/task-menu/filtering/filter-groups-samples-or-cells)
  * [Split by attribute](/icm/analyses/analysis-functionality/task-menu/filtering/split-by-attribute)
  * [Downsample cells](/icm/analyses/analysis-functionality/task-menu/filtering/downsample-cells)
* [Normalization and scaling](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling)
  * [Normalization](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/normalization)
  * [Normalize to housekeeping genes](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/normalize-to-housekeeping-genes)
  * [Normalize to baseline](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/normalize-to-baseline)
  * [Scran deconvolution](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/scran-deconvolution)
  * [TF-IDF normalization](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/tf-idf-normalization)
  * [Impute missing values](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/impute-missing-values)
  * [Impute missing values](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/impute-missing-values)
  * [Impute low expression](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/impute-low-expression)
  * [SCTransform](/icm/analyses/analysis-functionality/task-menu/normalization-and-scaling/sctransform)
* [Batch removal](/icm/analyses/analysis-functionality/task-menu/batch-removal)
  * [General linear model](/icm/analyses/analysis-functionality/task-menu/batch-removal/general-linear-model)
  * [Harmony](/icm/analyses/analysis-functionality/task-menu/batch-removal/harmony)
  * [Seurat3 integration](/icm/analyses/analysis-functionality/task-menu/batch-removal/seurat3-integration)
* [Statistics](/icm/analyses/analysis-functionality/task-menu/statistics)
  * [Differential Analysis](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis)
    * [DESeq2](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/deseq2-r-vs-deseq2)
    * [Hurdle model](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/hurdle-model)
    * [ANOVA/LIMMA-trend/LIMMA-voom](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/anova-limma-trend-limma-voom)
    * [Welch's ANOVA](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/kruskal-wallis)
    * [Kruskal-Wallis / Wilcoxon](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/kruskal-wallis-1)
    * [Poisson/Negative binomial/GSA (Gene Specific Analysis)](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/gsa)
    * [Troubleshooting](/icm/analyses/analysis-functionality/task-menu/statistics/differential-analysis/troubleshooting)
  * [Descriptive statistics](/icm/analyses/analysis-functionality/task-menu/statistics/correlation-analysis-1)
  * [Correlation](/icm/analyses/analysis-functionality/task-menu/statistics/correlation-analysis-1)
  * [QTL analysis](/icm/analyses/analysis-functionality/task-menu/statistics/correlation-analysis-1-1)
  * [Differential methylation](/icm/analyses/analysis-functionality/task-menu/statistics/differential-methylation)
  * [Detect differential methylation](/icm/analyses/analysis-functionality/task-menu/statistics/detect-differential-methylation)
  * [Compute biomarkers](/icm/analyses/analysis-functionality/task-menu/statistics/compute-biomarkers)
  * [Survival Analysis with Cox regression and Kaplan-Meier analysis](/icm/analyses/analysis-functionality/task-menu/statistics/survival-analysis-with-cox-regression-and-kaplan-meier-analysis)
  * [Classification](/icm/analyses/analysis-functionality/task-menu/statistics/survival-analysis-with-cox-regression-and-kaplan-meier-analysis-1)
  * [Spatially Variable Genes](/icm/analyses/analysis-functionality/task-menu/statistics/spatially-variable-genes)
* [Exploratory analysis](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis)
  * [Graph-based clustering](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/graph-based-clustering)
  * [K-means clustering](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/k-means-clustering)
  * [Compare clusters](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/compare-clusters)
  * [PCA](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/pca)
  * [t-SNE](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/t-sne)
  * [UMAP](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/umap)
  * [Hierarchical clustering / heatmap](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/hierarchical-clustering)
  * [AUCell](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/aucell)
  * [Find multimodal neighbors](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/find-multimodal-neighbors)
  * [SVD](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/svd)
  * [CellPhoneDB](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/cellphonedb)
  * [BANKSY - Spatial Domain Identification](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/banksy-spatial-domain-identification)
  * [Merge differential expression results](/icm/analyses/analysis-functionality/task-menu/exploratory-analysis/merge-differential-expression-results)
* [Region analysis](/icm/analyses/analysis-functionality/task-menu/region-analysis)
  * [Get Regional Methylation](/icm/analyses/analysis-functionality/task-menu/region-analysis/get-regional-methylation)
  * [Annotate Regions](/icm/analyses/analysis-functionality/task-menu/region-analysis/annotate-regions)
* [Trajectory analysis](/icm/analyses/analysis-functionality/task-menu/trajectory-analysis)
  * [Trajectory Analysis (Monocle 2)](/icm/analyses/analysis-functionality/task-menu/trajectory-analysis/trajectory-analysis-monocle-2)
  * [Trajectory Analysis (Monocle 3)](/icm/analyses/analysis-functionality/task-menu/trajectory-analysis/trajectory-analysis-monocle-3)
* [Variant Analysis](/icm/analyses/analysis-functionality/task-menu/variant-analysis)
  * [Annotate Variants](/icm/analyses/analysis-functionality/task-menu/variant-analysis/annotate-variants)
  * [Annotate Variants (SnpEff)](/icm/analyses/analysis-functionality/task-menu/variant-analysis/annotate-variants-snpeff)
  * [Annotate Variants (VEP)](/icm/analyses/analysis-functionality/task-menu/variant-analysis/annotate-variants-vep)
  * [Filter Variants](/icm/analyses/analysis-functionality/task-menu/variant-analysis/filter-variants)
  * [Summarize Cohort Mutations](/icm/analyses/analysis-functionality/task-menu/variant-analysis/summarize-cohort-mutations)
  * [Combine Variants](/icm/analyses/analysis-functionality/task-menu/variant-analysis/combine-variants)
  * [Filter variants with databases](/icm/analyses/analysis-functionality/task-menu/variant-analysis/filter-variants-with-databases)
  * [Create genotype matrix](/icm/analyses/analysis-functionality/task-menu/variant-analysis/create-genotype-matrix)
* [Combine multiomics data](/icm/analyses/analysis-functionality/task-menu/combine-multiomics-data)
  * [Combine 5-base methylation and variant data](/icm/analyses/analysis-functionality/task-menu/combine-multiomics-data/combine-5-base-methylation-and-variant-data)
* [Motif Detection](/icm/analyses/analysis-functionality/task-menu/motif-detection)
* [Biological interpretation](/icm/analyses/analysis-functionality/task-menu/biological-interpretation)
  * [Gene set enrichment](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/gene-set-enrichment)
  * [GSEA](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/gsea)
  * [Gene set ANOVA](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/gsea-1)
  * [Correlation Engine atlases](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/correlation-engine-atlases)
  * [Correlation Engine pathway](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/correlation-engine-pathways)
  * [Get targeted mRNA](/icm/analyses/analysis-functionality/task-menu/biological-interpretation/get-targeted-mrna)
* [Classification](/icm/analyses/analysis-functionality/task-menu/classification)
  * [Classify cell type](/icm/analyses/analysis-functionality/task-menu/classification/classify-cell-type)
  * [Train classifier](/icm/analyses/analysis-functionality/task-menu/classification/train-classifier)
  * [ScType](/icm/analyses/analysis-functionality/task-menu/classification/sctype)
* [Conversion](/icm/analyses/analysis-functionality/task-menu/conversion)
* [Region analysis](/icm/analyses/analysis-functionality/task-menu/region-analysis)
  * [Annotate regions](/icm/analyses/analysis-functionality/task-menu/region-analysis/annotate-regions)


# AI analysis methods summary

The AI analysis methods summary generates a structured draft of your analysis workflow in Illumina Connected Multiomics (ICM). It follows the task graph and uses task settings to summarize the methods used in your analysis.

You can use the summary to:

* Draft methods sections for manuscripts
* Document and report analyses
* Share workflows for handoffs and reproducibility

{% hint style="warning" %}
Review the generated text and citations before reuse. AI-generated content may be inaccurate.
{% endhint %}

#### Generate AI analysis methods summary

You can generate a methods summary at any point in an analysis. On the analysis page, click the **Summarize methods** button at the bottom of the page.

<figure><img src="/files/LbIzPGFSwmYGEx0BgYz8" alt=""><figcaption></figcaption></figure>

Select the task nodes to include in the summary. Then click **Generate summary**.

<figure><img src="/files/HchEFrsTTpiuyJaXlFvy" alt=""><figcaption></figcaption></figure>

The generated summary describes the selected analysis methods in plain language. It also includes relevant citations. The citations are pulled directly from the software documentation and are manually curated.

<figure><img src="/files/7A0zgLUIM1mRAoms6E6x" alt=""><figcaption></figcaption></figure>

To color the text by task, enable **Show task associations**.

<figure><img src="/files/le0o8BbTHpbFjp1noMuX" alt=""><figcaption></figcaption></figure>

To export the summary, click **Copy to clipboard**. To send feedback, use the thumbs up or thumbs down buttons in the summary panel.

<figure><img src="/files/4unjKvZwhsxnCS2fWfNV" alt=""><figcaption></figcaption></figure>


# AI analysis suggestions

Provides AI-powered recommendations for the next analysis step based on your current data.

{% hint style="warning" %}
This AI feature is in Beta testing. AI-generated content may be inaccurate.
{% endhint %}

## How AI suggestions are generated

The AI suggestions tool relies on a database of thousands of historical datasets. Once a node has been selected, the tool will determine the three most probable next steps in order of likelihood. It also provides an rationale for each choice to inform the user on how to proceed with their analysis.

## Generating an AI-guided analysis

* Click on a node to open the task menu and select '**Suggest next step**' (Figure 1).

<figure><img src="/files/00pWOTZ4VEeG7eCWzgzY" alt=""><figcaption><p>Figure 1. The AI suggestion tool is accessible from the context menu. Click <strong>Suggest next step</strong> to generate the suggestions.</p></figcaption></figure>

* Three suggestions will appear in the task menu and on the task graph. Each suggestion has an accompanying rationale, accessible by clicking '**Show more**'. Once the user has made a decision on how to proceed they can click '**Accept**' (Figure 2).

<figure><img src="/files/PgtcuBn84QEWI2VRJwZx" alt=""><figcaption><p>Figure 2. After clicking on <strong>Suggest next step</strong> the tool provides three potential next steps in order of likelihood. The nodes suggested will appear on the task graph. At this stage the use can select <strong>Accept</strong> for one of the suggested tasks or click <strong>Dismiss suggestions</strong>. The tool provides a rationale for each suggested task, click <strong>Show more</strong> to read the full explanation.</p></figcaption></figure>

* Once the user clicks '**Accept**' the task setup page will open. Note the disclaimer at the top of the page. Currently the AI tool does not recommend analysis-specific parameters. These remain the default set for the data type (Figure 3).

<figure><img src="/files/OqOqhiubo7mPageJ9Bjk" alt=""><figcaption><p>Figure 3. The task setup menu opens, note the disclaimer at the top of the page.</p></figcaption></figure>

* Clicking '**Finish**' runs the task as it normally would. A new node appears on the task graph.

### Additional notes

* The user is encouraged to provide feedback using the thumbs up/down icons (Figure 4). They can also provide written feedback for their choice. Clicking '**Cancel**' will still send the thumbs up/down selection as feedback. If you don't want to share your feedback click the 'x' in the feedback box.

<figure><img src="/files/qM9boBBHuM5il5oz0ImV" alt=""><figcaption><p>Figure 4. Feedback for each task suggested by the AI tool can be provided by clicking on the thumbs up/down icons. Users are also encouraged to provide feedback using the text box</p></figcaption></figure>

* Clicking '**Dismiss suggestions**' will trigger a similar feedback process.
* The tool will provide a confidence score for the choices presented. This can be low, medium or high; based on the historical data. Note that a low score can indicate all the options proposed are equally valid as a subsequent step, depending on the specific analysis (Figure 5).

<figure><img src="/files/80sS2dsogdFFrskCxOp2" alt=""><figcaption><p>Figure 5. Confidence score assigned to the AI suggestions.</p></figcaption></figure>


# Task actions

Left single clicking on any task (the rectangles) in the analysis pipeline will cause a Task Actions section to appear in the pop-up menu. This allows users to:

* Rerun tasks: rerun the selected task, the task dialog will pop-up and users can change parameters of the task. Previous downstream analysis of the selected task will not be rerun.
* Rerun with downstream tasks: rerun the selected task, the task dialog will pop-up, users can change the parameters of the current task and the downstream analysis will be rerun with the same configuration as the previous one.
* Edit description: the description of the task can be replaced by manually typing in string.
* Change color: choose a color to apply only on the selected task by clicking on **Apply.** Click **Apply to downstream** to change the selected task and the downstream pipeline color to the newly selected color.
* Delete task: this option is only available if the user is the owner of the project or the owner of the task. When a task is deleted, all downstream tasks, including tasks from other users, will be deleted. Users may check the box to choose to delete the task's output files. If delete output files is not checked, the task will be removed from the pipeline, but the output files of the task will remain on the disk.
* Restart task: this option is only available on failed tasks and requires an admin role to perform, but does not require that you have a user account. Since you are logged in as an admin, restarting a task will not take up a concurrent seat and the disk space consumed by the output files will count towards the original owner of the task's storage space.


# Data summary report

The *Data summary report* in Connected Multiomics provides an overview of all tasks performed as part of a pipeline. This is particularly useful for report writing, record keeping and revisiting projects after a long period of time.

## Viewing the Data Summary Report

Click on an output data node under the *Analyses* tab of a project and choose **Data summary report** from the context sensitive menu on the right. The report will include details from all of the tasks upstream of the selected the node. If tasks have been performed downstream of the selected data node, they will not be included in the report.

<div align="left"><figure><img src="/files/tRa3lXWgdfgrJMrLLzph" alt=""><figcaption></figcaption></figure></div>

Each task will appear as a separate section on the *Data summary report*.

<div align="left"><figure><img src="/files/lZqo6kHo0Tr98W7rMW60" alt=""><figcaption></figcaption></figure></div>

Click the arrow ( ![](/files/fBQyURnEW8WNT7OoKRgw) / ![](/files/a0C51ChzT9RV6oYEqR5g)) to expand and collapse each section. When expanded, the task name, user that performed the task, start date and time, duration and the output file size are displayed. To view or hide a table of task settings, click **Show/hide details**.

## Saving the Data Summary Report

The *Data summary report* can be saved in different formats via the web browser. The instructions below are for Google Chrome. If you are using a different browser, consult your browser's help for equivalent instructions.

### Save as a PDF

On the *Data summary report*, expand all sections and show all task details. Right-click anywhere on the page and choose **Print...** from the menu or use **Ctrl+P** (**Command+P** on Mac).

<div align="left"><figure><img src="/files/PZyGWUiQPlMK7FkTXPmY" alt="" width="233"><figcaption></figcaption></figure></div>

In the print dialog, set the destination to **Save as PDF**. Select the **Background graphics** checkbox (optional), click the blue **Save** button and choose a file location on your local machine.

<div align="left"><figure><img src="/files/FQgURLgTPOqEBnGDxFLv" alt="" width="278"><figcaption></figcaption></figure></div>

The PDF can be attached to an email and/or opened in a PDF viewer of your choice.

### Save as HTML

On the *Data summary report*, right-click anywhere on the page and choose **Save as…** from the menu or use **Ctrl+S** (**Command+S** on Mac). Choose a file location on your local machine and set the file type to **Web Page, Complete**.

The HTML file can be opened in a browser of your choice.


# QA/QC

Connected Multiomics contains a number of quality control tools and reports that can be used to evaluate the current status of your analysis and decide downstream steps. Quality control tools are organized under the Quality Assurance / Quality Control (QA/QC) section of the context-sensitive menu and are available for different type of data nodes.

This section will illustrate:

* [Feature distribution](/icm/analyses/analysis-functionality/task-menu/qa-qc/feature-distribution)
* [Imported count matrix report](/icm/analyses/analysis-functionality/task-menu/qa-qc/feature-distribution-1)
* [Single-cell QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/single-cell-qa-qc)
* [Cell barcode QA/QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/cell-barcode-qa-qc)
* [5-base Methylation QC](/icm/analyses/analysis-functionality/task-menu/qa-qc/5-base-methylation-qc)

In addition to the tools listed above, many other functionalities can also be interpreted in sense of quality control. For instance, principal components analysis, hierarchical clustering (on sample level), variant detection report, and quantification report.


# Feature distribution

The Feature distribution plot visualizes the distribution of features in a counts matrix data node.

## Running Feature distribution

To run Feature distribution:

* Click a counts data node
* Click the **QA/QC** section of the toolbox
* Click **Feature distribution**

A new task node is generated with the Feature distribution report.

## Feature distribution plot configuration

The Feature distribution task report plots the distribution of all features (genes or proteins) in the input data node with one feature per row. Features are ordered by average value in descending order.

<div align="left"><figure><img src="/files/1foqV3MfJOKVkt5BuFnN" alt=""><figcaption></figcaption></figure></div>

The plot can be configured using the panel of the left-hand side of the page.

#### Filter

Using the filter, you can choose which features are shown in the task report.

The *Manual* filter lets you type a feature ID (such as a protein ID) and filter to matching features by clicking + . You can add multiple feature IDs to filter to multiple features.

<div align="left"><figure><img src="/files/ucmudo8ZlcmLvZlcjsQF" alt=""><figcaption></figcaption></figure></div>

The *List* filter lets you filter to the features included in a feature list. To learn more about feature lists, please see [List management](/icm/analyses/analysis-functionality/settings/lists).

#### Plot type

Distributions can be plotted as histograms, which is the default setting, with the x-axis being the expression value and the y-axis the frequency, or as a strip plot, where the x-axis is the expression value and the position of each cell/sample is shown as a thin vertical line, or strip, on the plot.

<div align="left"><figure><img src="/files/ldTw3iefzYKim3dMX4Ob" alt=""><figcaption></figcaption></figure></div>

To switch between plot types, use the *Plot type* radio buttons.

Mousing over a dot in the histogram plot gives the range of feature values that are being binned to generate the dot and the number of cells/samples for that bin in a pop-up.

<figure><img src="/files/CcKNxBwxoEVyM7UVbqmb" alt=""><figcaption></figcaption></figure>

Mousing over a strip shows the sample ID and feature value in a pop-up. If there are multiple cells/samples with the same value, only one strip will be visible for those cells/samples and the mouse-over will indicate how many cells/samples are represented by that one strip.

<figure><img src="/files/Jwx9UeJdrkxOlpjL5P5h" alt=""><figcaption></figcaption></figure>

Clicking a strip will highlight that cell/sample in all of the plots on the page. The grey dot in each strip plot shows the median value for that feature. To view the median value, mouse over the dot.

#### Page

To navigate between pages, use the Previous and Next buttons or type the page number in the text field and click Enter on your keyboard.

The number of features that appear in the plot on each page is set by the *Items per page* drop-down menu. You can choose to show 10, 25, or 50 features per page.

<div align="left"><figure><img src="/files/Na2DTn0uIxMBOuPqFQQT" alt=""><figcaption></figcaption></figure></div>

#### Color by

You can add attribute information to the plots using the *Color by* drop-down menu.

For histogram plots, the histograms will be split and colored by the levels of the selected attribute. You can choose any categorical attribute.

<div align="left"><figure><img src="/files/9ARQCOGWqkoApHm4y8Cd" alt=""><figcaption></figcaption></figure></div>

For strip plots, the sample/cell strips will be colored by the levels or values of the selected attribute. You can choose any categorical or numeric attribute.

<div align="left"><figure><img src="/files/NdStbloaNUtw9qf6o0eH" alt=""><figcaption></figcaption></figure></div>


# Imported count matrix report

The imported count matrix report is a summary report on sample distribution information of imported counts matrix data e.g. ILMN miRNA count matrix data

* Click an imported miRNA data node
* Click the **QA/QC** section of the toolbox
* Click **Imported count matrix report**

A new task node is generated with the Imported count matrix report. Double click on the report to open it.

<figure><img src="/files/1oFjZzAtDZax5Qz3LKTA" alt=""><figcaption></figcaption></figure>

In the Feature distribution table title, it displays the size of the matrix, number of samples and number of features. Each row is a sample in the table, columns contains descriptive statistics of features in the sample.

If there are less than 30 samples in the data node, a bar chart is presented. Each bar is a sample. The X-axis is the read count range, Y axis is the number of features within the range. Hovering your mouse over the bar displays the following information:

<div align="left"><figure><img src="/files/S4ogPAI8VwKrj0kiUVzL" alt=""><figcaption></figcaption></figure></div>

* Sample name
* Range of read counts, “\[ “represent inclusive, “)” represent exclusive, e.g. \[0,0] means 0 read counts; (0,10] means the range is greater than 0 count but less than and equal to 10 counts.
* Number of features within the read count range
* Percentage of the features within the read count range

A Box-whisker plot is displayed below the bar chart. In the box-whisker plot, each box is a sample on X-axis, the box represents 25th and 75th percentile, the whiskers represent 10th and 90th percentile, Y-axis represents the feature counts, when you hover over each box, detailed sample information is displayed

<div align="left"><figure><img src="/files/4RBKqRObevnI2OlS8I63" alt="" width="563"><figcaption></figcaption></figure></div>

* Sample name
* Range of read counts, “\[ “represent inclusive, “)” represent exclusive
* Number of features within the read count range in the sample


# Single-cell QA/QC

The Single-cell QA/QC task in Connected Multiomics enables you to visualize several useful metrics that will help you include only high-quality cells. To invoke the Single-cell QA/QC task:

* Click a **Single cell counts** data node
* Click the **QA/QC** section of the task menu
* Click **Single cell QA/QC**

By default, all samples are used to perform QA/QC. You can choose **Split by sample** in *Grouping* option to perform QA/QC separately for each sample.

You will be prompted to choose the genome assembly and annotation file by the Single cell QA/QC configuration dialog.

<div align="left"><figure><img src="/files/OTjmyc2UJIVocjEvd0NG" alt=""><figcaption></figcaption></figure></div>

Note, it is still possible to run the task without specifying an annotation file. If you choose not to specify an annotation file, the detection of mitochondrial counts will not be possible. The annotation file should match the same annotation file used in the upstream analysis.

The Single cell QA/QC task report opens in a new data viewer session. Four dot and violin plots showing the value of every cell on the canvas: counts per cell, detected features per cell, the percentage of mitochondrial counts per cell (when annotation file contains the genes on MT chromosome), and the percentage of ribosomal counts per cell (human and mouse only).

<figure><img src="/files/ws8bkTyfhcaVgzNsXQZ7" alt=""><figcaption></figcaption></figure>

If your cells do not express any mitochondrial genes or an appropriate annotation file was not specified, the plot for the percentage of mitochondrial counts per cell will be non-informative.

Mitochondrial genes are defined as genes located on a mitochondrial chromosome in the gene annotation file. The mitochondrial chromosome is identified in the gene annotation file by having "M" or "MT" in its chromosome name. If the gene annotation file does not follow this naming convention for the mitochondrial chromosome, Connected Multiomics will not be able to identify any mitochondrial genes.

Ribosomal genes are defined as genes that code for proteins in the large and small ribosomal subunits. Ribosomal genes are identified by searching their gene symbol against a list of 89 L & S ribosomal genes taken from [HGNC](https://www.genenames.org/). The search is case-insensitive and includes all known gene name aliases from HGNC. Identifying ribosomal genes is performed independent of the gene annotation file specified.

Total counts are calculated as the sum of the counts for all features in each cell from the input data node. The number of detected features is calculated as the number of features in each cell with greater than zero counts. The percentage of mitochondrial counts is calculated as the sum of counts for known mitochondrial genes divided by the sum of counts for all features and multiplied by 100. The percentage of ribosomal counts are calculated as the sum of counts for known ribosomal genes divided by the sum of counts for all features and multiplied by 100.

Each point on the plots is a cell. All cells from all samples are shown on the plots. The overlaid violins illustrate the distribution of cell values for the y-axis metric.

The appearance of a plot can be configured by selecting a plot and adjusting the *Configure* settings in the panel on the left. Here are some suggestions, but feel free to explore the other options available:

* Open **Axes** and change the Y-axis scale to **Logarithmic**. This can be helpful to view the range of values better, although it is usually better to keep the Ribosomal counts plot in linear scale.

<figure><img src="/files/3rQh5sE7eyKV4DmBnLFw" alt=""><figcaption></figcaption></figure>

* Within **Style** switch on *Summary* **Box & Whiskers**. Inspecting the median, Q1, Q3, upper 90%, and lower 10% quantiles of the distributions can be helpful in deciding appropriate thresholds.

<figure><img src="/files/16woOtihqib8SvL3STOU" alt=""><figcaption></figcaption></figure>

High-quality cells can be selected using **Select & Filter**, which is pre-loaded with the selection criteria, one for each quality metric.

<figure><img src="/files/MhKLSTSlRMsrH20G6fWw" alt=""><figcaption></figcaption></figure>

Hovering the mouse over one of the selection criteria reveals a histogram showing you the frequency distribution of the respective quality metric. The minimum and maximum thresholds can be adjusted by clicking and dragging the sliders or by typing directly into the text boxes for each selection criteria.

<figure><img src="/files/oIKQIlHGCjUnUWpPVG01" alt=""><figcaption></figcaption></figure>

Alternatively, Pin histogram to view all of the distributions at one time to determine thresholds with ease.

<figure><img src="/files/oMp8GCbG1Hi4RbgtaXMy" alt=""><figcaption></figcaption></figure>

Adjusting the selection criteria will select and deselect cells in all three plots simultaneously. Depending on your settings, the deselected points will either be dimmed or gray. The filters are additive. Combining multiple filters will include the intersection of the three filters. The number of cells selected is shown in the figure legend of each plot.

<figure><img src="/files/QhgnIVh8G2XRSJt1r4eV" alt=""><figcaption></figcaption></figure>

To filter the high-quality cells, click the include selected cells icon <img src="/files/kzv2UrEA4UBu8MC2dOd5" alt="" data-size="original"> in **Filter** in the top right of **Select & Filter**, and click **Apply observation filter**...

<div align="left"><figure><img src="/files/v5DhwwczBRvPJZTNpHX3" alt=""><figcaption></figcaption></figure></div>

Select the **input data node** for the filtering task and click **Select**.

<figure><img src="/files/I3iZRuFO4pQORe1hIrXk" alt=""><figcaption></figcaption></figure>

A new data node, *Filtered counts*, will be generated under the *Analyses* tab.

Double click the *Filtered counts* data node to view the task report. The report includes a summary of the count distribution across all features for each sample; a detailed breakdown of the number of cells included in the filter for each sample; and the minimum and maximum values for each quality metric (expressed genes, total counts, etc) across the included cells for each sample.

<figure><img src="/files/LdKgtHwGim8aPPmHxUZf" alt=""><figcaption></figcaption></figure>


# Cell barcode QA/QC

The Cell barcode QA/QC task lets you determine whether a given cell barcode is associated with a cell. This is an important QC step in all droplet-based single cell RNA-seq experiments, where all barcodes are sequenced.

To invoke Cell barcode QA/QC:

* Click a **Single cell counts** data node
* Click the **QA/QC** section of the task menu
* Click **Cell barcode QA/QC**

The task can be performed with or without the EmptyDrops method enabled.

## Cell Barcode QA/QC without EmptyDrops

To perform the task without the EmptyDrops method enabled, leave the checkbox unchecked and click **Finish**.

Note: Data imported from DRAGEN result is recommended to use this option since barcode with 0 counts are filtered out.

<figure><img src="/files/fblEzBI8kRfuIR64tUDb" alt=""><figcaption></figcaption></figure>

The Cell barcode QA/QC task report is a plot. X-axis is the barcodes ranked by their UMI counts. Y-axis is the UMI counts in the barcode. This type of plot is often referred to as a knee plot.

<figure><img src="/files/csXzBkpInUX3dpuIPdVI" alt=""><figcaption></figcaption></figure>

The knee plot is used to choose a cutoff point between barcodes that correspond to cells and barcodes that do not if the imported raw count data without any barcode filtering performed upstream. Connected Multiomics automatically calculates an inflection point, shown by the vertical line on the graph. Barcodes designated as cells are shown in blue while barcodes designated as without cells (background) are shown in grey.

The cutoff can be adjusted by dragging the vertical line across the graph or by using the text fields in the *Filter* panel on the left-hand side of the plot. Using the *Filter* panel, you can specify the number of cells or the percentage of reads in cells and the cutoff point will be adjusted to match your criteria. The number of cells and the percentage of counts in cells is adjusted as the cutoff point is changed. To return to the automatically calculated cutoff, click **Reset sample filter**.

The percentage of counts in cells and median counts per cell are useful technical quality metrics that can be consulted when optimizing sample handling, cell isolation techniques, and library preparation.

One knee plot is generated for each sample. In projects with multiple samples, *Next* and *Back* buttons will appear at the top left of the plot, to enable navigation between sample knee plots. Manual filters must be set separately for each sample. This is typically used when the user expects a certain number of cells to be processed, like in experiments where droplets were loaded with a predefined number of cells.

To return to the knee plot view, click **Back to filter**. To apply the filter and run the Filter barcodes task, click **Apply filter**. A Filtered counts data node will be generated.

## Cell Barcode QA/QC with EmptyDrops

{% hint style="warning" %}
If your data has already been filtered to remove barcodes with low total counts, this method will not be suitable. This method requires empty barcodes to be present in the single cell count matrix, in order to estimate the ambient RNA profile.
{% endhint %}

The EmptyDrops method (1) uses a statistical test to identify which barcodes correspond to real cells and empty droplets. An ambient RNA expression profile is estimated from barcodes below a specified total UMI count threshold, using the Good-Turing algorithm. The expression profile of each barcode above the low-count threshold is then tested for deviations from the ambient profile. Real cells are expected to have a low p-value, indicating a significant deviation from the expected background noise level. False discovery rate (FDR) correction is applied to all the p-values and those falling equal to or below the specified FDR level are detected as real cells. This can allow for the detection of additional cells that would otherwise be discarded due to a low total UMI count.

In addition, a knee point threshold will be calculated to identify cells with a very high total UMI count. It's possible that some barcodes with a high total UMI count will not pass the EmptyDrops significance test. This could be due to biases in the ambient RNA profile, leading to a non-significant difference between a barcode's expression profile vs the ambient profile. To protect against this issue, it is advisable to use the EmptyDrops results in conjunction with the knee point filter, on the assumption that barcodes with a very high total UMI count will always correspond to real cells. Note, the knee point will be more conservative than the inflection point calculated by Connected Multiomics when the EmptyDrops method is not enabled.

To perform the task with the EmptyDrops method, **check** the checkbox, configure the additional options, and click **Finish.**

<figure><img src="/files/DlhjTTGuIjHbHjRLOckJ" alt=""><figcaption></figcaption></figure>

**Ambient count threshold**

Barcodes with a total UMI count equal to or below this threshold will be used to create the ambient RNA expression profile to estimate background noise. The default is set to 100, which is reasonable for most data.

**FDR threshold**

Barcodes equal to or below this FDR threshold show a significant deviation from the ambient profile and can therefore be considered real cells. Increasing this value will result in more cells, but will also increase the number of potential false positives.

**Random generator seed**

This is used for performing Monte Carlo simulations to determine p-values. To reproduce results, use the same random seed for all runs.

There are additional metrics on the left of the plot in the report.

<figure><img src="/files/rWBMKLy5ylmu3F0q9pM8" alt=""><figcaption></figcaption></figure>

The number of actual cells detected by the EmptyDrops test and the knee point filter are shown above the Venn diagram on the left. In the above example plot 3,189 barcodes are above the knee point filter (represented by the vertical blue line on the plot) and 2,657 barcodes passed the significance test in EmptyDrops. The overlap between these sets of barcodes is represented by the Venn diagram.There are 1,583 barcodes pass the significance test in EmptyDrops and have a high total UMI count above the knee point filter; 1,606 barcodes have a very high total UMI count with no significant difference from the ambient profile in EmptyDrops; 1,074 barcodes fall below the knee point but are still significantly different from the ambient profile.

The number of cells included by the knee point filter can be adjusted either by click on the plot to change the position of the vertical blue line or by typing a different number of cells into the text box on the left.

The total number of cells is shown in the text box on the left. By default, this will be all of the cells detected by the knee point filter plus the extra cells detected by EmptyDrops. In the example, there are 3,189 cells with a high total UMI count plus the additional 1,074 cells from EmptyDrops (total = 4,263).

Different sections of the Venn diagram can be selected/deselected to include/exclude barcodes. For example, clicking the '1,606' section of the Venn diagram will deselect those barcodes. Now, the only cells that will pass the filter will be the significant ones from EmptyDrops.

<div align="left"><figure><img src="/files/yrcIvkT872k12nCWrr4T" alt=""><figcaption></figcaption></figure></div>

## References

1. Lun, A., Riesenfeld, S., Andrews, T. *et al.* EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data. Genome Biol. 2019; 20: 63.


# 5-base Methylation QC

The **5-base methylation QC** task in the Connected Multiomics enables you to visualize sample-level QC metrics that describe reads mapping quality and CpG methylation calling. The QC metrics are extracted from the DRAGEN analysis metric files that were ingested into the study as required files for 5-base DNA Prep data analysis in the Connected Multiomics. To invoke the **5-base methylation QC** task:

* At **Analyses** page, click on the **5-base Methylation** node.
* Click **QA/QC** section in the context-sensitive task menu on the right.
* Click **5-base methylation QC**.

There is no parameters setting required for the 5-base methylation QC task. After click on the 5-base methylation QC task from the context-sensitive task menu, a task node called **5-base methylation QC report** is initiated. When completed, double-click on the **5-base methylation QC report** task node to open the QC report in a data viewer. The QC report consists of plots and tables organized in 2 sheets. Click on sheet name at the bottom of the data viewer to navigate from one sheet to another.

## Metrics

Sheet **Metrics** shows sample-level QC metrics plot. Each sample is a data point, they are randomly spead out on x-axis. The QC metric is represented by y-axis. Each plot is overlay with a violin plot to show distribution of the QC metrics.

<figure><img src="/files/ZtvyPbij7Mhec7AZ8d7L" alt=""><figcaption></figcaption></figure>

* Percent methylation in samples: Percentages of CpG methylation in samples.
* Percent methylation in unmethylated control: Percentage of CpG methylation in the unmethylated control (lambda). Low value indicates good quality.
* Percent methylation in methylated control: Percentage of CpG methylation in the methylated control (pUC19). High value indicates good quality.
* Percent duplicate reads: Percentage of duplicate marked reads, as a result of PCR amplification.
* Percent mapped reads: Percentage of mapped reads, indicate the alignment rate.
* Average autosomal coverage: Mean autosomal coverage across the whole genome. Higher coverage indicates the counts of methylated/unmethylated more accurately reflects the true methylation amount at any particular site.
* QC metrics table: Text representations of the QC metrics plots.

In Metric sheet, samples can be selected using **Selection > Select & Filter. T**he **Select & Filter** dialog is pre-loaded with the selection criteria, one for each QC metric.

<figure><img src="/files/I4KeQj3PNIRfB8PC5WjR" alt=""><figcaption></figcaption></figure>

Hovering the mouse over one of the selection criteria reveals a histogram showing you the frequency distribution of the respective QC metric. The minimum and maximum thresholds can be adjusted by clicking and dragging the sliders or by typing directly into the text boxes for each selection criteria.

<figure><img src="/files/2rfw412p3kPMdIaIXAJA" alt=""><figcaption></figcaption></figure>

Adjusting the selection criteria will select and deselect samples in all 6 plots simultaneously. Depending on your settings, the deselected points will either be dimmed or gray. The filters are additive. Combining multiple filters will include the intersection of the the filters. The number of samples selected is shown in the figure legend of each plot.

To filter the dataset to the selected samples, click the **include selected points** icon ( <img src="/files/yijHJ6ZMhncjLGHZQiWl" alt="" data-size="line"> ) in **Filter** in the top right of **Select & Filter**, and click **Apply observation filter**...

<figure><img src="/files/jsy5NiRCzCbbo0XNCIvc" alt=""><figcaption></figcaption></figure>

Select the **input data node** for the filtering task and click **Select**.

<figure><img src="/files/9NT12kg1HvqxapGILTuI" alt=""><figcaption></figcaption></figure>

A new data node, *Filtered samples*, will be generated under the *Analyses* tab.

## M-bias

Sheet **M-bias** shows M-bias plots for methylation level and coverage across positions on read1 and read2. The M-bias should be consistent across all positions. It is common for the first/last 10 bases to have un-even methylation due to end-repair and sequencing artifacts.

<figure><img src="/files/1CRSwBFJGrh7c126LIoR" alt=""><figcaption></figcaption></figure>

All plots in one data viewer screen can be downloaded into local computer as a single image by clicking **Export** button on the top of the screen. To download an individual plot into local computer, select the plot, click **Plot** button from the left panel within the plot, then click **Export**, follow the wizard to set image file format, image size, and resolution.

<figure><img src="/files/C7Y2loGYG5S28TIcD3ic" alt=""><figcaption></figcaption></figure>


# Annotation/Metadata

This section has tools that are useful in managing and understanding single cell data, especially for downstream analysis. To invoke Annotation/Metadata tools, click on any **Single cell counts** data node. These include the following tasks:

* [Annotate cells](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells)
* [Annotate features](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells-1)
* [Publish cell attributes to project](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/publish-cell-attributes-to-project)
* [Attribute report](/icm/analyses/analysis-functionality/task-menu/annotation-metadata/annotate-cells-2)


# Annotate cells

If you have attribute information about your cells, you can use the Annotate cells task in Connected Multiomics to apply this information to the data. Once applied, these can be used like any other attributes, and thus can be used for cell selection, classification and differential analysis.

To run Annotate cells:

* Click a Single cell counts data node
* Click the **Annotation/Metadata** section in the toolbox
* Click **Annotate cells**

You will be prompted to specify annotation input options:

* Single file (all samples): requires one .txt file for all cells in all samples. Each row in the file represents a barcode and at least one barcode column which will match the barcodes in your data. It also requires a column containing Sample ID which must match the Sample name in the Metadata tab of your project.
* File per sample: requires the format of all of the annotation files to be the same. Each file has barcodes on rows, it requires one barcode column that will match the barcodes in your data in that sample. All files should have the same set of columns, column headers are case sensitive.

Browse to each sample file on the server to specify annotation files for all of the samples in the dialog.

<figure><img src="/files/8ctbbhpk9Fdao5p9uOue" alt=""><figcaption></figcaption></figure>

If you would like to annotate your matrix features with a gene annotation file, you can choose an annotation file at the bottom on the dialog. You can choose any gene/feature annotation available on the server. If a feature annotation is selected, the percentage of mitochondrial reads will be calculated using the selected annotation file.

<figure><img src="/files/Gu8v7UeYGTCSfPX2cCmk" alt=""><figcaption></figcaption></figure>

* Click **Next** to continue

The next dialog page previews the attributes found in the annotations text file.

<figure><img src="/files/CSdEDQZXdiKG8X4YfqIO" alt=""><figcaption></figcaption></figure>

You can choose which attributes to import using the check-boxes, change the names of attributes using the text fields, and indicate whether an attribute with numbers is categorical or numeric.

* Click **Finish** to import the attributes

A new data node, Annotated single cell counts, will be generated . The annotations will be available in downstream analysis tasks.




---

[Next Page](/llms-full.txt/1)

