> For the complete documentation index, see [llms.txt](https://help.multiomics.illumina.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.multiomics.illumina.com/icm/analyses/walkthroughs/perturb-seq.md).

# Perturb-seq

## Illumina Connected Multiomics

Illumina Connected Multiomics (ICM) is available for further tertiary analysis of Illumina Single Cell Transcriptomics Perturb-seq data and other multiomic data.

## Getting Started <a href="#getting-started" id="getting-started"></a>

Refer to the following links to the ICM user guide to get started with ICM:

* [Registration and Login](https://help.connected.illumina.com/icm/introduction/icm)
* [Data Inputs](https://help.connected.illumina.com/icm/introduction/data-inputs)
* [Creating a Study from a BioInsight Platform Core Project](https://help.connected.illumina.com/icm/studies/create-study)
* [Viewing Results and Navigating in ICM](https://help.connected.illumina.com/icm/studies/enter-study)

## Demo Data

Demo data that can be used to follow along with this walkthrough is found in the Connected Multiomics Demo Data repository and can be found under Single cell > Perturb-seq.

Each sample will need the following files as input (when specifying sample1 as the sample id):

* sample1.scRNA.filtered.matrix.mtx.gz
* sample1.scRNA.filtered.barcodes.tsv.gz
* sample1.scRNA.filtered.features.tsv.gz
* sample1.scRNA.feature\_barcode\_reference.csv
* sample1.scRNA.positive\_cell\_guide\_assignments.csv

This data is one sample consisting of Human A549 cell line. As determine by DRAGEN, each cell barcode present is a positive cell assignment (cells below a certain threshold are filtered out). The num\_transcripts (number of transcripts) represents coverage of the guide as read counts.

## Default Single Cell Perturb-seq Analysis

The default Perturb-seq analysis runs a pre-defined pipeline on each sample and presents analysis results in visualizations in a *Data viewer*.

### Creating a Default Analysis

After adding data to a study, follow the following steps to create a Default analysis.

* Select the samples to include in the analysis
* Click on *+ Create analysis*

<figure><img src="/files/sPZDWFiHCukYekvFikRm" alt=""><figcaption><p>Select samples and Create analysis</p></figcaption></figure>

* In the pop-up window, provide a name for the analysis
* select *Default: Illumina Single Cell Transcriptomics Perturb-seq* from the dropdown as the *Analysis type*
* click on the *Run Analysis* button

<figure><img src="/files/EAtekvtlRUO9A0au94Ut" alt=""><figcaption><p>Define Default or Custom analysis</p></figcaption></figure>

### View Default Analysis Results

The analysis status will change to *Complete* when the analysis has finished.

* Click on the analysis tile to open results

The analysis opens to the analysis task graph.

<figure><img src="/files/SXUghi9r3OJo1f30l3FU" alt=""><figcaption><p>The task graph shows the analysis pipeline</p></figcaption></figure>

The first task run on the imported data, *Split by feature type,* is to split the data into different features: *CRISPR Direct Capture (gRNAs)* and *Gene Expression*. This is because having *gRNAs* included in standard analyses for both clustering and differential expression can lead to unexpected clustering results. The remainder of the pipeline follows the standard scRNA-seq workflow. Additional details for each task are available via the [Single Cell walkthrough.](/dragen-single-cell-rna/tertiary-analysis/illumina-connected-multiomics.md)

### View Summary report

* Double-click the *Summary report* to open a Data viewer session.

<figure><img src="/files/8gnGzWJdka838JZEhT1D" alt=""><figcaption><p>Double-click to view the Summary report</p></figcaption></figure>

The Summary report shows the *Gene expression* tab and *CRISPR Direct Capture* tab from the Perturb-seq samples; these can be toggled at the bottom of the data viewer session.

#### Gene expression tab

The *Gene expression* tab consists of one UMAP colored by cluster IDs, a biomarker table, a cell composition pie chart and the distributions of *Total count* and *Expressed genes* within different clusters.

<figure><img src="/files/oA8QrGacRNoiKnX58GXk" alt=""><figcaption><p>Gene expression data from the Perturb-seq samples</p></figcaption></figure>

The plots can be configured by selecting the *Configure* option from the toolbox on the left within each plot. Here is one example that converts the above UMAP to a feature plot by recoloring single cells with a ‘feature’: one of the *gRNAs* (*PDCD10\_4*). In the feature plot, all cells in red carry the perturbation of the *gRNA*, while all the non-perturbation cells are in grey for this specific guide.

<figure><img src="/files/7ztHAQQ8v1SowWWGYCO3" alt=""><figcaption></figcaption></figure>

#### CRISPR Direct Capture tab

The CRISPR Direct Capture tab consists of the frequency of the total number of features (guides) in the cells as well as the frequency of each of the top 10 features with highest sum.

<figure><img src="/files/q4WuNVWfrfLHHXid88Oc" alt=""><figcaption><p>CRISPR Direct Capture data from the Perturb-seq samples</p></figcaption></figure>

## Additional information for custom analysis

### Example pipeline

A typical custom pipeline includes both the gene expression analysis and CRISPR Direct capture analysis.

<figure><img src="/files/5qPFxhJwcGaBcxefmo8q" alt=""><figcaption><p>Example Custom Perturb-seq analysis</p></figcaption></figure>

### [**Single cell QA/QC task**](/icm/analyses/analysis-functionality/task-menu/qa-qc/single-cell-qa-qc.md)

It is optional to filter only high quality cells based on the total count, detected features, % mitochondria, and % ribosomal counts.

* Select the **Single cell QA/QC task**, under QA/QC in the task menu
  * This results in the *Filtered cells* results node

### [Filter cells](/icm/analyses/analysis-functionality/task-menu/filtering/filter-groups-samples-or-cells.md#filter-by-metadata) (observations)

It is optional to filter the cells. In the next step (shown on the pipeline as the *Filter observations* task), we will filter the cells by metadata to include cells with a number of features = 1.0. We chose to do this because this will filter our the cells to include 1 guide per cell and remove the cells with more than 1 guide or cells with no guides to limit combinatorial effects. There are cases where you might keep multi-guide cells and this should be based on your experimental design and research question.

* From the *Filtered cells* node, open the **Filter cells** task under *Filtering* in the task menu
* Filter by *Metadata* to *Include* the *num\_features = 1.0*
* Click **Finish**

<figure><img src="/files/9n0DNKmRXt0Pl2keBwWT" alt=""><figcaption><p>num_features is CRISPR metadata. We are not filtering gene expression data in this step</p></figcaption></figure>

This results in a counts node filtered to cells including 1.0 number of features (CRISPR guides).

<figure><img src="/files/jbELORBAe5uUKgMwPs5c" alt=""><figcaption><p>Counts node with cells including 1.0 CRISPR feature guide per cell</p></figcaption></figure>

### [Split by feature type](/icm/analyses/analysis-functionality/task-menu/pre-analysis-tools/split-matrix.md)

Split by feature type is used to split and analyze the *CRISPR Direct Capture* and *Gene expression* features independently.

* Select the **Split by feature type** task under *Pre-analysis tools* in the task menu

This results in a node containing the CRISPR Direct Capture features and node containing the Gene expression features.

<figure><img src="/files/GsqF6fuKPGKyTvBCbPEf" alt=""><figcaption><p>Split by feature type to analyze the CRISPR and Gene expression data</p></figcaption></figure>

### [Create the CRISPR Direct Feature Data viewer session](#crispr-direct-capture-tab)

* Click the **Data viewer** tab

<figure><img src="/files/IHXIdlvpxvDHGGY1WmMD" alt=""><figcaption><p>Click the Data viewer tab</p></figcaption></figure>

* Click **Create new view**

<figure><img src="/files/cLVgi9IMc5ZVqtZGlSx7" alt=""><figcaption><p>Create new Data Viewer session</p></figcaption></figure>

* Click **Setup**
* **+ New plot**

<figure><img src="/files/O1ImqhsE6sO9E0BSam8y" alt=""><figcaption><p>Add new plot to Data viewer session</p></figcaption></figure>

* Choose **Bar chart** as the plot type
* Select *Filtered Features* node as the data to plot
* Choose *num\_features* as the data attribute
* Modify the **Axes** configurations as shown below to match the settings from the CRISPR Direct Feature Data viewer session tab

<figure><img src="/files/itjcdjmIHMiKYQuV6gD0" alt=""><figcaption><p>Modify Axes settings to change the visualization</p></figcaption></figure>

* Add another bar chart but this time choose *ATM/design\_3* as the content data to plot
* Click **Add**
* Modify the **Axes** settings to match below

<figure><img src="/files/KSlaoaFXJFwfAq1sh3A3" alt=""><figcaption></figcaption></figure>

* **Duplicate** the plot 9 times

<figure><img src="/files/g9a7yTbsWS44174s328o" alt=""><figcaption><p>Duplicate the plot</p></figcaption></figure>

* Change the **Axes** *content data* to the other top features.

Ctrl+Click and drag to resize views without snapping. Click the **Save** button to Save over an existing session or Save As to save a new Data viewer session.

<figure><img src="/files/e609gacSvEXDsF33I7NW" alt=""><figcaption><p>Resize views and Save the Data viewer session</p></figcaption></figure>

### Select cells by criteria

**Selection** > **Select & Filter** can be used to select cells with specific targets. Our input data from DRAGEN contains a Feature call label and Target gene name. Either of these can be used to select cells matching the criteria.

* Select the criteria and enter the name of interest to select the cells

<figure><img src="/files/Cv1I2NtGTKKsAcuHWi0v" alt=""><figcaption></figcaption></figure>

These cells could be labeled and classified using **Selection** > **Classify**.

### Classify the Non-Targeting control and Perturbed cells population

Depending on your experimental design and research question there can be different populations of cells that could be used as controls. This could be non-targeting controls, cells with no detected guide, or all cells carrying guides other than the guide of interest. In this data, we have already filtered to the cells containing one CRISPR guide, so we will not use cells with no detected guide as the control. Below we will demonstrate classifying the non-targeting controls and perturbed cell populations.

* [Select the Non-Targeting Control cell population](#select-cells-by-criteria)
* Classify the selection as Non-targeting control
* Navigate to **Selection > Invert selection**

<figure><img src="/files/YQ3NITkxjQi0Cfrt1yCi" alt=""><figcaption></figcaption></figure>

* **Classify** the selection as Perturbed cells

Use the **Apply classifications** button to make the cell-level attribute (Classification) available from all data nodes within the analysis, including for use with differential analysis.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://help.multiomics.illumina.com/icm/analyses/walkthroughs/perturb-seq.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
