Use Pycytominer in Galaxy for processing high dimensional image-based readouts

Author(s) orcid logoRiccardo Massei avatar Riccardo Massei
Overview
Creative Commons License: CC-BY Questions:
  • How do I process high-dimensional readouts using Galaxy?

  • How can I process CellProfiler or DeepProfiler feature readouts with Pycytominer in Galaxy?

Objectives:
  • Learn how to execute the five main Pycytominer steps

  • Create a complete workflow for processing data with Pycytominer

Requirements:
Time estimation: 1 hour
Level: Intermediate Intermediate
Supporting Materials:
Published: Oct 8, 2026
Last modification: Oct 8, 2026
License: Tutorial Content is licensed under Creative Commons Attribution 4.0 International License. The GTN Framework is licensed under MIT
version Revision: 1

High-content imaging screens generate thousands of microscopy images capturing how cells respond to genetic or chemical perturbations. In particular, Cell Painting assays are a high-content/high-throughput imaging methods designed to reveal a broad range of cellular phenotypes. Image analysis makes it possible to extract numerical information from cell shape, intensity, texture, granularity, producing what are known as “morphological profiles”. These features combined form morphological profiles offering a window into a variety of biological processes, such as how cells react to genetic modifications, drug exposure, and shifts in their environment (Seal et al. 2025). Extracting biological meaning from these images requires transforming raw features and readouts into clean, comparable profiles. Because of the sheer number of values involved, these results can be extremely large, and specific frameworks are needed to process such morphological profiles correctly.

In this context, Pycytominer is a Python toolkit for processing high dimensional readouts from high-throughput image-based profiling experiments (Serrano et al. 2025).

pycytominer-logo.png.

In this tutorial, you will learn how to run a Pycytominer pipeline using Galaxy. We will follow the different steps explained in the Pycytominer documentation. If you want a more comprehensive explanation of each step, please feel free to visit the main Pycytominer main documentation page or the GitHub repository! each step, please feel free to visit the main Pycytominer main documentation page or the GitHub repository!

Agenda

In this tutorial, we will deal with:

  1. Getting data
    1. Step 1: Aggregate — From Cells to Wells
    2. Step 2: Annotate — Adding Experimental Context
    3. Step 3: Normalize by Removing Technical Variation
    4. Step 4: Feature Selection — Keeping Only Informative Features
    5. Step 5: Consensus — Collapsing Replicates
    6. A full workflow for table readouts processing
  2. Conclusion

Getting data

A synthetic dataset necessary for this tutorial can be created following the instructions of the Pycytominer documentation.

Experimental design:

Property Value
Plates (biological replicates) 1
Wells per plate 6 (2 × DMSO vehicle control, 2 × Compound A, 2 × Compound B)
Cells per well ~100
Total single-cell measurements ~600
Morphological features 11 (across three compartments)

For simplicity, we provide the generated files for this tutorial.

Hands On: Data Upload
  1. If you are logged in, create a new history for this tutorial

    To create a new history simply click the new-history icon at the top of the history panel:

    UI for creating new history

  2. Download the following image-based profiles and import them into your Galaxy history.

    If you are importing the files via URL:

    • Copy the link location
    • Click galaxy-upload Upload at the top of the activity panel

    • Select galaxy-wf-edit Paste/Fetch Data
    • Paste the link(s) into the text field

    • Press Start

    • Close the window

    Galaxy upload link

    If you are importing the files from the shared data library:

    As an alternative to uploading the data from a URL or your computer, the files may also have been made available from a shared data library:

    1. Go into Libraries (left panel)
    2. Navigate to the correct folder as indicated by your instructor.
      • On most Galaxies tutorial data will be provided in a folder named GTN - Material –> Topic Name -> Tutorial Name.
    3. Select the desired files
    4. Click on Add to History galaxy-dropdown near the top and select as Datasets from the dropdown menu
    5. In the pop-up window, choose

      • “Select history”: the history you want to import the data to (or create a new one)
    6. Click on Import

  3. Confirm the datatypes are correct (tabular for both profiles)

    • Click on the galaxy-pencil pencil icon for the dataset to edit its attributes
    • In the central panel, click galaxy-chart-select-data Datatypes tab on the top
    • In the galaxy-chart-select-data Assign Datatype, select datatypes from “New Type” dropdown
      • Tip: you can start typing the datatype into the field to filter the dropdown menu
    • Click the Save button

Step 1: Aggregate — From Cells to Wells

Aggregation collapses single-cell measurements into a single profile per well or per sample by computing a summary statistic (such as the median) across all cells.

Hands On: Aggregate plate readouts with Pycytominer
  1. Aggregate readouts ( Galaxy version 1.6.1+galaxy0) with the following parameters to aggregate readouts:
    • param-file “Input feature-readouts table”: 01_single_cells.tsv file
    • “Aggregation Column”: Select “c1:Metadata_Plate” and “c2:Metadata_Well”
    • “Aggregation function”: Mean
  2. Rename galaxy-pencil the generated file to 01_output_aggregate.tsv.

    • Click on the galaxy-pencil pencil icon for the dataset to edit its attributes
    • In the central panel, change the Name field
    • Click the Save button

  3. Click on the visualise icon galaxy-visualise of the file to visually inspect the image-based profiles using the Tabulator visualisation plugin.

The 600 single-cell measurements are now aggregated into 6 profiles, one per well: for each feature, the values of the ~100 cells in a well are summarized into a single value (here, the mean).

01-aggregate.png.

Step 2: Annotate — Adding Experimental Context

Annotation merges these profiles with experimental metadata (i.e. plate and well identifiers, and other conditions) so each profile is linked to what was done to the cells.

Hands On: Annotate readouts with metadata using Pycytominer
  1. Annotate readouts with metadata ( Galaxy version 1.6.1+galaxy0) with the following parameters:
    • param-file “Input feature-readouts table”: 01_output_aggregate.tsv file
    • “Column describing the wells in the feature-readouts table”: Select “c2:Metadata_Well”
    • param-file “Input platemap table”: 01_platemap.tsv file
    • “Column describing the wells in the platemap”: Select “c1:well_position”
  2. Rename galaxy-pencil the generated file to 02_output_annotated.tsv.
  3. Click on the visualise icon galaxy-visualise of the file to visually inspect the image-based profiles using the Tabulator visualisation plugin.

Three additional columns are now added to the table: “Metadata_treatment”, “Metadata_cell_line” and “Metadata_concentration_um”. Pycytominer adds the Metadata_ prefix to the plate map columns (treatment → Metadata_treatment) to distinguish them from the morphological features. All this information is important to give more context to the data. Metadata_treatment allows us to identify the DMSO control wells used for normalization.

02-annotate.png.

Step 3: Normalize by Removing Technical Variation

Normalization rescales features so they can be compared with each other. Without it, features with large absolute values (e.g. cell area) would dominate any downstream distance calculation, regardless of whether they carry biological signal. A common approach is to standardize each feature against control samples: here, every well is expressed relative to the DMSO control wells. In experiments with several plates or batches, normalizing each plate against its own controls also corrects technical variation, such as differences in staining, imaging conditions or cell density.

Hands On: Normalize readouts with Pycytominer
  1. Normalize readouts ( Galaxy version 1.6.1+galaxy0) with the following parameters:
    • param-file “Input feature-readouts table”: 02_output_annotated.tsv file
    • “Column with normalization values”: Select “c1:Metadata_treatment”
    • “Value”: Type “DMSO”
    • “Normalization method”: Select “Standardize”
  2. Rename galaxy-pencil the generated file to 03_output_normalized.tsv.
  3. Click on the visualise icon galaxy-visualise of the file to visually inspect the image-based profiles using the Tabulator visualization plugin.

With the standardize normalization method, feature becomes a z-score based on the mean and standard deviation of the DMSO wells. So DMSO wells end up around 0, and treated wells show how many standard deviations they differ from the control. All features are now expressed in the same unit (standard deviations from the DMSO control), so they can be compared with each other.

03-normalize.png.

Step 4: Feature Selection — Keeping Only Informative Features

Feature selection removes uninformative or redundant features, such as those with low variance, high correlation with other features, or missing values, yielding a compact and reliable feature set.

Hands On: Select informative features with Pycytominer
  1. Select informative features ( Galaxy version 1.6.1+galaxy0) with the following parameters:
    • param-file “Input feature-readouts table”: ‘03_output_normalized.tsv` file
    • “Operations”: Select “Variance Threshold” and “Blocklist”
  2. Rename galaxy-pencil the generated file to 04_output_features.tsv.
  3. Click on the visualise icon galaxy-visualise of the file to visually inspect the image-based profiles using the Tabulator visualization plugin.

Variance Threshold removes features that barely vary across samples, and Blocklist removes features that are known to be noisy or uninformative in image-based profiling. Thanks to the Variance Threshold operation, the table now has 15 columns instead of 16: Cells_AreaShape_EulerNumber was removed because it has the same value in every well, so it carries no information.

04-features.png.

Step 5: Consensus — Collapsing Replicates

The Compute Consensus tool collapses replicate profiles into one consensus profile per treatment group by computing the mean across all replicates.

Hands On: Compute consensus profiles with Pycytominer
  1. Compute consensus ( Galaxy version 1.6.1+galaxy0) with the following parameters:
    • param-file “Input feature-readouts table”: 04_output_features.tsv` file
    • “Column with unique condition”: Select “c1:Metadata_treatment”, “c2:Metadata_cell_line” and “c3:Metadata_concentration_um”
    • “Reduction operation”: Select “Mean”
  2. Rename galaxy-pencil the generated file to 05_output_consensus.tsv.
  3. Click on the visualize icon galaxy-visualise of the file to visually inspect the image-based profiles using the Tabulator visualization plugin.

The table now has 3 profiles, one per treatment (DMSO, Compound A and Compound B).

05-consensus.png.

A full workflow for table readouts processing

You can now create a workflow from the different Pycytominer steps in your history:

Hands On: Extract Pycytominer workflow from history
  1. Now we can extract the workflow for batch processing:

    1. Clean up your history: remove any failed (red) jobs from your history by clicking on the galaxy-delete button.

      This will make the creation of the workflow easier.

    2. Click on galaxy-history-options (History options) at the top of your history panel and select Extract workflow.

      `Extract Workflow` entry in the history options menu

      The central panel will show the content of the history in reverse order (oldest on top), and you will be able to choose which steps to include in the workflow.

    3. Replace the Workflow name to something more descriptive.

    4. Rename each workflow input in the boxes at the top of the second column.

    5. If there are any steps that shouldn’t be included in the workflow, you can uncheck them in the first column of boxes.

    6. Click on the Create Workflow button near the top.

      You will get a message that the workflow was created.

    • Name it “pycytominer-full-steps”.
    • Uncheck 01_platemap.tsv and 01_single_cells.tsv as inputs (the workflow is supposed to be applied to the image-based profiles directly).
  2. Edit the workflow you just created:

    • Select “Input dataset” from the list of tools. The step param-file 8: Input Dataset appears.
    • Select “Input dataset” from the list of tools. The step param-file 9: Input Dataset appears.
    • Change the “Label” of param-file 8: Input Dataset to input table readouts.
    • Change the “Label” of param-file 9: Input Dataset to input plate metadata.
    • Connect the output of param-file 8: input table readouts to the input of tool 3: Aggregate readouts.
    • Connect the output of param-file 9: input plate metadata to the “Input plate Table” input of tool 4: Annotate readouts with metadata.
    • Mark the results of tool 7: Compute Consensus Profile as the primary outputs of the workflow (by clicking on the checkboxes of the outputs).

You have now a Pycytominer automatized workflow in Galaxy!

06-final-workflow.png.

Conclusion

The following tutorial uses high-content imaging analysis as an example, but the same Pycytominer tools can be used in many other contexts for data wrangling, normalization, and annotation… Find your own solution!