Skip to content

Transcriptomic POC - #49

Draft
TamaraNaboulsi wants to merge 5 commits into
mainfrom
poc/transcriptomic
Draft

TamaraNaboulsi wants to merge 5 commits into
mainfrom
poc/transcriptomic

Conversation

@TamaraNaboulsi

Copy link
Copy Markdown
Member

This is the (very basic) initial change to handle transcriptomic tracks. The changes here are by no means final and are open to changing or scraping altogether based on better understanding of the project and/or better implementation ideas. This is how I understood the starting point of the backend side and it is limited to track discovery at this point. It includes the addition of a new model to represent the transcriptomic configuration, an update to the endpoint track_categories to return a new category of tracks (under the parent category of Genomic), a new endpoint which returns the configuration itself, and some seeding/preparation code that adds the POC pig data into an existing sqlite db.

The flow I envisioned:

  • Genome browser loads for a specific genome, track_categories is called and returns a new section representing the transcriptomic tracks. In contrary to the other (called 'inline' in this POC) tracks, I did not see a need to return all the tracks at this point as the UI doesn't need to list them on the right side panel like the others. The idea is that for users who aren't interested in transcriptomic, this saves a bit of loading.
  • If users click on the 'configure' in the transcriptomic section, the new endpoint transcriptomic/[genome-uuid]/configuration is called and that basically returns the json data. This returns everything for now, with the possibility of later choosing what we want returned at this point, especially for genomes with huge amounts of tracks.
  • The UI would then use this json data to draw the appropriate checkboxes in both the main and the filters sections.

Notes & caveats:

  • I am treating the 'tissue' as any other metadata here, as this grouping might be subject to change according to the recent meeting with genebuild and the summary document created after. The idea is that some configuration will be included in the json file in the future to specify what grouping the UI should use. For this POC we can assume the grouping is by 'tissue'. I believe (and I think it's confirmed in the documents so far) that the main categorization is biosample-run, with everything else being another metadata field and it's all just a way to present this data on the UI.
  • My understanding is that we have 2 types of filters here: the ones that come from metadata (tissue, age, sex, etc..) and the ones that require filtering of the track file itself. The first type reduces the selection of tracks but can easily be known from the json and will lead to specific track files. The fact that the files are divided by run (and are not 1 big file) makes this a bit simpler as the metadata themselves are related to the file as a whole. The second type of filters however will need to be passed along to the file parser so it can know what to extract from the files themselves. We still need some clarification on these kinds of filters.
  • To make things a bit quicker, the seeding script and the related transcriptomic.py were written by codex, and they include a bit of preparation of the json data, like including counts for the filters.
  • This POC has some minimal release handling through datasets, using the same logic currently used for the other tracks. It assumes the pig json handed over to be belonging to a single dataset. Would love some input on this from Daniel as my understanding of this is a bit lacking.

Next steps:

  • Input from Andres on what he needs from the backend. Currently the configuration endpoint returns the whole selection from the json. But since this will probably be added into an ingestion process (track api pipeline?), we can do some preprocessing at that point to return the selection a bit more structured and organized. Of course this assumes we know what the structure is supposed to look like.
  • Input from Daniel about the role of metadata and maybe some discussion on versioning.
  • Come up with questions to go back to genebuild with.

Testing:
I have tested this code on the latest (release 30) tracks.sqlite3 db provided by Automation, applying the seeding script on it to populate it with the pig data. On a local implementation of the endpoint, these are currently the returned values.

http://localhost:8000/track_categories/da38d82b-df50-419b-af74-886463b7bfa3
{
  "track_categories": [
    {
      "label": "Assembly",
      "track_category_id": "assembly",
      "type": "Genomic",
      "track_list": [
        {
          "track_id": "7d77026d-5c93-4549-8040-e44f97892e90",
          ...
        },
        ...
      ]
    },
    {
      "label": "Genes & transcripts",
      "track_category_id": "genes-transcripts",
      "type": "Genomic",
      "track_list": [
        {
          "track_id": "c38cd21f-dd61-48a4-b522-5291c2cdfdda",
          ...
        },
        ...
      ]
    },
    {
      "label": "Regulation",
      "track_category_id": "regulatory-features",
      "type": "Regulation",
      "track_list": [
        {
          "track_id": "4650ecb2-dead-4489-96b2-6d968dca3535",
          ...
        }
      ]
    },
    {
      "label": "Short variants by resource",
      "track_category_id": "short-variants",
      "type": "Variation",
      "track_list": [
        {
          "track_id": "15e088de-3fdb-4233-86eb-54a1e285d09b",
          ...
        }
      ]
    },
    {
      "label": "Transcriptomic data",
      "track_category_id": "transcriptomic",
      "type": "Genomic",
      "track_list": [],
      "configuration": {
        "href": "http://localhost:8000/transcriptomic/da38d82b-df50-419b-af74-886463b7bfa3/configuration?dataset_id=7cf49ddb-4de1-4e9c-9f6a-4a9552f4db5d"
      }
    }
  ]
}
http://localhost:8000/transcriptomic/da38d82b-df50-419b-af74-886463b7bfa3/configuration
{
  "track_count": 94,
  "filters": {
    "age": [
      {
        "value": "1 day",
        "count": 1
      },
      {
        "value": "1 month",
        "count": 2
      },
      {
        "value": "10 month",
        "count": 2
      },
      ...
    ],
    "breed": [
      {
        "value": "dly",
        "count": 1
      },
      {
        "value": "large white pig",
        "count": 2
      },
      {
        "value": "sus scrofa domesticus",
        "count": 3
      }
    ],
    "developmental_stage": [
      {
        "value": "35 days post gestation",
        "count": 3
      },
      {
        "value": "piglet",
        "count": 6
      },
      {
        "value": "adult",
        "count": 7
      },
      ...
    ],
    "organism": [
      {
        "value": "sus scrofa",
        "count": 88
      },
      {
        "value": "sus scrofa domesticus",
        "count": 6
      }
    ],
    "sex": [
      {
        "value": "female",
        "count": 19
      },
      {
        "value": "male",
        "count": 43
      },
      {
        "value": "neuter",
        "count": 8
      }
    ],
    "tissue": [
      {
        "value": "adipocyte",
        "count": 2
      },
      {
        "value": "adipose",
        "count": 5
      },
      {
        "value": "colon",
        "count": 14
      },
      {
        "value": "fat",
        "count": 4
      },
      ...
    ],
    ... [MORE FILTERS] ...
  },
  "selections": [
    {
      "schema_version": "transcriptomic-selection-0.1",
      "selection_id": "DRR169425",
      "selection_level": "run",
      "data_type": "rnaseq",
      "display_label": "DRR169425",
      "parent_sample_id": "SAMD00156644",
      "target_genome": {
        "species": "Sus scrofa",
        "assembly": "GCA_000003025.6"
      },
      "metadata": {
        "organism": "Sus scrofa",
        "sex": "female",
        "tissue": "MGC",
        "library_strategy": "RNA-Seq"
      },
      "provenance": {
        "biosample_accession": "SAMD00156644",
        "run_accessions": [
          "DRR169425"
        ]
      },
      "track_products": [
        {
          "track_type": "rnaseq_coverage",
          "format": "BigWig",
          "status": "available",
          "track_file": "/hps/nobackup/flicek/ensembl/genebuild/jackt/expression_tracks/bigwig/DRR169425_Aligned.sortedByCoord.out.bw",
          "md5": "e8f49128c2935c2c59207f8ae0bf824a",
          "track_id": "a6ad48fb-4394-55c7-8f75-e23d7a717621"
        }
      ]
    },
    ...
  ]
}

@veidenberg

veidenberg commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

In general looks good. Few small suggestions:

  • configuration can be removed from track_categories, since the knowledge how/when to call the config endpoint should belong to client (like for all the other client-serving APIs). Altrenatively, we can add something like configuration_type=transcriptomic, if we want to support more config-driven track types in the future (and config endpoint could then live at /api/configuration/transcriptomic, or api/tracks/configuration/transcriptomic if it's part of Track API). Also note that the client does not currently support dataset IDs (the latest dataset is always returned), this can be added later when we implement archives.
  • track_file should be removed from the config payload. Client doesn't need it and it exposes full filepaths in the web browser. The filepath should be part of /api/tracks/<track_id> payload.
  • All the transcriptomic tracks should be registered in Track API, like the other track types. This will give Genome Browser the necessary information to draw it (datafile path, zoom levels, rendering program names, etc.; optionally also description/url for ens-client, but it could also come from config endpoint). However, transcriptomic track records need some special tratment to exclude these from track_categories payload Already implemented. Nice 👍

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants