Skip to content

Bug in modkit extract full, canonical_base does not match reference sequence #715

Description

@awjga

Hi,

The results of my data after running modkit extract full does not match my reference.

From my reference, position 138 and 140 is T and 139 is C.
However, from the results of modkit extract full some read that contains position139 is identified as T in the out.
I see the same issue across the whole transcript, some positions are wrongly identified as T.

Looking at some of the reads and comparing the query_kmer with the reference sequence, it seems to be caused by an indels that is unmapped to the reference.
This is part of my reference:
135-144 (starts from 1)
cgcctctggc

For my downstream analysis, can I just discard the positions that have the wrong reference?

Below is the output of my results:

<style> </style>
5f0af38a-7492-48aa-a3de-1ed2ee404383 181 138 gene + + + 2 62 9 3149 3243 0.9863281 17802 13 . TCGCCTCTGGC T T FALSE 0 3149
f41ca671-d4c5-42d6-8ee6-1b954a8d964f 127 139 gene + + + 9 16 22 3149 3140 0.9082031 17802 7 . CGCCCTTGGCC T T FALSE 0 3149
ad238b5c-a836-478d-b239-a65408a0af2c 128 139 gene + + + 1 39 9 3149 3185 0.9042969 17802 11 . CGCCTTCGGCC T T FALSE 0 3149
bffc67f7-c72c-4a60-a7b4-cb0b857fb67d 129 139 gene + + + 3 42 12 3149 3171 0.9902344 17802 9 . CGCCTTCGGCC T T FALSE 0 3149
efa15c98-4958-40a0-9aef-d730efe32625 63 140 gene + + + 0 39 76 3149 3079 0.9941406 17802 31 . GCCTCTGGCCC T T FALSE 0 3149

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions