skbio.sequence.GrammaredSequence#

class skbio.sequence.GrammaredSequence(sequence, metadata=None, positional_metadata=None, interval_metadata=None, lowercase=False, validate=True, copy=None)[source]#

Store sequence data conforming to a character set.

This class is intended to be inherited from to create grammared sequences with custom alphabets. It is an abstract base class (ABC) that cannot be directly instantiated.

Raises:
ValueError

If sequence characters are not in the character set [1].

See also

DNA
RNA
Protein

References

[1]

Cornish-Bowden, A. (1985). Nomenclature for incompletely specified bases in nucleic acid sequences: recommendations 1984. Nucleic Acids Res, 13(9), 3021.

Examples

GrammaredSequence can be subclassed to create custom sequence types.

A minimal alphabet

This example demonstrates RY-coding: Representing a nucleotide sequence with just two states: R (purines, including A and G) and Y (pyrimidines, including C and T/U). This technique reduces the alphabet to binary thus facilitating computation, and has practical benefits in phylogenetic analysis.

A minimum RY sequence type only needs to declare its definite characters: R and Y (even though they are degenerate characters in the IUPAC DNA alphabet).

>>> from skbio.sequence import GrammaredSequence
>>> from skbio.util import classproperty
>>> class RYSequence(GrammaredSequence):
...     @classproperty
...     def definite_chars(cls):
...         return set("RY")

The new type validates its alphabet just like other grammared sequences:

>>> seq = RYSequence("RYYRYR")
>>> str(seq)
'RYYRYR'
>>> seq.has_definites()
True
>>> seq = RYSequence("ACGT")
Traceback (most recent call last):
    ...
ValueError: Invalid characters in sequence: ['A', 'C', 'G', 'T']...

Choice of characters spans ASCII code points 0-127. Even unprintable characters are valid. The following example uses 0 and 1 as characters.

>>> class Sequence01(GrammaredSequence):
...     @classproperty
...     def definite_chars(cls):
...         return {chr(0), chr(1)}
>>> import numpy as np
>>> seq = Sequence01(np.array([0, 1, 1, 0, 0, 0, 1], dtype=np.uint8))
>>> seq.values.tobytes()
b'\x00\x01\x01\x00\x00\x00\x01'

Enriching grammar

More specialized grammars can additionally define gap characters, degenerate characters, a wildcard, and other properties when they are useful. The following example adds gap character ^ to the sequence type.

>>> class RYSequence(GrammaredSequence):
...     @classproperty
...     def definite_chars(cls):
...         return set("RY")
...
...     @classproperty
...     def gap_chars(cls):
...         return set("^")

Then one can perform pairwise alignment of two RY sequences and construct a tabular alignment. (Note: pair_align does not need a defined gap character, but TabularMSA does.)

>>> from skbio import TabularMSA
>>> from skbio.alignment import pair_align
>>> seq1 = RYSequence('YRYRRRYYRY')
>>> seq2 = RYSequence('RYRRRYRYYY')
>>> path = pair_align(seq1, seq2).paths[0]
>>> msa = TabularMSA.from_path_seqs(path, (seq1, seq2))
>>> msa
TabularMSA[RYSequence]
----------------------
Stats:
    sequence count: 2
    position count: 12
----------------------
YRYRRRYYRY^^
^RYRRRY^RYYY

Adding utilities

You can add custom class properties and methods to perform specific operations. The following code lets one construct an RY sequence from a DNA sequence. Only the four canonical nucleotides are recognized. Otherwise, an error will be raised.

>>> class RYSequence(GrammaredSequence):
...     @classproperty
...     def definite_chars(cls):
...         return set("RY")
...
...     @classproperty
...     def code_map(cls):
...         return bytes.maketrans(b"ACGT", b"RYRY")
...
...     @classmethod
...     def from_dna(cls, seq):
...         return cls(seq.values.tobytes().translate(cls.code_map))
>>> from skbio import DNA
>>> RYSequence.from_dna(DNA('GAATTC'))
RYSequence
--------------------------
Stats:
    length: 6
    has gaps: False
    has degenerates: False
    has definites: True
--------------------------
0 RRRYYY

Extending existing sequence types

One may subclass an existing sequence type and modify its grammar and operations. The following example creates a methylated DNA type by introducing a new definite character: Z, representing 5-methylcytosine (5mC). It also introduces methods for demethylation of Z to C, and for bisulfite treatment to preserve the methylation state for DNA sequencing.

>>> class MethylatedDNA(DNA):
...     @classproperty
...     def definite_chars(cls):
...         return DNA.definite_chars | {'Z'}
...
...     @classproperty
...     def complement_map(cls):
...         return DNA.complement_map | {'Z': 'G'}
...
...     def demethylate(self):
...         chars = self.values.copy()
...         chars[chars == b'Z'] = b'C'
...         return DNA(chars)
...
...     def bisulfite_convert(self):
...         chars = self.values.copy()
...         chars[chars == b'C'] = b'T'
...         chars[chars == b'Z'] = b'C'
...         return DNA(chars)
...
...     def transcribe(self):
...         return self.demethylate().transcribe()

The subclass accepts both ordinary DNA characters and the added 5mC state:

>>> seq = MethylatedDNA("ACZCG")
>>> str(seq)
'ACZCG'

It also provides a conversion specific to this sequence type. In this simplified example, unmethylated cytosine is converted to uracil (read as thymine during sequencing) while 5mC is retained as cytosine:

>>> converted = seq.bisulfite_convert()
>>> str(converted)
'ATCTG'
>>> type(converted) is DNA
True

MethylatedDNA inherits the rest of the DNA interface, but adding a character can require revisiting inherited operations. Here, complement_map and transcribe are modified to establish that 5mC should be considered as cytosine in these operations.

Other inherited operations may likewise need to be reviewed or overridden. For example, a subclass should decide how a newly introduced nucleotide impacts GC-content calculation, degeneracy handling, and any other operation whose meaning depends on the alphabet. Therefore, be very careful with extending existing sequence types, and consider limiting downstream analysis to what you can oversee.

Attributes

alphabet

All valid characters in the alphabet.

canonical_chars

Characters in the conventional core alphabet.

default_gap_char

Gap character to use when constructing a new gapped sequence.

definite_chars

Characters representing definite states.

degenerate_chars

Degenerate characters representing sets of definite characters.

degenerate_map

Mapping of degenerate to definite characters.

gap_chars

Characters representing gaps in the sequence.

noncanonical_chars

Non-canonical characters.

nondegenerate_chars

Non-degenerate characters.

wildcard_char

Character representing any other non-gap character in the alphabet.

Attributes (inherited)

default_write_format

Default write format for this object: fasta.

interval_metadata

IntervalMetadata object containing info about interval features.

metadata

dict containing metadata which applies to the entire object.

observed_chars

Set of observed characters in the sequence.

positional_metadata

pd.DataFrame containing metadata along an axis.

values

Array containing underlying sequence characters.

Methods

definites

Find positions containing definite characters in the sequence.

degap

Return a new sequence with gap characters removed.

degenerates

Find positions containing degenerate characters in the sequence.

expand_degenerates

Yield all possible definite versions of the sequence.

find_motifs

Search the biological sequence for motifs.

gaps

Find positions containing gaps in the biological sequence.

has_definites

Determine if sequence contains one or more definite characters.

has_degenerates

Determine if sequence contains one or more degenerate characters.

has_gaps

Determine if the sequence contains one or more gap characters.

has_nondegenerates

Determine if sequence contains one or more non-degenerate characters.

nondegenerates

Find positions containing non-degenerate characters in the sequence.

to_definites

Convert degenerate and noncanonical characters to alternative characters.

to_regex

Return regular expression object that accounts for degenerate chars.

Methods (inherited)

concat

Concatenate an iterable of Sequence objects.

count

Count occurrences of a subsequence in this sequence.

distance

Compute the distance to another sequence.

find_with_regex

Generate slices for patterns matched by a regular expression.

frequencies

Compute frequencies of characters in the sequence.

has_interval_metadata

Determine if the object has interval metadata.

has_metadata

Determine if the object has metadata.

has_positional_metadata

Determine if the object has positional metadata.

index

Find position where subsequence first occurs in the sequence.

iter_contiguous

Yield contiguous subsequences based on included.

iter_kmers

Generate k-mers of length k from this sequence.

kmer_frequencies

Return counts of words of length k from this sequence.

lowercase

Return a case-sensitive string representation of the sequence.

match_frequency

Return count of positions that are the same between two sequences.

matches

Find positions that match with another sequence.

mismatch_frequency

Return count of positions that differ between two sequences.

mismatches

Find positions that do not match with another sequence.

read

Create a new GrammaredSequence instance from a file.

replace

Replace values in this sequence with a different character.

to_indices

Convert the sequence into indices of characters.

write

Write an instance of GrammaredSequence to a file.

Special methods (inherited)

__bool__

Return truth value (truthiness) of sequence.

__contains__

Determine if a subsequence is contained in this sequence.

__copy__

Return a shallow copy of this sequence.

__deepcopy__

Return a deep copy of this sequence.

__eq__

Determine if this sequence is equal to another.

__ge__

Return self>=value.

__getitem__

Slice this sequence.

__getstate__

Helper for pickle.

__gt__

Return self>value.

__iter__

Iterate over positions in this sequence.

__le__

Return self<=value.

__len__

Return the number of characters in this sequence.

__lt__

Return self<value.

__ne__

Determine if this sequence is not equal to another.

__reversed__

Iterate over positions in this sequence in reverse order.

__str__

Return sequence characters as a string.

Details

alphabet[source]#

All valid characters in the alphabet.

This includes gap, definite, and degenerate characters.

Returns:
set

Valid characters.

Notes

This property is derived from degenerate_chars, definite_chars, and gap_chars.

canonical_chars[source]#

Characters in the conventional core alphabet.

Such as the four nucleotides and the 20 basic amino acids.

Added in version 0.7.5.

Returns:
set

Canonical characters.

Notes

This property is derived by excluding noncanonical_chars from definite_chars.

default_gap_char[source]#

Gap character to use when constructing a new gapped sequence.

This character is used when it is necessary to represent gap characters in a new sequence. For example, a majority consensus sequence will use this character to represent gaps.

Returns:
str or None

Default gap character, or None if gaps are not defined.

See also

gap_chars

Notes

When a subclass defines a non-empty gap_chars without defining this property, the first gap character in sorted order will be designated as the default gap character during class creation.

definite_chars[source]#

Characters representing definite states.

Returns:
set

Definite characters.

Notes

This character set is the minimum requirement for creating a subclass of GrammaredSequence.

degenerate_chars[source]#

Degenerate characters representing sets of definite characters.

Returns:
set

Degenerate characters.

See also

degenerate_map

Notes

This property is derived from degenerate_map.

degenerate_map[source]#

Mapping of degenerate to definite characters.

Returns:
dict of set

Mapping of each degenerate character to the set of definite characters it represents. Default is an empty dictionary.

Notes

Each degenerate character may represent an arbitrary number of definite characters.

gap_chars[source]#

Characters representing gaps in the sequence.

Returns:
set

Characters defined as gaps. Default is an empty set.

noncanonical_chars[source]#

Non-canonical characters.

They are definite characters outside the conventional core alphabet of a sequence type.

Returns:
set

Non-canonical characters. Default is an empty set.

Notes

This character set serves as an exclusion from definite characters to obtain canonical characters.

nondegenerate_chars[source]#

Non-degenerate characters.

Returns:
set

Non-degenerate characters.

Warning

nondegenerate_chars is deprecated as of 0.5.0. It has been renamed to definite_chars.

See also

definite_chars
wildcard_char[source]#

Character representing any other non-gap character in the alphabet.

Returns:
str of length 1

Wildcard character. Default is None. When set, it must be a definite or degenerate character in the alphabet.