skbio.sequence.GrammaredSequence#
- class skbio.sequence.GrammaredSequence(sequence, metadata=None, positional_metadata=None, interval_metadata=None, lowercase=False, validate=True, copy=None)[source]#
Store sequence data conforming to a character set.
This class is intended to be inherited from to create grammared sequences with custom alphabets. It is an abstract base class (ABC) that cannot be directly instantiated.
- Raises:
- ValueError
If sequence characters are not in the character set [1].
References
[1]Cornish-Bowden, A. (1985). Nomenclature for incompletely specified bases in nucleic acid sequences: recommendations 1984. Nucleic Acids Res, 13(9), 3021.
Examples
GrammaredSequencecan be subclassed to create custom sequence types.A minimal alphabet
This example demonstrates RY-coding: Representing a nucleotide sequence with just two states:
R(purines, includingAandG) andY(pyrimidines, includingCandT/U). This technique reduces the alphabet to binary thus facilitating computation, and has practical benefits in phylogenetic analysis.A minimum RY sequence type only needs to declare its definite characters:
RandY(even though they are degenerate characters in the IUPAC DNA alphabet).>>> from skbio.sequence import GrammaredSequence >>> from skbio.util import classproperty
>>> class RYSequence(GrammaredSequence): ... @classproperty ... def definite_chars(cls): ... return set("RY")
The new type validates its alphabet just like other grammared sequences:
>>> seq = RYSequence("RYYRYR") >>> str(seq) 'RYYRYR' >>> seq.has_definites() True
>>> seq = RYSequence("ACGT") Traceback (most recent call last): ... ValueError: Invalid characters in sequence: ['A', 'C', 'G', 'T']...
Choice of characters spans ASCII code points 0-127. Even unprintable characters are valid. The following example uses 0 and 1 as characters.
>>> class Sequence01(GrammaredSequence): ... @classproperty ... def definite_chars(cls): ... return {chr(0), chr(1)}
>>> import numpy as np >>> seq = Sequence01(np.array([0, 1, 1, 0, 0, 0, 1], dtype=np.uint8)) >>> seq.values.tobytes() b'\x00\x01\x01\x00\x00\x00\x01'
Enriching grammar
More specialized grammars can additionally define gap characters, degenerate characters, a wildcard, and other properties when they are useful. The following example adds gap character
^to the sequence type.>>> class RYSequence(GrammaredSequence): ... @classproperty ... def definite_chars(cls): ... return set("RY") ... ... @classproperty ... def gap_chars(cls): ... return set("^")
Then one can perform pairwise alignment of two RY sequences and construct a tabular alignment. (Note:
pair_aligndoes not need a defined gap character, butTabularMSAdoes.)>>> from skbio import TabularMSA >>> from skbio.alignment import pair_align >>> seq1 = RYSequence('YRYRRRYYRY') >>> seq2 = RYSequence('RYRRRYRYYY') >>> path = pair_align(seq1, seq2).paths[0] >>> msa = TabularMSA.from_path_seqs(path, (seq1, seq2)) >>> msa TabularMSA[RYSequence] ---------------------- Stats: sequence count: 2 position count: 12 ---------------------- YRYRRRYYRY^^ ^RYRRRY^RYYY
Adding utilities
You can add custom class properties and methods to perform specific operations. The following code lets one construct an RY sequence from a DNA sequence. Only the four canonical nucleotides are recognized. Otherwise, an error will be raised.
>>> class RYSequence(GrammaredSequence): ... @classproperty ... def definite_chars(cls): ... return set("RY") ... ... @classproperty ... def code_map(cls): ... return bytes.maketrans(b"ACGT", b"RYRY") ... ... @classmethod ... def from_dna(cls, seq): ... return cls(seq.values.tobytes().translate(cls.code_map))
>>> from skbio import DNA >>> RYSequence.from_dna(DNA('GAATTC')) RYSequence -------------------------- Stats: length: 6 has gaps: False has degenerates: False has definites: True -------------------------- 0 RRRYYY
Extending existing sequence types
One may subclass an existing sequence type and modify its grammar and operations. The following example creates a methylated DNA type by introducing a new definite character:
Z, representing 5-methylcytosine (5mC). It also introduces methods for demethylation ofZtoC, and for bisulfite treatment to preserve the methylation state for DNA sequencing.>>> class MethylatedDNA(DNA): ... @classproperty ... def definite_chars(cls): ... return DNA.definite_chars | {'Z'} ... ... @classproperty ... def complement_map(cls): ... return DNA.complement_map | {'Z': 'G'} ... ... def demethylate(self): ... chars = self.values.copy() ... chars[chars == b'Z'] = b'C' ... return DNA(chars) ... ... def bisulfite_convert(self): ... chars = self.values.copy() ... chars[chars == b'C'] = b'T' ... chars[chars == b'Z'] = b'C' ... return DNA(chars) ... ... def transcribe(self): ... return self.demethylate().transcribe()
The subclass accepts both ordinary DNA characters and the added 5mC state:
>>> seq = MethylatedDNA("ACZCG") >>> str(seq) 'ACZCG'
It also provides a conversion specific to this sequence type. In this simplified example, unmethylated cytosine is converted to uracil (read as thymine during sequencing) while 5mC is retained as cytosine:
>>> converted = seq.bisulfite_convert() >>> str(converted) 'ATCTG' >>> type(converted) is DNA True
MethylatedDNAinherits the rest of theDNAinterface, but adding a character can require revisiting inherited operations. Here,complement_mapandtranscribeare modified to establish that 5mC should be considered as cytosine in these operations.Other inherited operations may likewise need to be reviewed or overridden. For example, a subclass should decide how a newly introduced nucleotide impacts GC-content calculation, degeneracy handling, and any other operation whose meaning depends on the alphabet. Therefore, be very careful with extending existing sequence types, and consider limiting downstream analysis to what you can oversee.
Attributes
All valid characters in the alphabet.
Characters in the conventional core alphabet.
Gap character to use when constructing a new gapped sequence.
Characters representing definite states.
Degenerate characters representing sets of definite characters.
Mapping of degenerate to definite characters.
Characters representing gaps in the sequence.
Non-canonical characters.
Non-degenerate characters.
Character representing any other non-gap character in the alphabet.
Attributes (inherited)
Default write format for this object:
fasta.IntervalMetadataobject containing info about interval features.dictcontaining metadata which applies to the entire object.Set of observed characters in the sequence.
pd.DataFramecontaining metadata along an axis.Array containing underlying sequence characters.
Methods
Find positions containing definite characters in the sequence.
Return a new sequence with gap characters removed.
Find positions containing degenerate characters in the sequence.
Yield all possible definite versions of the sequence.
Search the biological sequence for motifs.
Find positions containing gaps in the biological sequence.
Determine if sequence contains one or more definite characters.
Determine if sequence contains one or more degenerate characters.
Determine if the sequence contains one or more gap characters.
Determine if sequence contains one or more non-degenerate characters.
Find positions containing non-degenerate characters in the sequence.
Convert degenerate and noncanonical characters to alternative characters.
Return regular expression object that accounts for degenerate chars.
Methods (inherited)
Concatenate an iterable of
Sequenceobjects.Count occurrences of a subsequence in this sequence.
Compute the distance to another sequence.
Generate slices for patterns matched by a regular expression.
Compute frequencies of characters in the sequence.
Determine if the object has interval metadata.
Determine if the object has metadata.
Determine if the object has positional metadata.
Find position where subsequence first occurs in the sequence.
Yield contiguous subsequences based on included.
Generate k-mers of length k from this sequence.
Return counts of words of length k from this sequence.
Return a case-sensitive string representation of the sequence.
Return count of positions that are the same between two sequences.
Find positions that match with another sequence.
Return count of positions that differ between two sequences.
Find positions that do not match with another sequence.
Create a new
GrammaredSequenceinstance from a file.Replace values in this sequence with a different character.
Convert the sequence into indices of characters.
Write an instance of
GrammaredSequenceto a file.Special methods (inherited)
Return truth value (truthiness) of sequence.
Determine if a subsequence is contained in this sequence.
Return a shallow copy of this sequence.
Return a deep copy of this sequence.
Determine if this sequence is equal to another.
__ge__Return self>=value.
Slice this sequence.
__getstate__Helper for pickle.
__gt__Return self>value.
Iterate over positions in this sequence.
__le__Return self<=value.
Return the number of characters in this sequence.
__lt__Return self<value.
Determine if this sequence is not equal to another.
Iterate over positions in this sequence in reverse order.
Return sequence characters as a string.
Details
- alphabet[source]#
All valid characters in the alphabet.
This includes gap, definite, and degenerate characters.
- Returns:
- set
Valid characters.
See also
Notes
This property is derived from
degenerate_chars,definite_chars, andgap_chars.
- canonical_chars[source]#
Characters in the conventional core alphabet.
Such as the four nucleotides and the 20 basic amino acids.
Added in version 0.7.5.
- Returns:
- set
Canonical characters.
See also
Notes
This property is derived by excluding
noncanonical_charsfromdefinite_chars.
- default_gap_char[source]#
Gap character to use when constructing a new gapped sequence.
This character is used when it is necessary to represent gap characters in a new sequence. For example, a majority consensus sequence will use this character to represent gaps.
- Returns:
- str or None
Default gap character, or None if gaps are not defined.
See also
Notes
When a subclass defines a non-empty
gap_charswithout defining this property, the first gap character in sorted order will be designated as the default gap character during class creation.
- definite_chars[source]#
Characters representing definite states.
- Returns:
- set
Definite characters.
Notes
This character set is the minimum requirement for creating a subclass of
GrammaredSequence.
- degenerate_chars[source]#
Degenerate characters representing sets of definite characters.
- Returns:
- set
Degenerate characters.
See also
Notes
This property is derived from
degenerate_map.
- degenerate_map[source]#
Mapping of degenerate to definite characters.
- Returns:
- dict of set
Mapping of each degenerate character to the set of definite characters it represents. Default is an empty dictionary.
Notes
Each degenerate character may represent an arbitrary number of definite characters.
- gap_chars[source]#
Characters representing gaps in the sequence.
- Returns:
- set
Characters defined as gaps. Default is an empty set.
- noncanonical_chars[source]#
Non-canonical characters.
They are definite characters outside the conventional core alphabet of a sequence type.
- Returns:
- set
Non-canonical characters. Default is an empty set.
See also
Notes
This character set serves as an exclusion from definite characters to obtain canonical characters.