skbio.sequence.Protein#
- class skbio.sequence.Protein(sequence, metadata=None, positional_metadata=None, interval_metadata=None, lowercase=False, validate=True, copy=None)[source]#
Store protein sequence data and optional associated metadata.
- Parameters:
- sequencestr, bytes-like, 1D ndarray (uint8 or ‘|S1’), or
Sequence Characters representing the protein sequence.
- metadatadict, optional
Arbitrary metadata which applies to the entire sequence.
- positional_metadatapd.DataFrame consumable, optional
Arbitrary per-character metadata. For example, quality scores of sequencing reads. Must be able to pass directly to the
pd.DataFrameconstructor.- interval_metadataIntervalMetadata, optional
Arbitrary interval metadata which applies to intervals within the a sequence to store interval features (such as domains of the protein sequence).
- lowercasebool or str, optional
If True, lowercase sequence characters will be converted to uppercase to ensure they are valid IUPAC protein characters. If False (default), characters will not be converted. If a string, in addition to the uppercase conversion, a boolean array indicating which positions were originally lowercase will be stored in the positional metadata under this key.
- validatebool, optional
If True (default), validation will be performed to ensure that all sequence characters are in the IUPAC protein character set. Turning off validation (False) will improve performance. If invalid characters are present, however, there is no guarantee that subsequent operations will retain the expected behavior. Only turn off validation if you are certain that the sequence characters are valid. To store sequence data that is not IUPAC-compliant, use
Sequence.- copybool, optional
Control copying of sequence data. See
Sequencefor details.Added in version 0.7.5.
- sequencestr, bytes-like, 1D ndarray (uint8 or ‘|S1’), or
See also
Notes
According to the IUPAC notation [1] , a protein sequence may contain the following 20 canonical amino acids:
Code
3-letter
Amino acid
AAla
Alanine
CCys
Cysteine
DAsp
Aspartic acid
EGlu
Glutamic acid
FPhe
Phenylalanine
GGly
Glycine
HHis
Histidine
IIle
Isoleucine
KLys
Lysine
LLeu
Leucine
MMet
Methionine
NAsn
Asparagine
PPro
Proline
QGln
Glutamine
RArg
Arginine
SSer
Serine
TThr
Threonine
VVal
Valine
WTrp
Tryptophan
YTyr
Tyrosine
And the following two non-canonical amino acids:
Code
3-letter
Amino acid
OPyl
Pyrrolysine
USec
Selenocysteine
The total of 22 amino acids listed above constitute the definite character set of the
Proteinsequence type.Additionally, the following four degenerate characters are defined, each of which representing two or more amino acids:
Code
3-letter
Amino acids
BAsx
D or N
ZGlx
E or Q
JXle
I or L
XXaa
All 22
Plus one stop character:
*(Ter), and two gap characters:-and..Characters other than the above 29 are not allowed. To include additional characters, you may create a custom alphabet using
GrammaredSequence. Directly modifying the alphabet ofProteinmay break functions that rely on the IUPAC alphabet.It should be noted that some functions do not support certain valid characters. For example, the BLOSUM and PAM substitution matrices do not contain
J(Xle). In such circumstances, unsupported characters will be replaced with the wildchard characterXto represent any of the definite amino acids.References
[1]Cornish-Bowden, A. (1985). Nomenclature for incompletely specified bases in nucleic acid sequences: recommendations 1984. Nucleic Acids Res, 13(9), 3021.
Examples
>>> from skbio import Protein >>> Protein('PAW') Protein -------------------------- Stats: length: 3 has gaps: False has degenerates: False has definites: True has stops: False -------------------------- 0 PAW
Convert lowercase characters to uppercase:
>>> Protein('paW', lowercase=True) Protein -------------------------- Stats: length: 3 has gaps: False has degenerates: False has definites: True has stops: False -------------------------- 0 PAW
Attributes
All valid characters in the alphabet.
Gap character to use when constructing a new gapped sequence.
Characters representing definite states.
Mapping of degenerate to definite characters.
Characters representing gaps in the sequence.
Non-canonical characters.
Return characters representing translation stop codons.
Character representing any other non-gap character in the alphabet.
Attributes (inherited)
Characters in the conventional core alphabet.
Default write format for this object:
fasta.Degenerate characters representing sets of definite characters.
IntervalMetadataobject containing info about interval features.dictcontaining metadata which applies to the entire object.Non-degenerate characters.
Set of observed characters in the sequence.
pd.DataFramecontaining metadata along an axis.Array containing underlying sequence characters.
Methods
Search the biological sequence for motifs.
Determine if the sequence contains one or more stop characters.
Create a new
Proteininstance from a file.Find positions containing stop characters in the protein sequence.
Write an instance of
Proteinto a file.Methods (inherited)
Concatenate an iterable of
Sequenceobjects.Count occurrences of a subsequence in this sequence.
Find positions containing definite characters in the sequence.
Return a new sequence with gap characters removed.
Find positions containing degenerate characters in the sequence.
Compute the distance to another sequence.
Yield all possible definite versions of the sequence.
Generate slices for patterns matched by a regular expression.
Compute frequencies of characters in the sequence.
Find positions containing gaps in the biological sequence.
Determine if sequence contains one or more definite characters.
Determine if sequence contains one or more degenerate characters.
Determine if the sequence contains one or more gap characters.
Determine if the object has interval metadata.
Determine if the object has metadata.
Determine if sequence contains one or more non-degenerate characters.
Determine if the object has positional metadata.
Find position where subsequence first occurs in the sequence.
Yield contiguous subsequences based on included.
Generate k-mers of length k from this sequence.
Return counts of words of length k from this sequence.
Return a case-sensitive string representation of the sequence.
Return count of positions that are the same between two sequences.
Find positions that match with another sequence.
Return count of positions that differ between two sequences.
Find positions that do not match with another sequence.
Find positions containing non-degenerate characters in the sequence.
Replace values in this sequence with a different character.
Convert degenerate and noncanonical characters to alternative characters.
Convert the sequence into indices of characters.
Return regular expression object that accounts for degenerate chars.
Special methods (inherited)
Return truth value (truthiness) of sequence.
Determine if a subsequence is contained in this sequence.
Return a shallow copy of this sequence.
Return a deep copy of this sequence.
Determine if this sequence is equal to another.
__ge__Return self>=value.
Slice this sequence.
__getstate__Helper for pickle.
__gt__Return self>value.
Iterate over positions in this sequence.
__le__Return self<=value.
Return the number of characters in this sequence.
__lt__Return self<value.
Determine if this sequence is not equal to another.
Iterate over positions in this sequence in reverse order.
Return sequence characters as a string.
Details
- alphabet[source]#
All valid characters in the alphabet.
This includes gap, definite, and degenerate characters.
- Returns:
- set
Valid characters.
See also
gap_charsdefinite_charsdegenerate_chars
Notes
This property is derived from
degenerate_chars,definite_chars, andgap_chars.
- default_gap_char[source]#
Gap character to use when constructing a new gapped sequence.
This character is used when it is necessary to represent gap characters in a new sequence. For example, a majority consensus sequence will use this character to represent gaps.
- Returns:
- str or None
Default gap character, or None if gaps are not defined.
See also
Notes
When a subclass defines a non-empty
gap_charswithout defining this property, the first gap character in sorted order will be designated as the default gap character during class creation.
- definite_chars[source]#
Characters representing definite states.
- Returns:
- set
Definite characters.
Notes
This character set is the minimum requirement for creating a subclass of
GrammaredSequence.
- degenerate_map[source]#
Mapping of degenerate to definite characters.
- Returns:
- dict of set
Mapping of each degenerate character to the set of definite characters it represents. Default is an empty dictionary.
Notes
Each degenerate character may represent an arbitrary number of definite characters.
- gap_chars[source]#
Characters representing gaps in the sequence.
- Returns:
- set
Characters defined as gaps. Default is an empty set.
- noncanonical_chars[source]#
Non-canonical characters.
They are definite characters outside the conventional core alphabet of a sequence type.
- Returns:
- set
Non-canonical characters. Default is an empty set.
See also
definite_charscanonical_chars
Notes
This character set serves as an exclusion from definite characters to obtain canonical characters.